Python技术迷

别再只会爬虫了!Python在Web开发里原来这么能打

别说,你们是不是也有过这种错觉:觉得 Python 学到会写爬虫脚本,能把网页扒下来存成 Excel,好像就“能接私活、能挣钱了”。

结果真干两单,客户下一句就把人问懵了——“这个数据能不能做个网页给我们内部用?”、“能不能让销售自己查?”……爬虫会了半天,发现最难的那块其实是 Web。

我前段时间在公司楼下抽烟,小李又来一句:哥,我会写爬虫了,Web 这块是不是得再学个 Java 啊?我当场就笑了,Python 在 Web 开发这块,可一点都不菜,真要用好了,根本不止“会爬虫”这么简单。

就趁这个事儿,跟你慢慢聊聊哈。

你想象一下一个完整一点的小系统:

  • 每天从多个网站爬行业新闻
  • 爬完按关键字打标签
  • 做个简单后台,运营能改标题、加备注
  • 对外给销售一个网址,输入关键字就能查

如果你只会爬虫,你能做到哪一步? 大概率是:把数据丢到本地 CSV,然后发给别人。

但如果你把 Python Web 搞明白了,这一整套,其实完全可以用一门语言打穿:爬虫 + Web API + 页面渲染 + 简单后台。

用 FastAPI 先撑起一块“门面”

说 Web 开发很玄乎?先整出一个“能跑起来”的服务再说,心理压力会小很多。

比如用 FastAPI,写个最小可用的 Web 服务,大概长这样:

from fastapi import FastAPI
from pydantic import BaseModel

app = FastAPI()

# 模拟一个内存数据库
ARTICLES: list[dict] = []

classArticleIn(BaseModel):
    title: str
    url: str
    tags: list[str] = []

@app.get("/")
defindex():
return {"msg": "服务还活着呢"}

@app.post("/articles")
defcreate_article(item: ArticleIn):
    item_dict = item.dict()
    item_dict["id"] = len(ARTICLES) + 1
    ARTICLES.append(item_dict)
return item_dict

@app.get("/articles")
deflist_articles(tag: str | None = None):
ifnot tag:
return ARTICLES
return [a for a in ARTICLES if tag in a["tags"]]

你看,这里面几个点其实就是 Web 开发最核心的东西:

  • 有一个“应用对象”(app = FastAPI())
  • 有“路由”——@app.get、@app.post 这些
  • 请求进来,调用一个 Python 函数,返回数据(这里是 JSON)

这玩意一跑,你的爬虫脚本就能从“本地脚本”升级为“对外服务”了。比如你爬完就 POST 一下 /articles,别人就可以通过 /articles?tag=xxx 来查。

这就是 Web:把你的代码变成“别人能随时来调用的东西”。

爬虫怎么跟 Web 连起来?

很多人写爬虫就是一个 main.py,定时跑一下,抓完写文件就结束。你把思路稍微转一下,其实可以变成这样:

  • 爬虫逻辑写在一个函数里
  • Web 提供一个接口,“点一下按钮”就触发爬虫
  • 任务记录下来,状态能查

比如非常简化一版的“手动触发爬虫”:

import httpx
from fastapi import FastAPI

app = FastAPI()

JOBS: dict[int, dict] = {}
JOB_ID = 0

asyncdefcrawl(keyword: str) -> list[dict]:
# 这地方本来应该是真爬,这里随便模拟一下
asyncwith httpx.AsyncClient() as client:
        resp = await client.get("https://httpbin.org/get", params={"q": keyword})
return [{"title": f"{keyword} - demo", "raw": resp.json()}]

@app.post("/crawl")
asyncdefstart_crawl(keyword: str):
global JOB_ID
    JOB_ID += 1
    job_id = JOB_ID
    JOBS[job_id] = {"status": "running", "result": None}

# 这里最严谨是丢到后台任务或队列里,这里先偷懒直接 await
    result = await crawl(keyword)
    JOBS[job_id]["status"] = "done"
    JOBS[job_id]["result"] = result
return {"job_id": job_id}

@app.get("/crawl/{job_id}")
defget_result(job_id: int):
return JOBS.get(job_id, {"error": "job not found"})

你看,这个就已经是一个简单的「爬虫服务」雏形了:

  • /crawl:开始一个抓取任务
  • /crawl/{job_id}:查询任务结果

以后你想接前端、接内部系统,都有统一出口。

Web 开发里那点“看起来很玄”的东西,其实也就这几块

我当时刚从脚本转 Web 的时候,最懵的是:“到底什么才叫会 Web 开发?”后来自己撸了两三套小系统,总结下来就几件事:

  1. 会定义 URL——请求怎么打到你这来
  2. 会解析请求参数、校验一下
  3. 会连数据库,把东西存起来
  4. 会把结果以页面 / JSON 的形式给出去
  5. 知道一点点性能、并发、日志、配置这些非功能的东西

一个一个拆开讲你就明白了。

连个数据库,数据才能“活”下来

你爬来的数据如果只是放 CSV,本质还是一次性的。上点强度的用法一定是存数据库、能增删改查。

用 SQLAlchemy + SQLite,写个最小模型试一下:

# models.py
from sqlalchemy import Column, Integer, String, create_engine
from sqlalchemy.orm import declarative_base, sessionmaker

Base = declarative_base()

classArticle(Base):
    __tablename__ = "articles"

    id = Column(Integer, primary_key=True, autoincrement=True)
    title = Column(String(256), nullable=False)
    url = Column(String(512), nullable=False)
    tag = Column(String(64), nullable=True)

engine = create_engine("sqlite:///./data.db", echo=False, future=True)
SessionLocal = sessionmaker(bind=engine, autoflush=False, autocommit=False)

配合 FastAPI 来一套简单的“写入 + 查询”:

# main.py
from fastapi import FastAPI, Depends
from pydantic import BaseModel
from sqlalchemy.orm import Session
from models import Base, engine, SessionLocal, Article

Base.metadata.create_all(bind=engine)

app = FastAPI()

classArticleIn(BaseModel):
    title: str
    url: str
    tag: str | None = None

defget_db():
    db = SessionLocal()
try:
yield db
finally:
        db.close()

@app.post("/articles/db")
defcreate_article_db(item: ArticleIn, db: Session = Depends(get_db)):
    obj = Article(title=item.title, url=item.url, tag=item.tag)
    db.add(obj)
    db.commit()
    db.refresh(obj)
return {"id": obj.id}

@app.get("/articles/db")
deflist_article_db(tag: str | None = None, db: Session = Depends(get_db)):
    query = db.query(Article)
if tag:
        query = query.filter(Article.tag == tag)
return query.all()

你哪怕刚开始只用 SQLite,也已经是“真正意义上的后端服务”了。 改成 MySQL、PostgreSQL,其实也就是改一下连接串、装个驱动那点事。

不只返回 JSON,页面也能自己写

公司里很多“内部系统”,压根不追求什么花里胡哨的前端框架,后端直接把 HTML 渲染出来就完了。

FastAPI 可以配 Jinja2 模板,用法跟 Flask 一样,很接地气:

from fastapi import FastAPI, Request
from fastapi.responses import HTMLResponse
from fastapi.templating import Jinja2Templates

app = FastAPI()
templates = Jinja2Templates(directory="templates")

fake_articles = [
    {"title": "Python 爬虫转 Web 开发", "tag": "python"},
    {"title": "FastAPI 实战记", "tag": "web"},
]

@app.get("/dashboard", response_class=HTMLResponse)
defdashboard(request: Request, tag: str | None = None):
    data = fake_articles
if tag:
        data = [a for a in fake_articles if a["tag"] == tag]
return templates.TemplateResponse(
"dashboard.html", {"request": request, "articles": data, "tag": tag}
    )

templates/dashboard.html 里写点最简单的模板:

<!doctype html>
<html>
<head>
<metacharset="utf-8">
<title>文章后台</title>
</head>
<body>
<h1>文章列表{% if tag %} - {{ tag }}{% endif %}</h1>
<ul>
      {% for a in articles %}
<li>{{ a.title }} [{{ a.tag or '无标签' }}]</li>
      {% else %}
<li>还没有数据</li>
      {% endfor %}
</ul>
</body>
</html>

你看,这东西跑起来,已经是个能给领导看的“后台页面”。 这也是 Python Web 很香的地方:一门语言搞爬虫、搞接口、搞页面都行。

写脚本的时候很多人不太在意性能,最多开几个线程。 但 Web 服务一上线,问题就会集中暴露:慢、卡、偶尔 502。

Python 在这块其实也不弱,ASGI + 协程该有的都有。FastAPI 默认就跑在 Uvicorn 上,天然支持 async/await,那些 I/O 密集的活(例如你并发爬几个站)反而会更舒服:

import asyncio
import httpx
from fastapi import FastAPI

app = FastAPI()

asyncdeffetch_one(url: str):
asyncwith httpx.AsyncClient(timeout=5) as client:
        r = await client.get(url)
return {"url": url, "status": r.status_code}

@app.get("/batch-check")
asyncdefbatch_check():
    urls = [
"https://example.com",
"https://httpbin.org/get",
"https://www.python.org",
    ]
    results = await asyncio.gather(*(fetch_one(u) for u in urls))
return results

爬虫那点“并发抓取”的经验,在 Web 场景里是能直接复用的: 你本来就对超时、重试、限速这些东西敏感,这些放到 Web 里简直是天然优势。

你要真想从“只会写爬虫脚本”进阶到“能搞个完整小系统”,可以按我说的这个顺序一点点加料:

  1. 先用 FastAPI/Flask 起一个最小服务:一个 GET,一个 POST,管它啥业务
  2. 再把你现有的爬虫逻辑搬进来,做成一个 /crawl 接口
  3. 接上数据库,哪怕 SQLite,保证数据是“可查询、可修改”的
  4. 用模板渲染一个后台页面,能浏览 / 搜索你爬到的东西
  5. 最后再考虑:日志怎么打、配置怎么分环境、部署到哪里跑(Gunicorn+Uvicorn、Nginx 这些)

这时候别人再问你“你会不会 Web 开发”,你就不是那种只会说“我会点爬虫”的状态了,而是可以很自然地回一句:“我用 Python 能从采集做到后端接口,简单后台也能搞。”

这俩听感完全不一样,机会也完全不一样。

-END-

我为大家打造了一份RPA教程,完全免费:songshuhezi.com/rpa.html

🔥虎哥私藏精品🔥

虎哥作为一名老码农,整理了全网最全《python高级架构师资料合集》,总量高达650GB