Python技术迷

别再重复造轮子了!Python这10个内置模块真能造!

脚本跑到凌晨两点挂了,原因不是接口超时,也不是磁盘满了,是路径拼错了。 更离谱的是,代码里还手写了一堆字符串切割、日志打印、参数解析、临时文件清理。看着像很勤快,实际就是在重复造轮子。Python 标准库里有些模块,真不是摆设。很多小工具、排障脚本、数据清洗脚本,用好它们,代码能少一半,坑也少一半。 我平时写 Python 脚本,第一反应不是先装第三方库。能用内置模块顶住的,先用内置模块。线上机器环境不一定允许你随便 pip install,尤其是那种临时排障脚本,越少依赖越稳。 第一个我会先看 pathlib 。 以前很多人这么写路径:
log_file = base_dir + "/" + day + "/" + "app.log"
这种代码我第一眼就不太信。Linux 下还好,换个环境、多个斜杠、目录不存在,问题就来了。 我一般这样写:
from pathlib import Path

day = "2026-06-26"
log_dir = Path("/data/logs") / day
log_dir.mkdir(parents=True, exist_ok=True)

log_file = log_dir / "app.log"

if log_file.exists():
    print(log_file.read_text(encoding="utf-8")[:500])
Path 最大的好处不是优雅,是少犯低级错。目录创建、文件判断、路径拼接都在一个对象上,不用脑子里来回模拟字符串。 第二个是 collections 。 统计日志里的错误码,别再手写字典判断了。
from collections import Counter

codes = []

with open("access.log", encoding="utf-8") as f:
    for line in f:
        if "status=" not in line:
            continue
        codes.append(line.split("status=")[1].split()[0])

print(Counter(codes).most_common(5))
这个脚本我真会这么写。先把现场数据捞出来,看哪个状态码最多,再决定翻不翻业务代码。很多时候不是代码复杂,是你一开始就看错方向。 第三个是 defaultdict ,也在 collections 里。 按用户聚合异常次数,别写一堆 if key not in dict 。
from collections import defaultdict

bad_users = defaultdict(int)

with open("pay.log", encoding="utf-8") as f:
    for line in f:
        if "PAY_FAIL" not in line:
            continue
        user_id = line.split("uid=")[1].split()[0]
        bad_users[user_id] += 1

for uid, count in sorted(bad_users.items(), key=lambda x: x[1], reverse=True)[:10]:
    print(uid, count)
这种东西适合临时排查。十分钟内把异常用户捞出来,比在那里讨论“可能是风控问题”靠谱。 第四个是 itertools 。 批量处理数据时,别一次把几十万行全塞进内存。尤其是导入、补数、刷缓存这类脚本,最怕写得很猛,跑起来把机器搞死。
from itertools import islice

def read_batch(fp, size=500):
    while True:
        batch = list(islice(fp, size))
        if not batch:
            break
        yield batch

with open("user_ids.txt", encoding="utf-8") as f:
    for batch in read_batch(f, 300):
        ids = [x.strip() for x in batch if x.strip()]
        print("处理一批:", len(ids))
        # 这里再去查库、调接口、刷缓存
批量脚本我一般都会留这个口子。别相信“数据量不大”,这句话线上最容易变。 第五个是 functools 。 有些接口查配置、查字典表、查地区编码,结果半天不变。你手写全局缓存也行,但大概率会写出一坨不好清的代码。
from functools import lru_cache

@lru_cache(maxsize=128)
def load_region_name(region_code):
    print("查一次配置:", region_code)
    # 现场脚本里这里可能是查库或读文件
    table = {"110000": "北京", "310000": "上海"}
    return table.get(region_code, "未知")

print(load_region_name("110000"))
print(load_region_name("110000"))
注意, lru_cache 别乱套在会变的数据上。缓存这东西,救命也坑人。配置一天变一次可以用,订单状态这种东西就别省那点查询了。 第六个是 contextlib 。 打开文件、临时切换目录、捕获可忽略异常,别到处写 try finally 。能收住资源的代码,后面排障时少很多脏东西。
from contextlib import suppress
from pathlib import Path

tmp_file = Path("/tmp/import.lock")

with suppress(FileNotFoundError):
    tmp_file.unlink()
这个我经常用在清理旧文件、删锁文件的脚本里。文件不存在不是异常,不值得刷一屏红日志。 第七个是 dataclasses 。 很多人写脚本也爱传一堆 dict,传到后面自己都忘了字段名。
from dataclasses import dataclass

@dataclass
class ImportRow:
    user_id: str
    phone: str
    source: str

row = ImportRow(user_id="u1024", phone="13800000000", source="crm")
print(row.user_id, row.source)
这玩意不是为了“面向对象”,就是为了少写错字段。尤其是字段清洗脚本,dict 一多, userId 、 user_id 、 uid 混在一起,后面查问题很烦。 第八个是 argparse 。 脚本别把日期、环境、文件路径写死。写死一次,别人复制一次,事故概率就高一点。
import argparse

parser = argparse.ArgumentParser()
parser.add_argument("--day", required=True)
parser.add_argument("--file", required=True)
parser.add_argument("--dry-run", action="store_true")

args = parser.parse_args()

print("日期:", args.day)
print("文件:", args.file)
print("只演练:", args.dry_run)
有 --dry-run 的脚本,我会更敢跑。先打印要处理什么,不真正写库。看一眼范围对不对,再去掉参数执行。这个习惯能挡住不少手抖。 第九个是 logging 。 不要再满代码 print() 。排障脚本可以 print,定时任务不行。真出问题时,你需要时间、级别、上下文。
import logging

logging.basicConfig(
    level=logging.INFO,
    format="%(asctime)s %(levelname)s trace=%(trace_id)s %(message)s"
)

logger = logging.LoggerAdapter(
    logging.getLogger("clean_job"),
    {"trace_id": "clean-20260626"}
)

logger.info("开始清理过期文件")
logger.warning("发现空目录,跳过")
日志不是写给正常时候看的,是写给半夜脑子不清醒的时候看的。字段越稳定,越好搜。 第十个是 sqlite3 。 这个模块经常被低估。临时对账、小批量落盘、脚本断点续跑,直接用它很舒服。别动不动就搞 MySQL 表,尤其是一次性任务。
import sqlite3

conn = sqlite3.connect("job_state.db")
conn.execute("""
create table if not exists done_user (
    user_id text primary key,
    handled_at text default current_timestamp
)
""")

def mark_done(user_id):
    conn.execute("insert or ignore into done_user(user_id) values (?)", (user_id,))
    conn.commit()

def is_done(user_id):
    row = conn.execute(
        "select 1 from done_user where user_id = ?",
        (user_id,)
    ).fetchone()
    return row is not None

for uid in ["u1", "u2", "u1"]:
    if is_done(uid):
        print("跳过已处理:", uid)
        continue
    print("处理:", uid)
    mark_done(uid)
这种断点记录很实用。脚本跑一半挂了,重新跑不会从头再干一遍。补数据、刷标签、批量通知,都能用。 再补两个我常顺手用的: json 和 hashlib 。 接口返回落文件,用 json ;算文件指纹、防重复,用 hashlib 。
import json
import hashlib
from pathlib import Path

data = {"uid": "u1024", "tags": ["vip", "active"]}
Path("user.json").write_text(
    json.dumps(data, ensure_ascii=False, indent=2),
    encoding="utf-8"
)

digest = hashlib.sha256(Path("user.json").read_bytes()).hexdigest()
print(digest[:16])
别小看这些模块。真正写脚本时,拼的是少出错、好回滚、能复跑、日志能查。不是写得多复杂。 Python 内置库最适合干这种脏活累活:查日志、清文件、导数据、做对账、补状态。能用标准库解决的,就别先掏第三方库。依赖越少,脚本越敢往线上机器放。