Python技术迷

Python 中 any() 和 all() 方法有什么作用?

线上脚本跑完,日志里没报错,结果 Excel 里混进去了 3 条脏数据。

这种问题我一般不先看循环,也不先怀疑 pandas,先看判断条件。尤其是那种一行写得很漂亮的代码:

if name and mobile and amount:
    rows.append(record)

这行看着挺顺手,实际很容易埋坑。

比如 amount 是 0 呢?

0 在 Python 里是假值,直接被挡掉了。你以为是在过滤空数据,实际上把合法的 0 金额订单也干掉了。

这时候就该把 any() 和 all() 拿出来了。

any() 看的是“有没有一个满足条件”。

all() 看的是“是不是全部满足条件”。

别急着背定义,看一段我平时更可能写的代码。

raw_orders = [
    {"order_id": "A001", "mobile": "13800000001", "amount": 19.9},
    {"order_id": "A002", "mobile": "", "amount": 35.0},
    {"order_id": "A003", "mobile": "13800000003", "amount": 0},
    {"order_id": "A004", "mobile": None, "amount": 88.0},
]

required_fields = ("order_id", "mobile", "amount")

clean_orders = []

for row in raw_orders:
    missing = [field for field in required_fields if row.get(field) in ("", None)]

if any(missing):
        print(f"[skip] order_id={row.get('order_id')} missing={missing}")
continue

    clean_orders.append(row)

print(clean_orders)

这里我没有写:

if row.get("order_id") and row.get("mobile") and row.get("amount"):

这种写法我现在基本不信。

因为 amount=0 会被误判。线上数据清洗最怕这种“看起来没毛病”的判断,它不报错,悄悄把数据删了,后面查账才发现不对。

any(missing) 的意思很直接:只要缺失字段列表里有一个东西,就说明这条数据不干净,跳过。

这就是 any() 最舒服的地方,它适合处理“只要有一个异常就拦住”的场景。

比如检查接口返回里有没有失败项:

pay_results = [
    {"user_id": 101, "status": "SUCCESS"},
    {"user_id": 102, "status": "SUCCESS"},
    {"user_id": 103, "status": "TIMEOUT"},
]

has_bad_case = any(item["status"] != "SUCCESS"for item in pay_results)

if has_bad_case:
    print("[warn] batch pay has failed or timeout record")

这段代码要是不用 any(),一般会写成这样:

has_bad_case = False

for item in pay_results:
if item["status"] != "SUCCESS":
        has_bad_case = True
break

不是不能写,就是啰嗦。

而且 any() 有个细节挺重要:它不是把所有数据都跑完才返回。只要遇到第一个 True,它就停了。

这个在大列表里有意义。

比如你扫一批日志,只要发现一条 ERROR,就没必要继续扫了:

recent_logs = [
"INFO task started",
"INFO load config ok",
"ERROR db connection timeout",
"INFO retry later",
]

if any("ERROR"in line for line in recent_logs):
    print("[alert] found error log")

看到 ERROR db connection timeout 这一行,any() 就可以收工了。

再看 all()。

all() 更适合那种“所有条件都过了,才允许继续”的场景。我平时写导入脚本时用得多。

defvalid_user(row):
    checks = [
        row.get("user_id") notin ("", None),
        row.get("mobile") notin ("", None),
        isinstance(row.get("age"), int),
0 <= row.get("age", -1) <= 120,
    ]

return all(checks)

users = [
    {"user_id": "U001", "mobile": "13800000001", "age": 18},
    {"user_id": "U002", "mobile": "", "age": 21},
    {"user_id": "U003", "mobile": "13800000003", "age": 151},
]

for user in users:
ifnot valid_user(user):
        print(f"[bad_user] {user}")
continue

    print(f"[ok_user] {user['user_id']}")

all(checks) 的意思是:这些检查项必须全部为真。

只要有一个是假,整条记录就不合格。

这里也有短路。all() 遇到第一个 False 就停,不会继续往后判断。

不过我得说一句,all() 别写得太满。

有些人喜欢一口气塞一长串:

if all([a > 0, b != "", c.startswith("X"), d in allow_list, e isnotNone]):
    ...

这种代码刚写完还行,过两周自己再看都烦。

我更喜欢把检查拆出来,尤其是线上排脏数据时,最好能知道哪一项没过。

可以这么写:

defcheck_import_row(row):
    rules = {
"sku_empty": row.get("sku") notin ("", None),
"price_invalid": isinstance(row.get("price"), (int, float)) and row["price"] >= 0,
"stock_invalid": isinstance(row.get("stock"), int) and row["stock"] >= 0,
    }

if all(rules.values()):
returnTrue, []

    failed = [name for name, passed in rules.items() ifnot passed]
returnFalse, failed

row = {"sku": "P10086", "price": -3, "stock": 12}

passed, failed_rules = check_import_row(row)

ifnot passed:
    print(f"[reject] sku={row.get('sku')} failed={failed_rules}")

这就比单纯一个 False 好查多了。

any() 和 all() 还有一个容易踩的小坑:空列表。

print(any([]))  # False
print(all([]))  # True

any([]) 是 False,这好理解,一个满足条件的都没有。

all([]) 是 True,第一次见可能有点别扭。它的意思可以理解成:没有任何一个元素违反条件,所以默认成立。

这个地方在权限判断里要小心。

比如这样写:

permissions = []

if all(p in permissions for p in ["read", "write"]):
    print("allow")

这段结果是 False,因为生成器里有两个判断。

但如果你写的是:

checks = []

if all(checks):
    print("allow")

它会输出 allow。

所以我一般会补一层判断:

checks = []

if checks and all(checks):
    print("allow")
else:
    print("deny")

别嫌这一句多,权限、金额、删除操作这种地方,我宁愿代码丑一点,也不想靠默认语义赌运气。

再补一个实际点的例子,批量校验接口返回。

responses = [
    {"trace_id": "t-001", "code": 0, "cost_ms": 83},
    {"trace_id": "t-002", "code": 0, "cost_ms": 96},
    {"trace_id": "t-003", "code": 500, "cost_ms": 20},
]

if all(resp["code"] == 0for resp in responses):
    print("[batch_ok] all api calls success")
else:
    bad_trace_ids = [
        resp["trace_id"]
for resp in responses
if resp["code"] != 0
    ]
    print(f"[batch_fail] bad_trace_ids={bad_trace_ids}")

这就是 all() 的正常用法:批量任务里,所有接口都成功,才认为这一批成功。

换成 any(),语义就变了:

if any(resp["code"] == 0for resp in responses):
    print("at least one success")

只要有一个成功就算成功。这在批量支付、批量发货这种业务里,通常是不对的。

我总结这俩方法,不想写成概念表,记两个判断顺序就行。

看到“只要有一个异常就报警、跳过、拦截”,先想 any()。

看到“必须全部通过才继续、提交、入库”,先想 all()。

另外,别把 any()、all() 当成炫技工具。判断很短,用它们挺舒服;判断带业务含义,最好拆成变量或者函数。代码不是越短越稳,尤其是跑在线上的清洗脚本和批处理任务,能打出哪条数据错了,比少写三行循环重要多了。