美团面试题:请介绍一下 Python 的线程同步?
作为一名 Python 开发工程师,线程同步这个话题我觉得是个绕不开的点,尤其是当我们开始接触多线程编程的时候,就会发现——一不小心,程序就可能因为线程竞争出问题。今天就来聊聊 Python 里的线程同步机制,顺便给大家几个面试必备的回答。
想象一下,你有一个多线程程序,每个线程都会修改同一个全局变量。由于 Python 线程的执行顺序是由操作系统调度的,不受我们控制,所以当多个线程同时修改这个变量时,可能会出现意想不到的结果,比如:
import threadingcounter = 0
def increment():
global counter
for _ in range(100000):
counter += 1
threads = []
for _ in range(5):
t = threading.Thread(target=increment)
t.start()
threads.append(t)
for t in threads:
t.join()
print("Final counter value:", counter) # 结果可能不是 500000
理论上,这个 counter 应该是 500000,但如果你运行几次,就会发现实际结果不稳定。这就是竞态条件(Race Condition),因为多个线程同时访问 counter,导致数据更新被覆盖或丢失。
为了解决这个问题,我们就需要用到线程同步机制。
1. Lock(锁)—— 最基础的同步工具
Python 的 threading.Lock 是最简单的同步工具,它保证同一时间只有一个线程可以访问共享资源,避免数据竞争。
import threadingcounter = 0
lock = threading.Lock()
def increment():
global counter
for _ in range(100000):
with lock: # 上锁
counter += 1
threads = []
for _ in range(5):
t = threading.Thread(target=increment)
t.start()
threads.append(t)
for t in threads:
t.join()
print("Final counter value:", counter) # 现在结果稳定是 500000
with lock: 这种方式比手动 lock.acquire() 和 lock.release() 更优雅,因为它能自动管理锁的释放。
但锁也有缺点,如果不小心忘记释放锁,或者多个线程同时等待锁而导致死锁(Deadlock),那问题就大了。所以如果能避免使用锁,尽量不要滥用。
2. Condition(条件变量)—— 适用于生产者-消费者模型
有时候,我们的线程需要等待某个条件满足才执行,比如典型的生产者-消费者问题。这时候 threading.Condition 就派上用场了。
import threading
import timecondition = threading.Condition()
items = []
def producer():
global items
for i in range(5):
time.sleep(1)
with condition:
items.append(i)
print(f"Produced {i}")
condition.notify() # 通知消费者
def consumer():
global items
for _ in range(5):
with condition:
while not items:
condition.wait() # 等待生产者通知
item = items.pop(0)
print(f"Consumed {item}")
t1 = threading.Thread(target=producer)
t2 = threading.Thread(target=consumer)
t1.start()
t2.start()
t1.join()
t2.join()
在这个例子里,消费者在 condition.wait() 处阻塞,直到生产者 condition.notify(),确保不会访问空的 items 列表。
3. Semaphore(信号量)—— 允许多个线程访问资源
threading.Semaphore 适用于资源有限的情况,比如限制最大线程数或数据库连接数等。
import threading
import timesemaphore = threading.Semaphore(2) # 最多允许2个线程同时执行
def worker(i):
with semaphore:
print(f"Thread-{i} is working...")
time.sleep(2)
print(f"Thread-{i} finished.")
threads = [threading.Thread(target=worker, args=(i,)) for i in range(5)]
for t in threads:
t.start()
for t in threads:
t.join()
这个例子限制了同时最多有两个线程在执行 worker 任务。信号量的计数值控制了可同时访问资源的线程数量。
4. Event(事件)—— 线程间的简单通信
threading.Event 适用于让某个线程等待另一个线程的信号。它的状态可以是设置(set)或未设置(clear),线程可以通过 event.wait() 来等待。
import threading
import timeevent = threading.Event()
def waiter():
print("Waiting for event...")
event.wait() # 阻塞,等待事件被 set()
print("Event received! Proceeding...")
def setter():
time.sleep(3)
print("Setting event...")
event.set() # 触发事件,让 waiter 继续执行
t1 = threading.Thread(target=waiter)
t2 = threading.Thread(target=setter)
t1.start()
t2.start()
t1.join()
t2.join()
这里 t1 线程会一直等到 event.set() 执行后才继续,适用于线程间的简单信号通信。
线程同步的最佳实践
避免死锁:多个线程相互等待锁释放可能会造成死锁,最好的办法是使用 try-finally或with lock确保释放锁。减少锁的使用范围:锁定的代码块越小,性能损失越少。 使用 Queue避免手动同步:如果你是在线程间传递数据,queue.Queue是更好的选择,它内部已经实现了锁机制,避免手动加锁的麻烦。
import queue
import threadingq = queue.Queue()
def producer():
for i in range(5):
q.put(i)
print(f"Produced {i}")
def consumer():
while not q.empty():
item = q.get()
print(f"Consumed {item}")
q.task_done()
t1 = threading.Thread(target=producer)
t2 = threading.Thread(target=consumer)
t1.start()
t1.join() # 确保生产者先执行
t2.start()
t2.join()
最后,来看几道相关面试题吧。
Q1:为什么 Python 需要线程同步?
A1:因为 Python 的多线程在访问共享资源时可能会引发竞态条件,导致数据不一致或程序错误。线程同步机制可以确保多个线程安全地访问共享资源。
Q2:Lock 和 RLock 的区别?
A2:Lock 在同一个线程里不能多次 acquire(),否则会死锁。而 RLock(可重入锁)允许同一个线程多次 acquire(),但必须释放相同次数。
Q3:Queue 如何帮助线程同步?
A3:queue.Queue 内部实现了锁,线程安全,适用于生产者-消费者模式,避免手动加锁带来的复杂性。
Python 的线程同步机制看似复杂,但只要理解了 Lock、Condition、Semaphore、Event 等工具的用途,写多线程代码就会得心应手。面试的时候,别光记概念,最好能手写点代码,这样才能真正掌握!
对编程、职场感兴趣的同学,大家可以联系我微信:golang404,拉你进入“程序员交流群”。
虎哥作为一名老码农,整理了全网最全《python高级架构师资料合集》。