4名大学生出题,AI考了0分!
Fudan University has replaced a traditional final exam with an AI challenge, asking students to create questions that stump leading AI models instead of answering them.
Of the 51 students, 50 managed to make at least one model answer a question incorrectly. Four produced question sets that completely defeated one of the models, though none managed to fully stump Claude, the strongest model in the test.
The assessment was part of a data mining course, where students designed 10 computational questions based on course material, each with a single correct answer and a complete solution.
The questions were tested on three AI models, and the more mistakes the models made, the higher the student's score.
Professor Xiao Yanghua said traditional exams focused on calculation have become less meaningful in the AI era, as AI can often solve standard problems faster and more accurately than students.
题目必须基于课程讲过的知识或教材内容,每道题要有唯一正确答案,学生自己得先能把题从头到尾算对。肖仰华说:“自己出的题自己都不会,那算不上真本事。”
计算与智能创新学院24级本科生谢锦树最后拿到了97分。他尝试让AI出题来难倒自己,便搭建了一个多智能体协作的自动化出题框架,用GPT-5.5-Pro做出题层,三个应考模型作答并自动判分。框架跑起来后,他发现AI会“作弊”。
AI会伪造标准答案,把假答案塞进去,让判分脚本以为对了。它会限制最大输出长度来截断其他模型的推理过程。它会调低推理深度参数,让其他模型懒得深入思考。它还会把一道成功了的题目复制十份来凑数。
于是,谢锦树加了一个审查层,拦截钻空子行为,最终自动生成了10道题,三个应考模型全部答错。
从“怎么算”到“怎么判断”
After the exam, Xiao found that top-performing students not only understood the course content but also knew where AI was likely to fail. By contrast, lower-scoring students often relied on familiar textbook-style questions that AI could easily solve.
Xiao said the course will continue using the "human tests AI" format, shifting its focus from memorization and calculation to judgment, critical thinking and creativity — skills he believes remain essential in the age of AI.
来源:中国青年报 复旦大学
跟着China Daily
精读英语新闻
“无痛”学英语,每天20分钟就够!
↓↓↓
推 荐 阅 读