dots3-note Preview:迈向服务真实生活的长程智能体,坚定的第一步
今天,我们开源了 dots3-note Preview。作为目前 dots3 系列中最轻量的模型,note 的总参数为 280B、激活参数为 16B,支持 512K 上下文,具备文本、视觉和语音的多模态理解能力,并针对复杂推理和长程 Agent 任务进行了优化。完整的 dots3 系列将包括 note、jazz 和 aria,以覆盖不同的任务复杂度、响应速度与计算成本。
我们一直希望能构建出让每个人受益的 AI,帮助人们解决生活中的各种问题。要达成这一目标的前提是:它能够感知真实世界,进行复杂推理,使用工具解决问题;更重要的是,它还要具有主动性和持续学习能力,能够随着世界的变化和与用户的交互而不断进化。当前的大模型距离这一目标仍有距离。
在 dots3-note Preview 中,我们首先面向这一目标构建了能同时感知文本、视觉和声音的多模态基座,随后在封闭、可验证的任务上提升了模型的推理和问题解决能力。在多项任务上,模型可比肩甚至超过参数规模数倍于自身的大尺寸模型。此外,我们还通过长程任务的强化学习训练,提升了模型从环境中学习和自主更新记忆的能力。最后,我们构建了一套模拟真实生活的环境,用于衡量并缩小封闭环境与开放的真实生活环境之间的差距。为了加速这一领域的发展,本次开源还发布了两个面向真实生活场景的评测环境:VibeSearchBench与 VibeLifeBench。
本次开源的模型已按照 Apache 2.0 协议发布于 Hugging Face和GitHub,模型结构已提交至Transformers。如需使用模型 API,可访问这里申请。详细的模型结构、训练和评测信息将与技术报告一起于一周内发布。
本文剩余部分展示了我们在 dots3-note Preview 中取得的进展,也向研究者们分享了探索过程中的发现以及下一步要解决的问题。我们希望这些工作能够促进并加速开源 AI 的发展。
dots3-note Preview 在推理、Agent 及多模态感知等能力上进行了全面优化,在多项任务上取得了比肩国内外更大尺寸优秀模型的效果:
随着人们希望 Agent 解决的任务越来越复杂、执行周期越来越长,现有 value-free 强化学习范式遇到了以下困难:Agent 完成一次探索需要十余个小时,训练效率难以接受;同时,稀疏的奖励也难以实现有效的信用分配。PPO 等 actor-critic 方法可以缓解这些问题,但 critic 通过固定算力的前向计算进行价值估计,无法像 actor 一样通过推理、反思和工具调用来分析当前状态并得出结论,因此在复杂问题上很难做出准确估计。
为此,我们提出 TEMPO(Test-time-scaled Value Estimation with Macro-step Policy Optimization)。TEMPO 将长程任务拆分为多个 macro-step,每个 macro-step 包含多轮模型与环境的交互。在每个 macro-step 结束时,同一个 Agent 从 actor 切换为 critic,通过 test-time scaling 的推理分析来估计当前状态的预期剩余回报,从而在长程任务结束前更新策略。在训练过程中,我们不仅通过强化学习训练模型如何行动,也训练它如何自我评价。通过 TEMPO,我们得以在单次运行需要数十小时的任务上进行有效的强化学习训练。
不同方法训练的模型在 ARC-AGI-3 任务上的评测结果存在显著差异:TEMPO 的平均 score 比 baseline checkpoint 提高了 31.5%,比 GRPO 提高了 20.6%;在达到相同关卡的情况下,TEMPO 模型以更少的通关步数获得了更高的分数。
Critic 的分析结果
例如在一个“放置骑士”的游戏中,Agent 需要在棋盘已有预置棋子的情况下继续放置指定数量的棋子,同时满足隐藏约束。
在一次训练中,两条轨迹分支连续运行 64 轮后都没有通过新的 level,从环境得分上看没有区别。但其中分支 B 将“棋子发生攻击关系”错误地理解成了任务目标,并继续沿着这个假设搜索;分支 A则已经识别出真正的冲突规则,并找到了一组接近可行解的候选布局。
Critic 读取两条轨迹,并结合任务状态和 privileged information 重新判断当前进展,对两条分支给出了明显不同的 value estimate。
递归自我评价
在 TEMPO 的训练过程中,我们发现 Agent 的自我评价能力显著强于预期。即使面对 Agent 当前还无法解决的问题,当它作为 critic 时,仍能从两个表面相似的状态中识别出真正突破了环境规律的那个状态,并给出差异显著的价值估计。对实际蒙特卡洛采样结果的分析表明,当允许 Agent 在测试时扩展计算时,它作为 critic 的能力非常强。这说明在我们关注的任务上,“评价比生成更简单”这一假设成立。
除了在强化学习训练阶段之外,我们也在推理阶段看到了模型自我评价的重要价值。在上个月举办的 IMO 2026 中,我们基于 dots3-note Preview 的分支版本构建了一套内部 harness,让 Agent 递归生成 proof,并通过工具调用对生成的 proof 进行自我评估和增强,充分发挥其自我评价和改进能力,最终取得了 IMO 官方认证的 42/42 满分金牌成绩。
我们发现,模型进行自我评估的能力非常重要。我们希望从可验证的长程任务走向真实世界、真实生活中的模糊长程任务。要实现这一目标,最关键的问题是这些任务缺少可验证的奖励信号;过去依赖人类专家评价模型结果的方式可能很难奏效,而模型进行递归自我评价(recursive self-critiquing)的能力可能是解决这一问题的关键,也是我们下一步的重点工作方向。
为了实现这一目标,我们意识到 Agent 必须具备一项重要能力——它必须能在测试时,从对外部世界的探索和对用户的了解中不断学习。为此,我们构建了数千个不依赖先验知识的新颖超长程环境,专门训练模型在陌生环境中在线学习知识、更新记忆的能力。我们发现,只要任务长度显著超过模型的上下文长度,通过上述强化学习方法训练模型解决问题后,模型就能自行学会生成有助于未来决策的记忆。
我们在陌生的虚拟环境 ARC-AGI-3 中验证了 dots3-note Preview 的这一能力。ARC-AGI-3 是 ARC Prize 今年推出的一项人工智能能力测试,其重点不再是静态推理题,而是测试 AI 能否像人一样在陌生环境中自主学习。其中,复杂任务需要 Agent 自主进行数千轮交互,完整运行时间可达 40–50 小时。
我们发现,dots3-note Preview 不仅表现出优秀的问题解决能力,也具备良好的“研究性”思维:它能通过与环境交互进行观察,主动提出并验证假设,更新记忆,并在未来的决策中使用这些记忆。同时,我们观察到,这一能力也能泛化到模型从未训练过的其他类型的陌生环境中。下方另一个案例展示了 Agent 在《杀戮尖塔 2》中持续运行并更新记忆的行为。
ARC-AGI-3 任务上的平均分数与推理成本之间的关系。其中,GPT-5.5 使用 Codex 评测,其余模型均使用 Claude Code。为避免缓存命中率波动带来的影响,输入价格不区分是否命中缓存。评测代码
ARC-AGI-3 的部分游戏完整运行时长可达数十小时。为便于展示,此处选择运行时长较短的游戏。
Memory 节选
提出并验证假设
建立假设
……
Goal hypothesis: MERGE the two blues (overlap same cell). Shortest path to overlap=10 moves, overlap at (1,5)
……
验证假设
……
Goal hypothesis: MERGE the two blues (overlap same cell). Shortest path to overlap=10 moves, overlap at (1,5)
……
纠正过时 Memoy
……
- COUPLING (re-verified): up/down PARALLEL; left -> L(min-col) col+1, R(max-col) col-1;
right -> L col-1, R col+1. (left/right mirror INVERTED vs L2/L3.)
……
杀戮尖塔完整测试长达数十小时,此处截取部分视频展示。
完整 Memory 节选
(向下滑动查看更多)
……
# MEMORY.md
##Game: mirror-coupled maze (two blue tokens)
- 64x64 ASCII color-coded frames. Grid of band-aligned cells. Cell size varies per level
(L0/L1=5x5 or 4x4; L2=4x4). Bands for L2 = list(range(2,63,4)) -> 16x16.
- Two blue tokens: LEFT and RIGHT. Coupling (all levels so far):
- up/down: both move SAME vertical dir (parallel).
- left/right MIRRORED: LEFT blue right->col+1,left->col-1; RIGHT blue right->col-1,left->col+1.
- A token moves 1 cell if target is walkable, else stays. Asymmetric walls desync the blues.
- WIN CONDITION (CONFIRMED L0,L1,L2): MERGE the two blues (same cell) -> current_level++.
- SOLVER LOOP per level: parse frame -> detect bands/cellsize from blue block -> build celltype
grid -> BFS joint state (mirror coupling) to merge -> execute in <=10-move chunks, verifying
blue positions vs predictions each step; re-BFS from actual state on any divergence.
##Levels 0,1: SOLVED
- L1: R = HAZARD (moving blue onto R kills it -> level RESET). M/u = walls. K/b = safe corridor.
Never move a blue onto R. Merge = win.
- Parser gotcha: split blues by col<32 FAILS when a blue straddles col 31/32. Use connected-components.
##KEY MECHANIC (marker/cursor), confirmed by action-delta on L2:
- Click a BB 2x2 marker -> it becomes an activated CURSOR (YY), blues FREEZE (turn to 'w').
Arrows then move ONLY the cursor (up=row-1,down=row+1,left=col-1,right=col+1), leaving empty
corridor behind as it moves.
- Click ANY empty cell -> drops the cursor as a wall (BB) at the CURSOR's current position, blues
UNLOCK (back to 'b'). Clicking a different BB marker switches control to it.
- Markers are WALLS for blues. A blue can only pass a gate where a marker WAS (relocated away).
- Pixel coords for clicks: px=2+col*4+1, py=2+row*4+1 (x=col-pixel, y=row-pixel).
##Level 2: SOLVED (2026-08-01)
- U and R are BOTH WALLS. Two blues start (11,5),(11,9) in two sealed lower regions.
Markers(walls) at (3,7)[bridge gate],(4,2)[left gate],(7,9)[right gate]. Upper region
(bridge + connecting corridors) is isolated; both blues connect to it ONLY through these 3
markers. All 3 must be opened (moved to leaves) for the blues to reach each other.
- SOLUTION: relocate each marker to a FINAL-GRAPH LEAF (dead-end, so walling it is harmless),
opening its original cell. Drops used: (4,2)->(8,1), (7,9)->(4,10), (3,7)->(2,3).
Verify final config reaches both blues via BFS first (all 3 open + 3 leaf-walls: L reaches R).
Then mirror-coupling BFS (state=(Lr,Lc,Rr,Rc), merge when same cell) -> 31-move path.
Executed in <=10-move chunks, every step matched prediction. Merge at (6,5) -> level 3.
- General marker-gate approach: open every marker that gates a region, drop each at a leaf of the
all-markers-open graph, then mirror-BFS merge.
##Level 3: SOLVED (2026-08-01)
- 5x5 cells, bands=range(4,59,5) -> 11x11. Pixel center = 4 + band*5 + 2. Two blues start L(6,2)
[min col], R(4,8). Marker at (5,5). R cells (mixed K+R) are HAZARDS: moving a blue onto an
R cell RESETS the whole level to start (NOT a wall - a blue falls in and resets). Hazard cells
at rows 1,4,6,9 (cols 4-6 in middle) etc.
- Merge impossible by pure movement (mirror coupling: LEFT col+RIGHT col invariant unless blocked;
row-diff only changes via blocking). Even opening marker (5,5) gives no merge.
- SOLUTION: relocate marker to a wall cell that lets the blues desync (e.g. (5,2)), then mirror-BFS
merge. Marker move: click marker (px=4+band*5+2), cursor arrows, click empty to drop.
- CRITICAL BUG HIT: BFS start must assign LEFT=blue with MIN col, RIGHT=MAX col (L=min(bp,key=
x:x[1]), R=max(...)). I hardcoded start=(4,8,6,2) which SWAPPED L/R -> merge path invalid.
Always derive L/R from actual blue positions each level. Verify every executed move vs prediction.
##Level 4 — SOLVED
- Executed 38-move BFS merge. All steps matched sim exactly. Merge at (1,7) -> current_level 4->5.
- Key mechanic recap: HOLD-type colored pressure-plate switches; a blue on color-X switch opens color-X gates; only 2 coupled blues so use walls to park one on a switch while moving other. left/right mirror INVERTED vs L2/L3 (right->L col-1, R col+1).
##Now on: Level 5 (in progress)
- 16x16, bands=range(2,63,4), px center=2+band*4+2. Blues start (5,4)[TL],(5,10)[TR).
- COUPLING (verified): up/down PARALLEL; left->L(min-col) col-1, R(max-col) col+1; right->L col+1, R col-1. (= levels 0-3 mirror, NOT inverted).
- R cells = HAZARD (moving a blue onto R RESETS level). M,p = walls.
- Board split by col7 (solid R hazard) into left(TL rows2-6 c2-6, BL rows8-12 c2-6) and right(TR,BR). Only connection = gap (10,7) (K), initially blocked by a 2x2 B marker.
- SWITCHES (verified): O(3,4)[TL]->opens O-gate(7,9-11)[TR-BR]; G(3,10)[TR] & G(10,4)[BL]->open G-gate(7,3-5)[TL-BL]. Gates open while a blue sits on matching switch (hold-type); cells become K when open.
- B MARKER mechanic: click B -> cursor (blues freeze to 'w', gates stay open if blues on switches). Cursor moves 1 cell/arrow through K only (CANNOT pass closed gates or R). Click empty cell -> drops WALL at CURSOR's pos (rendered 2x2 'BB', blues unfreeze). Wall placement limited to cursor-reachable region.
- Cursor CANNOT reach TL normally (G-gate blocks) BUT can pass G-gate when it's OPEN (blue frozen on a G-switch).
- SOLUTION PLAN (needs reset first to restore B at (10,7)):
1. reset; up;up -> blues (3,4),(3,10) on both switches, BOTH gates open.
2. click B(10,7) px(31,43) -> cursor active, blues frozen on switches, gates open.
3. move cursor to (4,4): left,up,up,left,left,up,up,up,up (goes BL up through open G-gate(7,4) into TL).
4. click empty (e.g. (12,12) px 52,52) -> wall at (4,4), blues unfreeze at (3,4),(3,10).
5. 26-move merge (coupling B, wall(4,4)): down,down,down,down,down,down,left,down,right,right,right,right,right,down,right,up,up,down,right,down,left,down,up,up,up,up -> MERGE at (8,5).
- 16x16 grid, bands=range(2,63,4) -> 16x16. Pixel center = 2 + band*4 + 2.
- NO markers this level. Blues start (12,1),(12,13). K=corridor; M,p=permanent walls.
- COUPLING (re-verified): up/down PARALLEL; left -> L(min-col) col+1, R(max-col) col-1; right -> L col-1, R col+1. (left/right mirror INVERTED vs L2/L3.)
- GATE/SWITCH mechanic (HOLD-type pressure plates): single cells u(12,3),u(1,3); O(6,14); G(6,8) are SWITCHES. A blue standing on a color-X switch OPENS all color-X gate cells (3-cell bars in barrier rows); gates CLOSE when blue leaves switch (verified: L on (12,3) opened u-gates (5,10-12)&(9,10-12); leaving closed them).
- Gates: U-gates (5,10-12),(9,10-12); O-gates (9,2-4); G-gates (5,3-5). Blues blocked by closed gates (walls, no reset).
- Regions: LL(rows10-13,0-6), LM(6-8,0-6), LU(1-4, full width), RL(10-13,8-14), RM(6-8,8-14). LU=RU (shared upper). LL exit=O-gate; LM exit up=G-gate, down=O-gate; RL exit=u-gate; RM exit up/down=u-gate. Switches: u(12,3) in LL; O(6,14),G(6,8) in RM; u(1,3) in LU.
- SOLUTION: BFS over (Lr,Lc,Rr,Rc), gate-open derived from current positions, each blue moves independently blocked by walls/closed-gates, merge when same cell. Executing 38-move path found (start from (12,4),(12,10) after my test moves). Merge at (1,7).
- Path: right,up,up,up,up,left,up,right,right,up,right,right,up,left,left,left,left,up,up,up,left,left,up,up,up,up,up,right,right,up,left,up,up,up,up,left,left,left
- Every gate-crossing step has switch-holder blocked (stays), so gate stays open during move.
……
要从虚拟环境走向真实任务,模型除了要能够进行自我评价和从环境中学习,还需要解决两个基本问题:如何感知世界,以及如何影响世界。本次发布的 dots3-note Preview 首次具备了同时理解文本、视觉和语音信息的能力,并在通用代码、工具调用和结果交付方面有了明显提升。
下面两个案例展示了 dots3-note Preview 如何结合这两种能力,完成更加开放、多模态的真实世界复杂任务。
以上场景,从虚拟环境到可控的编程环境,都是可验证的任务——环境封闭、状态可枚举、结果对错分明。真实生活并非如此,它与前述任务存在三个根本区别:
用户不会明确描述需求:现实中,用户意图往往在多轮交互中逐步形成,并随中间结果继续变化。
任务需要通过多次交互完成:现实任务可能持续数天甚至数周,需要模型跨阶段保留并更新任务状态。
外部环境是动态的:天气、价格、库存和航班等状态会异步变化,且这些变化未必会被主动告知。
要让模型具备在这样的世界中工作的能力,首先需要构建能够承载这些任务的环境。为此,我们搭建了面向真实生活的复杂场景模拟环境:覆盖出行、就医、采购、履约、社交等大量真实服务,并接入了大量可调用的工具;每个场景都有自己的后端状态机与模拟时间线,用户消息、业务通知和外部状态变化按阶段自然发生,而不是一次性呈现给模型。用户由 persona-driven 用户模拟器扮演,其意图随交互逐步披露,并会根据中间结果继续变化。
在这套环境上,我们训练模型跨阶段维护任务状态、在外部信息变化时主动回溯并修正既有计划的能力。
下面三个案例展示了 dots3-note Preview 在这类模拟的真实生活长程任务中的表现:
为了让社区能够复现和衡量这类能力,我们同时开源了 VibeSearchBench 与 VibeLifeBench:
VibeSearchBench:评估意图逐步披露条件下的多轮搜索,包含 20 个领域的 200 项任务。每项任务由模糊的初始请求和 persona-driven user simulator 构成,用户约束随多轮交互逐步披露;Agent 可以执行 search、visit 和 code,并通过预测知识图谱与 ground-truth graph 的 node/triplet matching 进行评估,主要指标为 Triplet F1。
VibeLifeBench:评估非静态环境中的跨阶段任务执行,包含 10 个领域的 20 项任务。每项任务跨越 20–30 个阶段和一条模拟时间线,覆盖大量真实服务环境;用户消息、业务通知与后端状态变化按阶段发生,并通过 1,247 项 atomic checks 直接评估跨阶段状态一致性、工具执行结果与最终交付。例如,前文的家庭旅行案例要求模型在机型、天气和航班状态变化后持续更新既有行程。
构建能够解决现实问题、帮助人们生活得更好的 AI,是 dots 一直以来的目标。我们始终相信,技术应当为人而建,最终也应该帮助人们解决真实世界、真实生活中的问题。
dots3-note Preview 是我们朝这个目标迈出的一小步,也是一个新的起点。从封闭环境中的可验证任务走向真实生活中开放、模糊的长程任务,我们仍有很长的路要走。这个模型还有许多尚未解决的问题,也有许多做得不够好的地方。我们选择开放这个模型,是希望听到来自社区和用户的真实反馈,并在持续学习中把每一个环节做得更好。
我们相信,开源模型权重只是开放的一部分。我们也愿意分享探索过程中的方法、经验、失败与反思,让更多人能够理解、使用并继续构建这些技术。也邀请所有认同这一方向的人与我们同行——共同构建真正服务于人、也属于每个人的 AI。
局限性
dots3-note Preview 是一个阶段性预览版本。当前的强化学习训练尚不充分,模型在幻觉抑制、文本与多模态能力的平衡以及稳定性方面仍有不足。
本文中的真实生活任务运行在模拟环境中。要将这些能力转化为可靠、完整的真实体验,除了基础模型,还需要 harness、connector、数据源、安全与权限机制,以及产品体验的共同完善。相关工作仍在持续推进。
Reasoning & Agentic 评测
Notes:
Results with * are from our own testing.
Terminal-Bench 2.1: We evaluated the model using the Terminus-2 harness with parser=json, a 10-hour timeout, temperature=0.7, top_p=0.95, max_tokens=81,920, and a 256K context window. Each task was allocated 8 CPUs and 16 GB of RAM.
ARC-AGI-2: We evaluated models on the official public evaluation set; unmarked results are official leaderboard scores from the private set. dots3-note Preview was evaluated with temperature=0.7, top_p=1.0, and max_tokens=384K. Other public-set results use the best available inference configuration for each model. † Claude Opus 4.8: The official leaderboard result uses reasoning_effort=high because no official result for the maximum reasoning setting is available.
ARC-AGI-3 (arcagi3 harness): We evaluated models on the official public evaluation set using the official evaluation code. Unmarked results are official leaderboard scores from the private set. Opus 4.8 and GPT-5.5 were both evaluated with reasoning effort set to high.
ARC-AGI-3 (general harness): We evaluated models on the official public evaluation set. GPT-5.5 was evaluated using Codex, while all other models were evaluated using Claude Code. All models used a prior-free harness, and the evaluation code has been open-sourced. Opus 4.8 and GPT-5.5 were both evaluated with reasoning effort set to high.
ClawEval: We evaluated models with ClawEval on the General set (199 tasks), including 38 multi-turn user-agent dialogues (C-series) and 161 single-turn tasks (T-series). Inference used a temperature of 1.0, top-p of 0.8, a maximum output length of 16,384 tokens, and a 262K context window. The model was given an OpenClaw-style system prompt.
WildClawBench: We evaluated models on the full WildClawBench set (60 tasks). Inference used a temperature of 0.7, top-p of 0.95, a maximum output length of 65,536 tokens, and a 320K context window. The model ran in the OpenClaw 2026.6.1 agent harness with the full tool profile. Trajectories were graded using a combination of deterministic checks and a GPT-5.4 judge.
VibeLifeBench: We evaluated all models using the OpenClaw harness with a 256K context window. Results are reported as avg@3 on benchmark version v1.0.0.
Agentic Search: We evaluated all models using our internal ReACT harness. Access to Hugging Face was blocked during evaluation to prevent test data leakage. GPT-OSS-120B served as the judge for BrowseComp, BrowseComp-zh, and LiveBrowseComp. Qwen3.5-397B-A17B served as the judge for VibeSearchBench, with Triplet F1 as the primary metric. Gemini-3.5-flash served as the judge for DeepSearchQA and WideSearch; we report F1 and item-level F1, respectively.
APEX-Agent: We evaluated our model in a multimodal setting using the Archipelago harness. Gemini-3.5-flash served as the judge, and we report pass@1.
Toolathlon-Verified: Our evaluation used a temperature of 1.0, top-p of 0.95, a maximum output length of 32K tokens, and a 384K context window. Results are reported as avg@3.
Skills-Bench: We evaluated our model on 79 tasks using Claude Code, excluding tasks that depend on external APIs. Results are averaged over three runs.
SWE series: We evaluated the SWE-bench suite with the live-swe-agent framework using a tailored instruction prompt. Inference used a temperature of 1.0, top_p=0.95, max_new_tokens=32K, and a 384K context window.
NL2repo: We evaluated models using the official harness. Each task was limited to 250 turns and 10 hours, with 4 CPU cores and 32 GB of RAM. Inference used a temperature of 1.0, top-p of 0.95, a maximum output length of 49,152 tokens, and a 384K context window. When the context limit was reached, older reasoning and long tool inputs and outputs were trimmed, while the latest 24 messages and complete tool-call/result pairs were preserved. Anti-hacking instructions and task-specific web-access rules were also applied.
IMOAnswerBench: We manually reviewed all problems in IMOAnswerBench and corrected the ground truth where needed. One problem was removed because its statement was corrupted.
Codeforces: We evaluated models on 114 problems from 14 Codeforces Division 1 contests (May–November 2025). For each problem, we generated 32 solutions, then randomly sampled and ordered 10 submissions without replacement. Solutions were submitted to the official Codeforces online judge. We scored the first accepted submission using the median human score after the same number of failed attempts. The average expected rating across all contests was estimated with 100,000 Monte Carlo trials per problem. Because of constraints such as official API rate limits, we generated only 10 solutions for Hy3 and Qwen397b-a17b; as with the other models, we used 100,000 Monte Carlo trials to randomly order the solutions and estimate the expected rating.
LiveCodeBench v6: Because of constraints such as official API rate limits, we performed only three sampling runs for Hy3.
多模态评测
Notes:
HiPhO: We used Gemini-3.5-flash for grading. Scores for each competition were normalized to a 100-point scale based on the official maximum score, then macro-averaged with equal weight across the original 13 competitions.
GDP.pdf: We used Seed 2.0 Lite as the visual judge, providing both images of the relevant pages and OCR text as context. Each rubric criterion was scored as pass or fail. For each task, we first computed the fraction of criteria passed, then averaged the results equally across all 100 tasks to obtain the final mean-criteria score.
MME-Video-V2: Audio input was enabled when evaluating both dots3 and Gemini models.
VideoZeroBench / MathVision / CharXiv-Reason: We first used GPT-4o-mini to extract the final answer, then determined correctness using GPT-OSS, GPT-4o, or rule-based code.
Other benchmarks: For multiple-choice tasks, GPT-OSS was used consistently to extract the selected option. For VQA-style tasks, GPT-OSS served as the judge for answer correctness.