【中文】论文速递|AREX:面向深度研究的递归自我改进智能体(智源研究院)
智源研究院(BAAI)提出 AREX,核心洞察是深度研究任务中“找到答案代价高、验证答案却可分解为逐条约束检查”的不对称性。AREX 在检索证据与审计候选答案是否违反约束之间交替迭代,实现递归自我改进,并配备自主上下文更新工具,把交互历史压缩为已验证证据与未解决约束。提供 4B 稠密与 122B-A10B MoE 两个版本,在 BrowseComp、WideSearch、DeepSearchQA、Humanity's Last Exam 等基准上大幅超越同规模基线,权重已开源。
【EN】Paper Brief | AREX: Towards a Recursively Self-Improving Agent for Deep Research (BAAI)
BAAI's AREX builds on an asymmetry in deep research: discovering an answer is costly, but verifying one decomposes into tractable constraint-wise checks. The agent recursively self-improves by alternating between researching evidence and auditing provisional answers for constraint violations, aided by an autonomous context-update tool that compresses interaction history into verified evidence and unresolved constraints. Released in 4B dense and 122B-A10B MoE variants, it substantially outperforms comparable-scale baselines on BrowseComp, WideSearch, DeepSearchQA, and Humanity's Last Exam, with weights publicly available.
来源 Source:
留言
NO REGISTRATION · 临时网名 + 邮箱即可开聊