【中文】论文速递|腾讯 WorkBuddy Bench:抗数据污染的多领域编码智能体基准
腾讯发布 WorkBuddy Bench,覆盖 Code、Web、Office、Security 四个领域的编码智能体评测套件。针对“基准题目被模型在训练中见过”的数据污染顽疾,任务全部从真实 commit、PR 或业务场景反向构造、改写为口语化需求,使网络检索无法直接命中答案。套件采用统一任务格式与可复现协议,按领域分别计分而非单一总分,并完整开源环境、评测框架、测试与参考解,以可审计性替代保密性来实现抗污染。
【EN】Paper Brief | Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction
Tencent releases WorkBuddy Bench, a coding-agent evaluation suite spanning Code, Web, Office, and Security. To combat contamination — benchmark tasks models have already seen in training — every task is reverse-engineered from a real commit, pull request, or business scenario and rewritten as a colloquial request, so web retrieval can't surface the answer. The suite uses a uniform task format and reproducible protocol, scores each domain separately rather than as one aggregate, and open-sources environments, harness, tests, and solutions — trading secrecy for auditability as its contamination defense.
来源 Source:
留言
NO REGISTRATION · 临时网名 + 邮箱即可开聊