【中文】论文速递|微软:LLM 会在“不断变化的用户意图”中迷失
微软研究发现大模型评测的一个盲区:现有基准多为单轮、需求一次说清的设定,而真实协作中用户目标是逐步披露、反复修改甚至中途转向的。团队提出一个转换框架,把现有静态基准改造成“意图随轮次演化”的多轮对话(无需新标注、保留原评测协议)。结果发人深省:在静态设定下表现强劲的模型,一旦用户意图开始演化便出现大幅下滑,且跨模型家族普遍如此——对要做“协作型智能体”的模型来说,追踪动态目标仍是根本短板。
【EN】Paper Brief | LLMs Get Lost in Evolving User Intent (Microsoft)
Microsoft researchers expose a blind spot in LLM evaluation: benchmarks are mostly single-turn and fully specified, while real collaboration involves goals that are incrementally revealed, revised, and sometimes redirected. Their framework converts existing static benchmarks into multi-turn conversations with evolving intent — no new annotation, original protocols preserved. The sobering result: models that excel in static settings suffer substantial drops once intent evolves, consistently across model families — faithfully tracking shifting objectives remains a fundamental gap for would-be collaborative agents.
来源 Source:
留言
NO REGISTRATION · 临时网名 + 邮箱即可开聊