【中文】论文速递|Show, Don't Tell:用生成像素而非文本评测空间认知(浙大 OmniAI)
浙大 OmniAI 团队指出现有空间推理基准存在“答案-接口错配”:图像生成模型天然用像素表达空间理解,却被迫用文本或坐标作答。他们提出 ProVisE 框架,让图像生成模型以视觉方式作答,再将视觉输出转换为可用标准指标评测的结构化预测,并构建含 470 样本、14 个空间子任务的 SpatialGen-Bench,评测 31 个系统。结果显示图像生成模型用像素作答时颇具竞争力——GPT Image 2 能答对 GPT-5.4 答错的空间题中的 37%,而文本输出的 VLM 在组合推理上仍占优。
【EN】Paper Brief | Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels, Not Text (ZJU-OmniAI)
ZJU-OmniAI identifies an answer-interface mismatch in spatial reasoning evaluation: image-generation models naturally express spatial understanding in pixels, yet benchmarks force text or coordinate answers. Their ProVisE framework lets image-generation models answer visually, then converts those outputs into structured predictions scoreable by standard metrics, alongside SpatialGen-Bench (470 samples, 14 spatial subtasks) used to evaluate 31 systems. Image-generation models prove competitive when answering in pixels: GPT Image 2 correctly solves 37% of the spatial cases GPT-5.4 misses, while text-output VLMs retain the edge in compositional reasoning.
来源 Source:
留言
NO REGISTRATION · 临时网名 + 邮箱即可开聊