--:--
今日 PV 0
独立 IP 0
在线 0
今日实时07/25LIVE
0今日访问
真人访问050%
搜索蜘蛛050%
当前在线0now
0PV 浏览量
0独立 IP
0蜘蛛抓取
24 小时访问趋势暂无数据
小伍的游乐场大数据

Show, Don't Tell:用生成像素而非文本评测空间认知(浙大)| Paper: Evaluating Spatial Cognition in Generative Pixels

燎原Ai Build阅读 1留言 0

【中文】论文速递|Show, Don't Tell:用生成像素而非文本评测空间认知(浙大 OmniAI)

浙大 OmniAI 团队指出现有空间推理基准存在“答案-接口错配”:图像生成模型天然用像素表达空间理解,却被迫用文本或坐标作答。他们提出 ProVisE 框架,让图像生成模型以视觉方式作答,再将视觉输出转换为可用标准指标评测的结构化预测,并构建含 470 样本、14 个空间子任务的 SpatialGen-Bench,评测 31 个系统。结果显示图像生成模型用像素作答时颇具竞争力——GPT Image 2 能答对 GPT-5.4 答错的空间题中的 37%,而文本输出的 VLM 在组合推理上仍占优。

【EN】Paper Brief | Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels, Not Text (ZJU-OmniAI)

ZJU-OmniAI identifies an answer-interface mismatch in spatial reasoning evaluation: image-generation models naturally express spatial understanding in pixels, yet benchmarks force text or coordinate answers. Their ProVisE framework lets image-generation models answer visually, then converts those outputs into structured predictions scoreable by standard metrics, alongside SpatialGen-Bench (470 samples, 14 spatial subtasks) used to evaluate 31 systems. Image-generation models prove competitive when answering in pixels: GPT Image 2 correctly solves 37% of the spatial cases GPT-5.4 misses, while text-output VLMs retain the edge in compositional reasoning.

来源 Source:

https://arxiv.org/abs/2607.21072

https://huggingface.co/papers/2607.21072

留言

NO REGISTRATION · 临时网名 + 邮箱即可开聊

无需注册 · 邮箱仅用于回复通知,绝不公开 · 广告与机器人会被蜜罐直接吞掉

Show, Don't Tell:用生成像素而非文本评测空间认知(浙大)| Paper: Evaluating Spatial Cognition in Generative Pixels | 小伍的游乐场