这篇论文测了Claude、GPT和Gemini最新版,看图能力比人强,但让它们判断图表有没有骗人,还是不行。有意思的发现。
最新研究测试了Anthropic Claude Opus 4.5、OpenAI GPT 5.2 Pro和Google Gemini 3 Flash在可视化评估上的能力。使用修改后的VLAT测试,发现这三款模型的可视化素养均超过人类平均水平。但在指令遵循方面,few-shot和chain-of-thought提示技术对提升可视化素养已无明显效果。在识别误导性可视化时,无专门提示下模型准确率偏低。结论认为LLM作为可视化评估者的能力仍需重新审视。
LLMs have Visualization Literacy: Now What? Experiments Exploring LLM Visualization Evaluation Capabilities
As Large Language Models (LLMs) become more popular within the visualization community, researchers increasingly leverage them for diverse visualization tasks such as design guideline suggestions and visualization evaluation. However, in order for LLMs to act as trustworthy and fair evaluators, we argue that LLMs would need to possess visualization literacy, be capable of following user instructions and uphold graphical integrity. We test the latest versions of the most prominent LLMs, specifically Anthropic's Claude (Opus 4.5), OpenAI's Generative Pretrained Transformers (GPT 5.2 Pro), and Google's Gemini (Gemini 3 Flash) on these features and find that while these models now possess visualization literacy, they still struggle with other features necessary for instruction following and graphical integrity. Using a modified Visualization Literacy Assessment Test (VLAT), our findings show that these recent LLMs have achieved greater than human-levels of visualization literacy in contrast to prior research. In order to test the models' abilities to follow instructions, we used few-shot and chain-of-thought prompting as proxies for instruction following tasks on evaluating visualization literacy and find that these specialized prompting techniques are becoming obsolete with respect to improving visualization literacy. Additionally, we experiment with the inherent ability of LLMs to evaluate misleading visualizations to test the models' abilities for upholding graphical integrity and find that without specialized or leading prompting techniques, the models struggle with being able to accurately identify whether a visualization is misleading or not. Our results further break down the performance of each model on these tasks, but the culmination of our findings force us to reconsider the current effectiveness of LLMs as visualization evaluators.