Skip to main content
All Reviews
AstronomyNiche
intermediate

AstroMMBench: A Benchmark for Evaluating Multimodal Large Language Models Capabilities in Astronomy

Shi, Jinghang et al. (2025)

Published
Sep 29, 2025
Journal
arXiv (Cornell University)
DOI
10.48550/arXiv.2510.00063

At a GlanceAI

AstroMMBench tests multimodal LLMs on 621 expert-reviewed astronomy image questions across six subfields.

SummaryAI

AstroMMBench addresses the gap between general-purpose multimodal AI benchmarks and the specialized visual reasoning needed in astronomy. It provides 621 expert-reviewed multiple-choice questions spanning six astrophysical subfields, then compares 25 open- and closed-source multimodal language models. Ovis2-34B achieved the highest reported overall accuracy of 70.5%, while large performance differences across subfields—especially cosmology and high-energy astrophysics—show where models remain unreliable. The benchmark offers a domain-specific resource for measuring and improving AI systems intended for astronomical research.

Method SnapshotAI

The authors build an expert-curated astronomy image benchmark and evaluate 25 multimodal large language models on it.

BackgroundAI

Basic familiarity with astronomy image interpretation and multimodal large language models is helpful.