LLMs in dynamical (and general) astronomy
How large language models are transforming astronomical research in celestial mechanics and dynamical astronomy.
Shi, Jinghang et al. (2025)
AstroMMBench tests multimodal LLMs on 621 expert-reviewed astronomy image questions across six subfields.
AstroMMBench addresses the gap between general-purpose multimodal AI benchmarks and the specialized visual reasoning needed in astronomy. It provides 621 expert-reviewed multiple-choice questions spanning six astrophysical subfields, then compares 25 open- and closed-source multimodal language models. Ovis2-34B achieved the highest reported overall accuracy of 70.5%, while large performance differences across subfields—especially cosmology and high-energy astrophysics—show where models remain unreliable. The benchmark offers a domain-specific resource for measuring and improving AI systems intended for astronomical research.
The authors build an expert-curated astronomy image benchmark and evaluate 25 multimodal large language models on it.
Basic familiarity with astronomy image interpretation and multimodal large language models is helpful.