Skip to main content
All Reviews
AstronomyMust Read
intermediate

Evaluating multimodal commercial and open-source large language models for dynamical astronomy: a benchmark study of resonant behavior classification

Smirnov, Evgeny & Carruba, Valerio (2026)

Published
Mar 28, 2026
Journal
Scientific Reports · Vol. 16 · No. 1
DOI
10.1038/s41598-026-45926-y

At a GlanceAI

Benchmark shows multimodal LLMs can classify orbital resonances from images with high accuracy without fine-tuning.

SummaryAI

The study introduces reproducible benchmarks for testing multimodal LLMs on classifying mean-motion and secular resonances from plots of resonant arguments. Commercial models achieved perfect scores on unambiguous cases and up to 94% F1 on a harder three-class dataset, while open-source models approached commercial performance on the full binary benchmark. The main weakness was identifying transient and resonance-sticking behavior, but the results indicate that even untuned, locally runnable models can be useful for dynamical-astronomy classification.

Method SnapshotAI

Benchmarking multimodal commercial and open-source LLMs on image-based classification of dynamical resonance behavior.

BackgroundAI

Basic knowledge of orbital dynamics, mean-motion and secular resonances, and machine-learning evaluation metrics.

Have no resources to use the commercial models? Use OSS! You can even launch them on your laptop. And they ARE good.

ES