Evaluating multimodal commercial and open-source large language models for dynamical astronomy: a benchmark study of resonant behavior classification(pdf)
Smirnov, Evgeny, Carruba, Valerio · 2026 · Scientific Reports
At a GlanceAI
Benchmark shows multimodal LLMs can classify orbital resonances from images with high accuracy without fine-tuning.
SummaryAI
The study introduces reproducible benchmarks for testing multimodal LLMs on classifying mean-motion and secular resonances from plots of resonant arguments. Commercial models achieved perfect scores on unambiguous cases and up to 94% F1 on a harder three-class dataset, while open-source models approached commercial performance on the full binary benchmark. The main weakness was identifying transient and resonance-sticking behavior, but the results indicate that even untuned, locally runnable models can be useful for dynamical-astronomy classification.
Have no resources to use the commercial models? Use OSS! You can even launch them on your laptop. And they ARE good.
— ES
- Method:AI
- Benchmarking multimodal commercial and open-source LLMs on image-based classification of dynamical resonance behavior.
- Background:AI
- Basic knowledge of orbital dynamics, mean-motion and secular resonances, and machine-learning evaluation metrics.