Skip to main content
All Collections

Astronomy

LLMs in dynamical (and general) astronomy

No spamNo marketingOnly data
ES

Evgeny Smirnov

10 papers · 2 Must Read · 2023–2026

Last updated Aug 17, 2026

All papers in the expert’s recommended reading order. The full collection as the expert intended it.

Introduction

How large language models are transforming astronomical research in celestial mechanics and dynamical astronomy.

At a Glance

LLM can classify resonant behavior without coding or knowledge with high accuracy

Summary

The author created LLM pipeline with just one prompt that classifies resonant arguments and identifies whether time series has libration or circulation. This is a pilot study and simple cases are considered, but the accuracy and F1 score of 100% are remarkable. What's more important: instead of coding and using special methods, it takes 15m to write a prompt and use it.

Forget about filtering, periodograms, and working with time series. Just use LLMs!

ES

Method:
LLM
Background:
Basic knowledge of LLMs

At a GlanceAI

Benchmark shows multimodal LLMs can classify orbital resonances from images with high accuracy without fine-tuning.

SummaryAI

The study introduces reproducible benchmarks for testing multimodal LLMs on classifying mean-motion and secular resonances from plots of resonant arguments. Commercial models achieved perfect scores on unambiguous cases and up to 94% F1 on a harder three-class dataset, while open-source models approached commercial performance on the full binary benchmark. The main weakness was identifying transient and resonance-sticking behavior, but the results indicate that even untuned, locally runnable models can be useful for dynamical-astronomy classification.

Have no resources to use the commercial models? Use OSS! You can even launch them on your laptop. And they ARE good.

ES

Method:AI
Benchmarking multimodal commercial and open-source LLMs on image-based classification of dynamical resonance behavior.
Background:AI
Basic knowledge of orbital dynamics, mean-motion and secular resonances, and machine-learning evaluation metrics.
3
Niche
advanced
★ Essential

Wu, Di, Zhang, Raymond, Zucchelli, Enrico M. et al. · 2025 · Scientific Reports

At a Glance

How good can LLMs solve space science university-level problems

Summary

Authors created a dataset of questions from Astrodynamics, tested a variety of LLMs including open-source ones on them, and evaluated their performance. The paper is a good example of a benchmark study and how to conduct it. Helpful for anyone doing benchmark stuff in astronomy.

A nice example of the usefulness of LLMs in astronomy + a good example of how to do a benchmark study in astronomy + LLM

ES

Method:
LLM (benchmark)
Background:
Deep knowledge of benchmarking in LLMs + the state of art
4
Niche
beginner

Smith, Michael J., Geach, James E. · 2023 · Royal Society Open Science

At a GlanceAI

Review advocating open, astronomy-specific GPT-like foundation models for multimodal astronomical data.

SummaryAI

This review charts three waves of neural networks in astronomy, from early multilayer perceptrons through convolutional and recurrent networks to unsupervised and generative deep learning. It argues that rapidly growing, multimodal astronomical datasets make GPT-like foundation models a promising next step for supporting many downstream astronomy tasks. The authors propose that the astronomy community collaboratively develop open-source foundation models rather than relying solely on systems driven by large technology companies.

Has some nice theoretical stuff.

ES

Method:AI
A historical and forward-looking review of neural-network methods in astronomy, culminating in a proposal for astronomy-specific foundation models.
Background:AI
Basic knowledge of machine learning, neural networks, and astronomical data analysis is helpful.
5
Worth Reading
intermediate

Haan, Tijmen de, Ting, Yuan-Sen, Ghosal, Tirthankar et al. · 2025 · Scientific Reports

At a GlanceAI

AstroSage, an 8B astronomy LLM, matches GPT-4o on AstroMLab-1 through domain-specific training.

SummaryAI

AstroSage-Llama-3.1-8B is a freely available language model specialized for astronomy, astrophysics, cosmology, and astronomical instrumentation. It was trained on astronomy arXiv papers, additional astronomical literature, and millions of synthetic question-answer pairs, and achieves 80.9% on the AstroMLab-1 benchmark. Its reported GPT-4o-level performance shows that targeted domain training can allow relatively small open models to rival much larger general-purpose systems, potentially broadening AI support for astronomy education and research.

An important conclusion and point from this research is that you don't need a frontier model to achieve acceptable results!

ES

Method:AI
Domain-specializing an 8B-parameter Llama model using astronomy literature and synthetic question-answer data, then benchmarking it on AstroMLab-1.
Background:AI
Basic familiarity with large language models, scientific benchmarks, and astronomy research literature.
6
Worth Reading
beginner

Tanoglidis, Dimitrios, Jain, Bhuvnesh · 2024 · Research Notes of the AAS

At a GlanceAI

GPT-4o and LLaVA-NeXT classify galaxy images and artifacts above 80% accuracy using prompts alone.

SummaryAI

The study tests whether large vision-language models can classify astronomical images without being trained on astronomy-specific labels. GPT-4o and the open-source LLaVA-NeXT achieve typically above 80% accuracy for low-surface-brightness galaxies, artifacts, and galaxy morphology using natural-language prompts. The results position multimodal LLMs as potentially useful research and teaching tools, while highlighting that open models such as LLaVA-NeXT still need improvement and may benefit from domain-specific fine-tuning.

Yep, that's the point: no training, no code, just use. Similar to Smirnov (2024).

ES

Method:AI
Evaluating GPT-4o and LLaVA-NeXT with natural-language prompts for zero-shot classification of astronomical images.
Background:AI
Basic knowledge of galaxy morphology, astronomical imaging, and vision-language models is helpful.
LLMs in dynamical (and general) astronomy | Marginalia