Skip to main content
All Reviews
AstronomyWorth Reading
intermediate

Achieving GPT-4o level performance in astronomy with a specialized 8B-parameter large language model

Haan, Tijmen de et al. (2025)

Published
Apr 21, 2025
Journal
Scientific Reports · Vol. 15 · No. 1
DOI
10.1038/s41598-025-97131-y

At a GlanceAI

AstroSage, an 8B astronomy LLM, matches GPT-4o on AstroMLab-1 through domain-specific training.

SummaryAI

AstroSage-Llama-3.1-8B is a freely available language model specialized for astronomy, astrophysics, cosmology, and astronomical instrumentation. It was trained on astronomy arXiv papers, additional astronomical literature, and millions of synthetic question-answer pairs, and achieves 80.9% on the AstroMLab-1 benchmark. Its reported GPT-4o-level performance shows that targeted domain training can allow relatively small open models to rival much larger general-purpose systems, potentially broadening AI support for astronomy education and research.

Method SnapshotAI

Domain-specializing an 8B-parameter Llama model using astronomy literature and synthetic question-answer data, then benchmarking it on AstroMLab-1.

BackgroundAI

Basic familiarity with large language models, scientific benchmarks, and astronomy research literature.

An important conclusion and point from this research is that you don't need a frontier model to achieve acceptable results!

ES