LLMs in dynamical (and general) astronomy
How large language models are transforming astronomical research in celestial mechanics and dynamical astronomy.
Wu, John F. et al. (2024)
A proposed framework evaluates how astronomers use and assess an arXiv-grounded RAG chatbot in real research settings.
The paper addresses the lack of a standard way to evaluate LLM tools specifically for astronomy research. It presents an experimental design built around a Slack chatbot whose answers are grounded in astronomy papers on arXiv, while recording anonymized questions, answers, retrieval evidence, ratings, and feedback. This setup supports ongoing, real-world evaluation of usefulness and failure modes, helping guide the development of more reliable LLM systems for astronomers.
Deploying a Slack-based RAG chatbot grounded in arXiv astronomy papers and collecting anonymized interaction and feedback data.
Basic knowledge of large language models, retrieval-augmented generation, and astronomy research literature.