Skip to main content
All Reviews
AstronomyNiche
intermediate

Designing an Evaluation Framework for Large Language Models in Astronomy Research

Wu, John F. et al. (2024)

Published
May 30, 2024
Journal
arXiv (Cornell University)
DOI
10.48550/arXiv.2405.20389

At a GlanceAI

A proposed framework evaluates how astronomers use and assess an arXiv-grounded RAG chatbot in real research settings.

SummaryAI

The paper addresses the lack of a standard way to evaluate LLM tools specifically for astronomy research. It presents an experimental design built around a Slack chatbot whose answers are grounded in astronomy papers on arXiv, while recording anonymized questions, answers, retrieval evidence, ratings, and feedback. This setup supports ongoing, real-world evaluation of usefulness and failure modes, helping guide the development of more reliable LLM systems for astronomers.

Method SnapshotAI

Deploying a Slack-based RAG chatbot grounded in arXiv astronomy papers and collecting anonymized interaction and feedback data.

BackgroundAI

Basic knowledge of large language models, retrieval-augmented generation, and astronomy research literature.