Skip to main content
All Reviews
AstronomyWorth Reading
beginner

At First Sight! Zero-shot Classification of Astronomical Images with Large Multimodal Models

Tanoglidis, Dimitrios & Jain, Bhuvnesh (2024)

Published
Oct 21, 2024
Journal
Research Notes of the AAS · Vol. 8 · No. 10
DOI
10.3847/2515-5172/ad887a

At a GlanceAI

GPT-4o and LLaVA-NeXT classify galaxy images and artifacts above 80% accuracy using prompts alone.

SummaryAI

The study tests whether large vision-language models can classify astronomical images without being trained on astronomy-specific labels. GPT-4o and the open-source LLaVA-NeXT achieve typically above 80% accuracy for low-surface-brightness galaxies, artifacts, and galaxy morphology using natural-language prompts. The results position multimodal LLMs as potentially useful research and teaching tools, while highlighting that open models such as LLaVA-NeXT still need improvement and may benefit from domain-specific fine-tuning.

Method SnapshotAI

Evaluating GPT-4o and LLaVA-NeXT with natural-language prompts for zero-shot classification of astronomical images.

BackgroundAI

Basic knowledge of galaxy morphology, astronomical imaging, and vision-language models is helpful.

Yep, that's the point: no training, no code, just use. Similar to Smirnov (2024).

ES