0

StarEmbed: Benchmarking Time Series Foundation Models on Astronomical Observations of Variable Stars

Current time series foundation model (TSFM) training corpora largely omit data with certain complexities like irregular temporal sampling. Astronomical time series of stellar fluxes (light curves) are available in immense quantities and exhibit irregular sampling, multiple…

Preview
Year
2025
Hosting
Full text hostedCC-BY-4.0

Cite

Notes

Only stored in your browser.

Attribution

Abstract & full text
arxiv.org/abs/2510.06200CC-BY-4.0
TL;DR
Semantic Scholar
Attribution policy →

Abstract

Current time series foundation model (TSFM) training corpora largely omit data with certain complexities like irregular temporal sampling. Astronomical time series of stellar fluxes (light curves) are available in immense quantities and exhibit irregular sampling, multiple variates, and heteroskedasticity. We introduce StarEmbed, the first public benchmark for light curves comprised of real observations of 40,000 stars expert-labeled across seven classes and evaluations in clustering, classification, and out-of-distribution (OOD) source detection. We benchmark TSFMs with differing architecture and training strategies as well as domain-specific transformers. Our results demonstrate that the Chronos family, despite being pre-trained on regularly sampled non-astronomical data, yields state-of-the-art (SOTA) performance in light curve clustering and OOD detection. While no TSFM strictly surpasses the classification performance of the long-established domain baseline, they do demonstrate excellent generalization abilities. StarEmbed marks a step toward universal light curve embeddings and improved TSFM performance on challenging data.