datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
galaxy-mentions
Galaxy Mentions
A dataset of nearly every galaxy ever discussed in >2 sentences on the arXiv.
Note: We have not yet thoroughly verified or inspected the contents of this
dataset, as always use with skepticism!
Datasets
galaxy_mentions: one row per accepted galaxy mention.
evidence_quotes: one row per supporting quote, linked by mention_id.
papers: one row per processed arXiv paper, including accepted/rejected/error status.
batches: one row per extraction batch… See the full description on the dataset page: https://huggingface.co/datasets/astronolan/galaxy-mentions.astro-mcq
Astro-MCQ Dataset
Astro-MCQ is the first dataset in the upcoming AstroBench collection, a suite of domain-specific benchmark datasets for evaluating small and large language models (SLMs and LLMs) in space mission engineering and astronautics.
Overview
Astro-MCQ is a multiple-choice question dataset designed to evaluate language model performance across key topics in astronautics, including:
Orbital mechanics
Space propulsion
Space environment and its effects
Spacecraft… See the full description on the dataset page: https://huggingface.co/datasets/patrickfleith/astro-mcq.
