datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
generated-csvsLAVIBData for LAVIB: A Large-scale Video Interpolation Benchmark (arxiv link: arxiv.org/abs/2406.09754)
astro_horoscopeastrorag_papersThousandWorlds
ThousandWorlds
ThousandWorlds is a benchmark for emulating exoplanet climates: 1689 simulations across 5 GCMs, 8 planet parameters, and atmospheric
variables on a 32 x 64 x 10 latitude-longitude-pressure grid. It includes three
nested benchmark subsets, two evaluation protocols, and ten released baseline
methods.
Explore the dataset + discovered exoplanets online with the ThousandWorlds Explorer!
Built by Hamza Ali Shahjahan!
Inputs are 8 continuous planet parameters plus… See the full description on the dataset page: https://huggingface.co/datasets/AstroAutomata/ThousandWorlds.Par-Four-Fineweb-Edu-Fortified-Chemistry-Physics-Astronomy-Math-ReasonExtracted Chemistry Physics Asronomy Math and Logic portions from the original.
Script used for the extraction:
https://huggingface.co/datasets/Josephgflowers/Par-Four-Fineweb-Edu-Fortified-Chemistry-Physics-Astronomy-Math-Reason/resolve/main/find-science-fine.py
hls-ast-sagehls
SAGE-HLS Dataset: AST-Guided HLS-C Code Generation
The SAGE-HLS dataset is a large-scale, synthesis-friendly dataset for natural language to HLS-C code generation, enhanced by AST (Abstract Syntax Tree) representations. It supports training LLMs to generate high-quality, synthesis-ready high-level synthesis (HLS) code from functional descriptions, with structure-aware guidance.
📦 Dataset Structure
Each sample in the dataset contains the following fields:
Field… See the full description on the dataset page: https://huggingface.co/datasets/mashnoor/hls-ast-sagehls.astex_diverse_setThe Astex Diverse dataset accompanying the PoseBench manuscript and benchmarking suite.
astra-benchmark
Dataset Card for Astra-Benchmark v1
The dataset used for astra-benchmark v1 consists of multiple project questions, each with its own unique identifier and associated metadata. The dataset is stored in a CSV file named project_questions.csv located in the root directory of the project.
Structure of project_questions.csv
The CSV file should contain the following columns:
id: Unique identifier for each project.
name: Name of the project.
type: Type of project (e.g.… See the full description on the dataset page: https://huggingface.co/datasets/hackerrank/astra-benchmark.Astronomy_Exoplanetastro-llms-benchmark-dataset
AstroLLMs Gold Benchmark Dataset
This dataset is a collection of queries that astronomers asked to an astronomy research Slack chatbot. Along with the questions, there are open coding labels determined by a team of researchers and expert astronomer answers to these queries. Astronomers were asked to respond using citations and without the help of Large Language Models. This dataset of answers and responses is called the "Gold Benchmark Dataset".
Dataset Structure
The… See the full description on the dataset page: https://huggingface.co/datasets/jhu-clsp/astro-llms-benchmark-dataset.Art-GenEvalGPT
Dataset Card
Dataset Details
Dataset Description
The dataset includes synthetic dialogues in the art domain that can be used for training a chatbot to discuss artworks within a museum setting. Leveraging Large Language Models (LLMs), particularly ChatGPT, the dataset comprises over 13,000 dialogues generated using prompt-engineering techniques. The dialogues cover a wide range of user and chatbot behaviors, including expert guidance, tutoring, and handling… See the full description on the dataset page: https://huggingface.co/datasets/Astound/Art-GenEvalGPT.Astro-mcqa
AstroMCQA Dataset
🚨 NEWS 🚨 Check out the new version of this dataset: https://huggingface.co/datasets/patrickfleith/astro-mcq
Purpose and scope
The primary purpose of AstroMCQA is for application developers in the domain of space engineering to be able to comparatively assess LLM performances on the specific task of multiple-choice question-answering
Intended Usage
Comparative assessement of differents LLMs, Model evaluation, audit, and model… See the full description on the dataset page: https://huggingface.co/datasets/patrickfleith/Astro-mcqa.en-fr-datasetConsistency_Point_Astro_GAMEModule Version: Consistency_Point_Astrocyte_20260625-132521_EDT
GAME Schema Version: v 1.0
Github Link: https://github.com/de-Boer-Lab/GAME-consistency-evaluators/tree/main/Consistency_evaluator_point_astrocyte
Additional information can be found on GitHub: Genomic API for Model Evaluation
AstroChat
AstroChat Dataset Description
Purpose and Scope
The AstroChat dataset is a collection of 901 dialogues, synthetically generated, tailored to the specific domain of Astronautics / Space Mission Engineering.
This dataset will be frequently updated following feedback from the community. If you would like to contribute, please reach out in the community discussion.
Intended Use
The dataset is intended to be used for supervised fine-tuning of chat LLMs (Large… See the full description on the dataset page: https://huggingface.co/datasets/patrickfleith/AstroChat.astroastro-llms-full-query-data
AstroLLMs Full Query Dataset
This dataset includes all of the data collected in a four-week deployment of a Large Language Model-powered Slack chatbot trained on astrophysics papers. Astronomers were invited to interact with the chatbot, ask questions, and leave feedback. This data includes 368 question-answer pairs, including feedback, reactions, and labeling.
Dataset Structure
The columns of this dataset are thread_ts (unique time stamp of the query), channel_id… See the full description on the dataset page: https://huggingface.co/datasets/jhu-clsp/astro-llms-full-query-data.AstroCaptionsAstroCaptions is an image captioning dataset made of both human labelled and synthetic captions. AstroCaptions is made of 44115 publicly available NASA archive images.
It contains both very recent photos and old archive pictures from the first Apollo missions. Many astronauts, NASA scientists and executives appear on these images.
Each image comes with a description, scraped from public NASA website. These provides both visual description of the image and contextual
information. The first… See the full description on the dataset page: https://huggingface.co/datasets/momentslab/AstroCaptions.DiorisisThis dataset is based on the Diorisis ancient greek corpus:
Vatri, A., & McGillivray, B. (2018). The Diorisis Ancient Greek Corpus: Linguistics and Literature. Research Data Journal for the Humanities and Social Sciences, 3(1), 55-65. https://doi.org/10.1163/24523666-01000013
AstralBenchAstralBench is a carefully curated subset of 50 high-quality problems, selected for benchmarking model performance. It covers diverse mathematical topics and difficulty levels, with current model performance ranging from 5% to 30% accuracy.
Source of AstralBench
Seed data: IMO AnswerBench, Project Euler, HMMT, SMT, USA-TSTST, USEMO, EGMO, CMO, Pumac, Putnam, open-rl, mit-math
AstralBench problems are selected from various sources. Problems that have non-int and symbolic answers are… See the full description on the dataset page: https://huggingface.co/datasets/nguyen599/AstralBench.AstroPyAutoPdfLenAstro
Stellar Classification Dataset - SDSS17
Welcome to the Dataset!
Get ready to explore the cosmos with the Stellar Classification Dataset from the Sloan Digital Sky Survey (SDSS) Data Release 17! This dataset contains 100,000 observations of celestial objects—stars, galaxies, and quasars—captured through their spectral characteristics. Whether you're an astronomer studying the universe, a data scientist building classification models, or a student curious about the night… See the full description on the dataset page: https://huggingface.co/datasets/AethronPhantom/Astro.dataset-20260112-prime-seed
dataset-20260112-prime-seed
Created on: 2026-01-12T13:01:59.128698+00:00
Session ID: 2026-01-12T13:01:59.128698+00:00-5807
dataset-20251213-simple
dataset-20251213-simple
Created on: 2025-12-13T06:08:24.408128+00:00
Session ID: 2025-12-13T06:08:24.408128+00:00-6243
dataset-20251215-alpha-seed
dataset-20251215-alpha-seed
Created on: 2025-12-15T05:33:47.445983+00:00
Session ID: 2025-12-15T05:33:47.445983+00:00-4565
travel_itineraries_dataset_Indiaastana_traffic
Flowmatic Smart City Dataset
Flowmatic smart-city dataset with hourly UTC CSV partitions. 3 hour file(s), 500 total row(s). Files live under data/hourly/ and are indexed in data/hourly/manifest.json.
Pipeline run: cmpmwqy3j005joo3crz8pe451Updated: 2026-05-26T17:27:13.937ZPartition scheme: hourly-utcManifest: data/hourly/manifest.json
Hourly CSV layout
Rows are grouped by eventTime into one CSV per UTC hour:
Directory: data/hourly/
File pattern: YYYY-MM-DDTHH.csv… See the full description on the dataset page: https://huggingface.co/datasets/pushthetempo/astana_traffic.hardware_prices
