datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
astro_horoscopeThousandWorlds
ThousandWorlds
ThousandWorlds is a benchmark for emulating exoplanet climates: 1689 simulations across 5 GCMs, 8 planet parameters, and atmospheric
variables on a 32 x 64 x 10 latitude-longitude-pressure grid. It includes three
nested benchmark subsets, two evaluation protocols, and ten released baseline
methods.
Explore the dataset + discovered exoplanets online with the ThousandWorlds Explorer!
Built by Hamza Ali Shahjahan!
Inputs are 8 continuous planet parameters plus… See the full description on the dataset page: https://huggingface.co/datasets/AstroAutomata/ThousandWorlds.Par-Four-Fineweb-Edu-Fortified-Chemistry-Physics-Astronomy-Math-ReasonExtracted Chemistry Physics Asronomy Math and Logic portions from the original.
Script used for the extraction:
https://huggingface.co/datasets/Josephgflowers/Par-Four-Fineweb-Edu-Fortified-Chemistry-Physics-Astronomy-Math-Reason/resolve/main/find-science-fine.py
astrorag_papersAstronomy_Exoplanetastro-llms-benchmark-dataset
AstroLLMs Gold Benchmark Dataset
This dataset is a collection of queries that astronomers asked to an astronomy research Slack chatbot. Along with the questions, there are open coding labels determined by a team of researchers and expert astronomer answers to these queries. Astronomers were asked to respond using citations and without the help of Large Language Models. This dataset of answers and responses is called the "Gold Benchmark Dataset".
Dataset Structure
The… See the full description on the dataset page: https://huggingface.co/datasets/jhu-clsp/astro-llms-benchmark-dataset.Astro-mcqa
AstroMCQA Dataset
🚨 NEWS 🚨 Check out the new version of this dataset: https://huggingface.co/datasets/patrickfleith/astro-mcq
Purpose and scope
The primary purpose of AstroMCQA is for application developers in the domain of space engineering to be able to comparatively assess LLM performances on the specific task of multiple-choice question-answering
Intended Usage
Comparative assessement of differents LLMs, Model evaluation, audit, and model… See the full description on the dataset page: https://huggingface.co/datasets/patrickfleith/Astro-mcqa.AutoPdfLenAstroChat
AstroChat Dataset Description
Purpose and Scope
The AstroChat dataset is a collection of 901 dialogues, synthetically generated, tailored to the specific domain of Astronautics / Space Mission Engineering.
This dataset will be frequently updated following feedback from the community. If you would like to contribute, please reach out in the community discussion.
Intended Use
The dataset is intended to be used for supervised fine-tuning of chat LLMs (Large… See the full description on the dataset page: https://huggingface.co/datasets/patrickfleith/AstroChat.en-fr-datasetConsistency_Point_Astro_GAMEModule Version: Consistency_Point_Astrocyte_20260625-132521_EDT
GAME Schema Version: v 1.0
Github Link: https://github.com/de-Boer-Lab/GAME-consistency-evaluators/tree/main/Consistency_evaluator_point_astrocyte
Additional information can be found on GitHub: Genomic API for Model Evaluation
AstroCaptionsAstroCaptions is an image captioning dataset made of both human labelled and synthetic captions. AstroCaptions is made of 44115 publicly available NASA archive images.
It contains both very recent photos and old archive pictures from the first Apollo missions. Many astronauts, NASA scientists and executives appear on these images.
Each image comes with a description, scraped from public NASA website. These provides both visual description of the image and contextual
information. The first… See the full description on the dataset page: https://huggingface.co/datasets/momentslab/AstroCaptions.astro-llms-full-query-data
AstroLLMs Full Query Dataset
This dataset includes all of the data collected in a four-week deployment of a Large Language Model-powered Slack chatbot trained on astrophysics papers. Astronomers were invited to interact with the chatbot, ask questions, and leave feedback. This data includes 368 question-answer pairs, including feedback, reactions, and labeling.
Dataset Structure
The columns of this dataset are thread_ts (unique time stamp of the query), channel_id… See the full description on the dataset page: https://huggingface.co/datasets/jhu-clsp/astro-llms-full-query-data.astroAstroMMBenchAstro
Stellar Classification Dataset - SDSS17
Welcome to the Dataset!
Get ready to explore the cosmos with the Stellar Classification Dataset from the Sloan Digital Sky Survey (SDSS) Data Release 17! This dataset contains 100,000 observations of celestial objects—stars, galaxies, and quasars—captured through their spectral characteristics. Whether you're an astronomer studying the universe, a data scientist building classification models, or a student curious about the night… See the full description on the dataset page: https://huggingface.co/datasets/AethronPhantom/Astro.AstroPytravel_itineraries_dataset_IndiaTTS_Astro_data
Burmese Mahabote Astrology
မြန်မာဗေဒင်ပညာအတွက် Dataset ရအောင် စတင်စမ်းသပ်ခြင်းဖြစ်ပါသည်။
Astro Dataset v1.0.0:
မဟာဘုတ်ဋ္ဌာနနှင့် ဘွားဇာတာအဟော
Notes:
လောလောဆယ် Dataset များကို ကောင်းစွာ မပြင်ဆင်ရသေးပါ။ ဟောစာတမ်းအတွက် သေချာမွမ်းမံပြင်ဆင်ရပါဦးမည်။
Author
@Guru-ThutaSann
License
Apache License 2.0
SlimOrca_astronomy_8kSlimOrca_astronomySlimOrca_astronomy_5kSlimOrca_astronomy_12kTermes_de_physique_et_astronomie
[!NOTE]
Dataset origin: https://termini.gov.lv/kolekcijas/10
SlimOrca_astronomy_50kSlimOrca_astronomy_40kAstroAIMetadata3SlimOrca_astronomy_25kastrophotonics_spectral_interference_patternsAstroAIMetadata
