datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
marathi-jyotish-astronomy
🕉️ Vedic Neural Geometry & Jyotish QA Dataset
📊 Dataset Overview
Property
Value
Configs (subsets)
14 domain-specific + 1 original QA config (v2.1)
Languages
Marathi, Sanskrit, English (mixed)
Format
CSV / Parquet, structured tabular
License
MIT
Domain
Vedic Geometry, Jyotish, Numerology, Vastu, Samudrika Shastra
Each config below is an independent, domain-specific table — schemas differ across
domains by design (e.g. a KP event-promise… See the full description on the dataset page: https://huggingface.co/datasets/kalpesh77/marathi-jyotish-astronomy.survey-sim-dataPar-Four-Fineweb-Edu-Fortified-Chemistry-Physics-Astronomy-Math-ReasonExtracted Chemistry Physics Asronomy Math and Logic portions from the original.
Script used for the extraction:
https://huggingface.co/datasets/Josephgflowers/Par-Four-Fineweb-Edu-Fortified-Chemistry-Physics-Astronomy-Math-Reason/resolve/main/find-science-fine.py
arxiv-astronomy-similarity
Dataset card for ArXiv Astronomy Similarity
This dataset is a random sample of ArXiv abstracts labeled as astro-ph. It is intended for training vector similarity models.
It also has a BEIR-compatible version of the test split that can be used to measure the accuracy of vector models trained with this data.
astronomy-video-text-benchmark
Astronomy Video Text Data Notes
Dataset summary
Preparation notes and schema examples for Astronomy tasks using Video Text data. Full source material is intentionally not bundled, so provenance and licensing remain explicit.
Included material
dataloader.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable records for checking the schema.
README.md — data… See the full description on the dataset page: https://huggingface.co/datasets/ntnguyenke/astronomy-video-text-benchmark.fiesta_training_dataThis repo contains training data for the fiesta surrogates.
The surrogates are based on
possis (https://arxiv.org/abs/2211.14348)
afterglowpy (https://github.com/geoffryan/afterglowpy, https://arxiv.org/abs/1909.11691)
pyblastafterglow (https://github.com/vsevolodnedora/PyBlastAfterglowMag/, https://arxiv.org/abs/2409.16852)
task666_mmmlu_answer_generation_astronomy
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task666_mmmlu_answer_generation_astronomy
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task666_mmmlu_answer_generation_astronomy.astronomymirabest-radio-astronomy-unofficial
MiraBest Radio Astronomy Dataset (Unofficial)
⚠️ IMPORTANT: This is an unofficial repository containing a processed version of the MiraBest dataset formatted for stable diffusion fine-tuning. This repository is not affiliated with the original authors.
Unofficial processing of the MiraBest radio astronomy dataset with original classification labels and natural language captions for diffusion fine-tuning. Original dataset by Porter & Scaife (2023).
Original Dataset
The… See the full description on the dataset page: https://huggingface.co/datasets/kwazzi-jack/mirabest-radio-astronomy-unofficial.Astronomy_Exoplanetastronomy-samples
Astronomy Audio Text Data Notes
Dataset summary
This data card accompanies a lightweight Astronomy loader for Audio Text metadata. It is meant for pipeline inspection, source adaptation, and reproducible split preparation.
Included material
prepare.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable records for checking the schema.
README.md — data card… See the full description on the dataset page: https://huggingface.co/datasets/tonyzhu91/astronomy-samples.astronomy-data-2023
Astronomy Image Text Data Notes
Dataset summary
This repository contains a preparation pipeline and a small metadata sample for Astronomy work with Image Text inputs. It does not claim to be a complete benchmark release; the loader documents how source data is normalized and validated.
Included material
build_dataset.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small… See the full description on the dataset page: https://huggingface.co/datasets/nikhilsharma98/astronomy-data-2023.dataset_034267395_astronomy_sensor_fusion
dataset_034267395_astronomy_sensor_fusion.py
Dataset Summary
A astronomy dataset with sensor fusion modality, stored in webdataset format.
Preprocessing & Augmentation
Preprocessing: aggressive
Augmentation: randaugment
Splits & Sampling
Split strategy: leave one out
Sampling: contrastive
Quality & Labeling
Quality filtering: moderate
Labeling: manual
Files… See the full description on the dataset page: https://huggingface.co/datasets/fengjchen/dataset_034267395_astronomy_sensor_fusion.astronomy-data
Astronomy Sensor Fusion Data Notes
Dataset summary
This data card accompanies a lightweight Astronomy loader for Sensor Fusion metadata. It is meant for pipeline inspection, source adaptation, and reproducible split preparation.
Included material
prepare.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable records for checking the schema.
README.md — data… See the full description on the dataset page: https://huggingface.co/datasets/pranavit/astronomy-data.dataset_031508479_astronomy_video_text
dataset_031508479_astronomy_video_text.py
Dataset Summary
A astronomy dataset with video text modality, stored in tfrecord format.
Preprocessing & Augmentation
Preprocessing: standard
Augmentation: heavy
Splits & Sampling
Split strategy: random 90 10
Sampling: weighted
Quality & Labeling
Quality filtering: lenient
Labeling: pseudo label
Files
dataset_031508479_astronomy_video_text.py — main… See the full description on the dataset page: https://huggingface.co/datasets/Doyunseo/dataset_031508479_astronomy_video_text.Baike-Astronomy-ZH天文学百科,包含 8 个子目录,约 1000 条词条、110,0000 个字符。
数据包含一级目录、二级目录、标题、内容。其中内容已经处理为单行,且文本普遍较长。
一个样例如下:
{
"top_category": "天文学",
"sub_category": "天体力学",
"title": "万有引力定律",
"content": "万有引力定律(汉语拼音:wàn yǒu yǐn lì zhī dìng lǜ),(universal gravitation,law of),自然界中任何两个质点都相互吸引,这个力同两个质点的质量的乘积成正比,同它们之间的距离的二次方成反比。如用m1、m2表示两质点的质量,r表示两质点间的距离,F表示作用力的值,则F=Gm1m2/r2,式中的G是比例常量,称万有引力常量或牛顿引力常量,数值因不同单位制而异,在国际单位制中G为6.672×1011牛顿·米2/千克2。这个定律由牛顿于1687年在《原理》上首次发表,它和牛顿运动定律一起,构成了牛顿力学特别是天体力学的基础。\n… See the full description on the dataset page: https://huggingface.co/datasets/Mxode/Baike-Astronomy-ZH.astronomy-pointcloud-text-clean
Astronomy Pointcloud Text Data Notes
Dataset summary
This repository contains a preparation pipeline and a small metadata sample for Astronomy work with Pointcloud Text inputs. It does not claim to be a complete benchmark release; the loader documents how source data is normalized and validated.
Included material
prepare.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small… See the full description on the dataset page: https://huggingface.co/datasets/felmend86/astronomy-pointcloud-text-clean.astronomy-pointcloud-text-mini
Astronomy Pointcloud Text Data Notes
Dataset summary
This repository contains a preparation pipeline and a small metadata sample for Astronomy work with Pointcloud Text inputs. It does not claim to be a complete benchmark release; the loader documents how source data is normalized and validated.
Included material
loader.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small… See the full description on the dataset page: https://huggingface.co/datasets/JacobSanc/astronomy-pointcloud-text-mini.toy-astronomy
Astronomy Audio Text Data Notes
Dataset summary
Preparation notes and schema examples for Astronomy tasks using Audio Text data. Full source material is intentionally not bundled, so provenance and licensing remain explicit.
Included material
load_data.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable records for checking the schema.
README.md — data… See the full description on the dataset page: https://huggingface.co/datasets/hajunpark/toy-astronomy.cs224n-astronomy
Astronomy Image Text Data Notes
Dataset summary
This data card accompanies a lightweight Astronomy loader for Image Text metadata. It is meant for pipeline inspection, source adaptation, and reproducible split preparation.
Included material
build_dataset.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable records for checking the schema.
README.md — data… See the full description on the dataset page: https://huggingface.co/datasets/kavitasing/cs224n-astronomy.dataset_031546470_astronomy_video_text
dataset_031546470_astronomy_video_text.py
Dataset Summary
A astronomy dataset with video text modality, stored in tfrecord format.
Preprocessing & Augmentation
Preprocessing: aggressive
Augmentation: light
Splits & Sampling
Split strategy: leave one out
Sampling: balanced
Quality & Labeling
Quality filtering: moderate
Labeling: manual
Files
dataset_031546470_astronomy_video_text.py — main artifact… See the full description on the dataset page: https://huggingface.co/datasets/satos-hisasa/dataset_031546470_astronomy_video_text.mm-astronomyA set of NER-related questions about multimessenger astronomy.
astronomy-dataset
Astronomy Multimodal3 Data Notes
Dataset summary
This repository contains a preparation pipeline and a small metadata sample for Astronomy work with Multimodal3 inputs. It does not claim to be a complete benchmark release; the loader documents how source data is normalized and validated.
Included material
build_dataset.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small… See the full description on the dataset page: https://huggingface.co/datasets/yangyang7626/astronomy-dataset.astronomy-samples
Astronomy Sensor Fusion Data Notes
Dataset summary
This repository contains a preparation pipeline and a small metadata sample for Astronomy work with Sensor Fusion inputs. It does not claim to be a complete benchmark release; the loader documents how source data is normalized and validated.
Included material
dataset.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small… See the full description on the dataset page: https://huggingface.co/datasets/bryanphanbid/astronomy-samples.astronomy-collection
Astronomy Multimodal3 Data Notes
Dataset summary
This data card accompanies a lightweight Astronomy loader for Multimodal3 metadata. It is meant for pipeline inspection, source adaptation, and reproducible split preparation.
Included material
build_dataset.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable records for checking the schema.
README.md — data… See the full description on the dataset page: https://huggingface.co/datasets/khanhvkv95/astronomy-collection.astronomy-text-tabular
Astronomy Text Tabular Data Notes
Dataset summary
This repository contains a preparation pipeline and a small metadata sample for Astronomy work with Text Tabular inputs. It does not claim to be a complete benchmark release; the loader documents how source data is normalized and validated.
Included material
loader.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small… See the full description on the dataset page: https://huggingface.co/datasets/amyhughes/astronomy-text-tabular.astronomy-image-text-benchmark
Astronomy Image Text Data Notes
Dataset summary
Preparation notes and schema examples for Astronomy tasks using Image Text data. Full source material is intentionally not bundled, so provenance and licensing remain explicit.
Included material
prepare.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable records for checking the schema.
README.md — data card… See the full description on the dataset page: https://huggingface.co/datasets/myrasingheli/astronomy-image-text-benchmark.astronomy-stack-dpo-textastronomy-text-tabular-mini
Astronomy Text Tabular Data Notes
Dataset summary
Preparation notes and schema examples for Astronomy tasks using Text Tabular data. Full source material is intentionally not bundled, so provenance and licensing remain explicit.
Included material
dataloader.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable records for checking the schema.
README.md —… See the full description on the dataset page: https://huggingface.co/datasets/irynakov86/astronomy-text-tabular-mini.astronomy-sensor-fusion-clean64
Astronomy Sensor Fusion Data Notes
Dataset summary
This repository contains a preparation pipeline and a small metadata sample for Astronomy work with Sensor Fusion inputs. It does not claim to be a complete benchmark release; the loader documents how source data is normalized and validated.
Included material
loader.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small… See the full description on the dataset page: https://huggingface.co/datasets/Felixinoue/astronomy-sensor-fusion-clean64.
