datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
minty-astro-ph
MINT-1T ArXiv Astro-ph
An astronomy-focused subset of mlfoundations/MINT-1T-ArXiv, filtered to include only papers from the astro-ph arXiv category (including cross-listed papers).
Overview
Papers
~845k
Total size
~804 GB
Format
WebDataset tar shards
Shards
287 (astro-ph-00000.tar to astro-ph-00286.tar)
Shard size
~3 GB each
Source
MINT-1T (Awadalla et al., 2024)
Data Format
Each tar shard contains paired files per paper:… See the full description on the dataset page: https://huggingface.co/datasets/Smith42/minty-astro-ph.ASTRA
ASTRA: Autonomous Spacecraft Tracking and Rendezvous Archive
Anonymous review copy for our NeurIPS 2026 Datasets and Benchmarks Track submission.
ASTRA is a large-scale multi-satellite benchmark for 6-DoF spacecraft pose estimation, covering twenty NASA-referenced spacecraft rendered in Unreal Engine 5 and paired with FRESCO and NVIDIA COSMOS real-style variants. See the accompanying paper for full details.
Contents
This anonymous release contains the test split of… See the full description on the dataset page: https://huggingface.co/datasets/astra-spacecraft/ASTRA.chatuniviAUDETER
Dataset for AUDETER: A Large-scale Dataset for Deepfake Audio Detection in Open Worlds
Dataset Description
We introduce AUDETER (AUdio DEepfake TEst Range), a large-scale, highly diverse deepfake audio dataset for comprehensive evaluation and robust development of generalised models for deepfake audio detection. It consists of over 4,500 hours of synthetic audio generated by 11 recent TTS models and 10 vocoders with a broad range of TTS/vocoder patterns, totalling 3… See the full description on the dataset page: https://huggingface.co/datasets/astham71/AUDETER.
