datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
minty-astro-ph
MINT-1T ArXiv Astro-ph
An astronomy-focused subset of mlfoundations/MINT-1T-ArXiv, filtered to include only papers from the astro-ph arXiv category (including cross-listed papers).
Overview
Papers
~845k
Total size
~804 GB
Format
WebDataset tar shards
Shards
287 (astro-ph-00000.tar to astro-ph-00286.tar)
Shard size
~3 GB each
Source
MINT-1T (Awadalla et al., 2024)
Data Format
Each tar shard contains paired files per paper:… See the full description on the dataset page: https://huggingface.co/datasets/Smith42/minty-astro-ph.reprocessed_singapore_national_speech_corpus
Dataset Card for Reprocessed National Speech Corpus
NOTE: This is an Reprocessed version KaraKaraWitch from Recursal.The official download can be found here.
Dataset Details
Dataset Description
Dataset Description:
The National Speech Corpus (NSC) is the first large-scale Singapore English corpus, sponsored by the Info-communications and Media Development Authority (IMDA) of Singapore. The objective is to serve as a primary resource of open speech data for… See the full description on the dataset page: https://huggingface.co/datasets/recursal/reprocessed_singapore_national_speech_corpus.OpenWebText2
Dataset Card for OpenWebText2
OpenWebText2 is a reasonably large corpus of scraped natural language data.
Original hosting for this dataset has become difficult because it was hosted alongside another controversial dataset. To the best of my knowledge, this dataset itself is not encumbered in any way. It's a useful size for smaller language modelling experiments and is sometimes used in existing papers which it may be desirable to replicate. It is uploaded here to facilitate those… See the full description on the dataset page: https://huggingface.co/datasets/segyges/OpenWebText2.VLM-SFTxlsum-subset
Dataset Card for "XL-Sum"
Dataset Summary
We present XLSum, a comprehensive and diverse dataset comprising 1.35 million professionally annotated article-summary pairs from BBC, extracted using a set of carefully designed heuristics. The dataset covers 45 languages ranging from low to high-resource, for many of which no public dataset is currently available. XL-Sum is highly abstractive, concise, and of high quality, as indicated by human and intrinsic evaluation.… See the full description on the dataset page: https://huggingface.co/datasets/1-800-SHARED-TASKS/xlsum-subset.physics-scenarios-packed
physics-scenarios-packed
Packed (tar.gz) version of a 2D rigid body physics dataset for training language models on next-frame prediction. 1,000,020 scenes × 200 frames simulated with Pymunk / Chipmunk2D.
This repo is bandwidth-friendly: each scenario type ships as a single .tar.gz. For the unpacked JSONL files see physics-scenarios-raw.
Scale
Train: 900,000 scenes (24 seen scenario types × 37,500 each)
Val: 100,020 scenes (30 scenario types × 3,334 each — includes 6… See the full description on the dataset page: https://huggingface.co/datasets/AlexWortega/physics-scenarios-packed.physics-scenarios-raw
physics-scenarios-raw
Raw (un-tarred) JSONL version of the 2D rigid body physics dataset. Each scene is a separate file under <split>/<scenario_type>/scene_<id>.jsonl. Streaming-friendly for HF datasets and curriculum sampling.
For the bandwidth-efficient packaged version, see physics-scenarios-packed.
Scale (this snapshot)
Train: 80,000 scenes
Val: 10,000 scenes
Test: 10,000 scenes
Frames per scene: 200
Format: one .jsonl per scene (1 header + 200 frame lines)
This… See the full description on the dataset page: https://huggingface.co/datasets/AlexWortega/physics-scenarios-raw.Danbooru2021-SQLite
Danbooru 2021 SQLite
Dataset Summary
This is the metadata of danbooru 2021 dataset in SQLite format.
https://gwern.net/danbooru2021
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation
Curation… See the full description on the dataset page: https://huggingface.co/datasets/cheryramneg/Danbooru2021-SQLite.
