CoolFace
8 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Smith42 /minty-astro-ph MINT-1T ArXiv Astro-ph An astronomy-focused subset of mlfoundations/MINT-1T-ArXiv, filtered to include only papers from the astro-ph arXiv category (including cross-listed papers). Overview Papers ~845k Total size ~804 GB Format WebDataset tar shards Shards 287 (astro-ph-00000.tar to astro-ph-00286.tar) Shard size ~3 GB each Source MINT-1T (Awadalla et al., 2024) Data Format Each tar shard contains paired files per paper:… See the full description on the dataset page: https://huggingface.co/datasets/Smith42/minty-astro-ph.imagetext-generation100K<n<1M1 likes4.6k downloads5mo agoHugging Face02recursal /reprocessed_singapore_national_speech_corpus Dataset Card for Reprocessed National Speech Corpus NOTE: This is an Reprocessed version KaraKaraWitch from Recursal.The official download can be found here. Dataset Details Dataset Description Dataset Description: The National Speech Corpus (NSC) is the first large-scale Singapore English corpus, sponsored by the Info-communications and Media Development Authority (IMDA) of Singapore. The objective is to serve as a primary resource of open speech data for… See the full description on the dataset page: https://huggingface.co/datasets/recursal/reprocessed_singapore_national_speech_corpus.audiotext-generation1M<n<10M7 likes362 downloads2y agoHugging Face03segyges /OpenWebText2 Dataset Card for OpenWebText2 OpenWebText2 is a reasonably large corpus of scraped natural language data. Original hosting for this dataset has become difficult because it was hosted alongside another controversial dataset. To the best of my knowledge, this dataset itself is not encumbered in any way. It's a useful size for smaller language modelling experiments and is sometimes used in existing papers which it may be desirable to replicate. It is uploaded here to facilitate those… See the full description on the dataset page: https://huggingface.co/datasets/segyges/OpenWebText2.texttext-generationn<1K18 likes132 downloads2y agoHugging Face04YangyiYY /VLM-SFTimagetext-generation1M<n<10M2 likes70 downloads2y agoHugging Face051-800-SHARED-TASKS /xlsum-subset Dataset Card for "XL-Sum" Dataset Summary We present XLSum, a comprehensive and diverse dataset comprising 1.35 million professionally annotated article-summary pairs from BBC, extracted using a set of carefully designed heuristics. The dataset covers 45 languages ranging from low to high-resource, for many of which no public dataset is currently available. XL-Sum is highly abstractive, concise, and of high quality, as indicated by human and intrinsic evaluation.… See the full description on the dataset page: https://huggingface.co/datasets/1-800-SHARED-TASKS/xlsum-subset.textsummarizationn<1K0 likes59 downloads2y agoHugging Face06AlexWortega /physics-scenarios-packed physics-scenarios-packed Packed (tar.gz) version of a 2D rigid body physics dataset for training language models on next-frame prediction. 1,000,020 scenes × 200 frames simulated with Pymunk / Chipmunk2D. This repo is bandwidth-friendly: each scenario type ships as a single .tar.gz. For the unpacked JSONL files see physics-scenarios-raw. Scale Train: 900,000 scenes (24 seen scenario types × 37,500 each) Val: 100,020 scenes (30 scenario types × 3,334 each — includes 6… See the full description on the dataset page: https://huggingface.co/datasets/AlexWortega/physics-scenarios-packed.texttext-generation100K<n<1M0 likes12 downloads5mo agoHugging Face07AlexWortega /physics-scenarios-raw physics-scenarios-raw Raw (un-tarred) JSONL version of the 2D rigid body physics dataset. Each scene is a separate file under <split>/<scenario_type>/scene_<id>.jsonl. Streaming-friendly for HF datasets and curriculum sampling. For the bandwidth-efficient packaged version, see physics-scenarios-packed. Scale (this snapshot) Train: 80,000 scenes Val: 10,000 scenes Test: 10,000 scenes Frames per scene: 200 Format: one .jsonl per scene (1 header + 200 frame lines) This… See the full description on the dataset page: https://huggingface.co/datasets/AlexWortega/physics-scenarios-raw.texttext-generation100K<n<1M0 likes11 downloads5mo agoHugging Face08cheryramneg /Danbooru2021-SQLite Danbooru 2021 SQLite Dataset Summary This is the metadata of danbooru 2021 dataset in SQLite format. https://gwern.net/danbooru2021 Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation Curation… See the full description on the dataset page: https://huggingface.co/datasets/cheryramneg/Danbooru2021-SQLite.imagetext-generation1M<n<10M0 likes4 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.