CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mlfoundations /dclm-pool-7b-2x3 likes149k downloads2y agoHugging Face02mlfoundations /datacomp_pools DataComp Pools This repository contains metadata files for DataComp. For details on how to use the metadata, please visit our website and our github repository. We distribute the image url-text samples and metadata under a standard Creative Common CC-BY-4.0 license. The individual images are under their own copyrights. Terms and Conditions We have terms of service that are similar to those adopted by HuggingFace (https://huggingface.co/terms-of-service), which covers… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/datacomp_pools.image19 likes87k downloads3y agoHugging Face03allenai /dolma3_pool⚠️ IMPORTANT NOTICE ⚠️ This is the Dolma 3 pool, pre–quality upsampling and mixing. If you are interested in the data used to train Olmo 3 7B and Olmo 3 32B, visit allenai/dolma3_mix-6T-1025. Dolma 3 Pool The Dolma 3 pool is a dataset of over 9 trillion tokens from a diverse mix of web content, academic publications, code, and more. For detailed documenation on Dolma 3 processing and data, please see our Dolma 3 Github repository. For more information on Dolma in general… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolma3_pool.texttext-generation10B<n<100B40 likes74k downloads7mo agoHugging Face04allenai /dolma3.5_pool⚠️ IMPORTANT NOTICE ⚠️ This is the Dolma 3.5 pool. It contains no quality upsampling or mixing. This is an updated version of the Dolma 3 pool with additional quality filtering and more data sources. If you are interested in the data used to train Olmo 3 7B and Olmo 3 32B, visit allenai/dolma3_mix-6T-1025. Dolma 3.5 Pool The Dolma 3.5 pool is a dataset of nearly 10 trillion tokens from a diverse mix of web content, academic publications, code, and more. For detailed… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolma3.5_pool.text-generation8 likes69k downloads2mo agoHugging Face05mlfoundations /dclm-pool-1b-1x3 likes35k downloads2y agoHugging Face06allenai /dolma3_dolmino_pool⚠️ IMPORTANT NOTICE ⚠️ This is the Dolma 3 Dolmino pool; it hasn't been mixed. If you are interested in the data used to train: Olmo 3 7B: allenai/dolma3_dolmino_mix-100B-1025 Olmo 3 32B: allenai/dolma3_dolmino_mix-100B-1125 Dolma 3 Dolmino dataset pool for Olmo 3 stage 2 annealing training This dataset contains the high-quality pool of data considered for the second stage of Olmo 3 7B. Dataset Sources Source Category Tokens Documents TinyMATH Mind… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolma3_dolmino_pool.text-generation8 likes35k downloads9mo agoHugging Face07mlfoundations /dcvlm_pool_large DCVLM-Pool (large) The raw candidate pool at the large scale of our DataComp-VLM benchmark: 1,949,321,868 samples / 166.7 TB across 166 source datasets, as WebDataset tar shards — ≈4× the medium pool. 🚚 Upload in progress This repo is being populated incrementally and is not yet complete — shards are still being uploaded. Sources already present are final and safe to use; sources with fewer shards than the counts quoted below have not finished uploading yet.… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dcvlm_pool_large.image-text-to-text1B<n<10B0 likes30k downloads5d agoHugging Face08mlfoundations /dclm-pool-7b-1x1 likes24k downloads2y agoHugging Face09mlfoundations /dcvlm_pool_medium DCVLM-Pool (medium) The raw candidate pool at the medium scale of our DataComp-VLM benchmark: 483,576,747 samples / 41.1 TB across 166 source datasets, as WebDataset tar shards — ≈4× the small pool. This pool is unfiltered and unmixed. It is the input to a data-curation experiment, not a training set. You choose the filters and the mixing ratios, and create another training set. If you instead want a ready-to-train dataset, use dcvlm-baseline-200b (our reference SoTA… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dcvlm_pool_medium.image-text-to-text100M<n<1B0 likes20k downloads1mo agoHugging Face10allenai /dolma3_longmino_pool⚠️ IMPORTANT NOTICE ⚠️ This is the Dolma 3 Longmino pool; it hasn't been mixed. If you are interested in the data used to train: Olmo 3 7B: allenai/dolma3_longmino_mix-50B-1025 Olmo 3 32B: allenai/dolma3_dolmino_mix-100B-1125 Dolma 3 Longmino Pool (639B) Dolma 3 Longmino Pool is the full pool of documents considered for stage 3 (long context) extension trainin of Olmo 3 7B. Dataset Sources Source Type Tokens Docs LC-s2pdf-REX 32k-64k Synth PDFs 24.1B… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolma3_longmino_pool.14 likes15k downloads9mo agoHugging Face11mlfoundations /dclm-pool-1b-5x1 likes12k downloads2y agoHugging Face12mlfoundations /dcvlm_pool_small DCVLM-Pool (small) The raw candidate pool at the small scale of our DataComp-VLM benchmark: 120,940,134 samples / ~187.5B tokens / 10.3 TB across 166 source datasets, as WebDataset tar shards. This pool is unfiltered and unmixed. It is the input to a data-curation experiment, not a training set. You choose the filters and the mixing ratios, and create another training set. If you instead want a ready-to-train dataset, use dcvlm-baseline-200b (our reference SoTA DCVLM-baseline… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dcvlm_pool_small.image-text-to-text100M<n<1B0 likes9.1k downloads2mo agoHugging Face13KRAFTON /Raon-OpenTTS-Pool Raon-OpenTTS-Pool Technical Report Raon-OpenTTS-Pool is a large-scale open English speech corpus for text-to-speech (TTS) training, constructed from 8 publicly available speech corpora and a set of web-sourced recordings. It is the training data behind Raon-OpenTTS, an open TTS model that performs on par with state-of-the-art closed-data systems. 615K hours of speech audio 239.7M speech segments 11 source datasets aggregated into a unified format All… See the full description on the dataset page: https://huggingface.co/datasets/KRAFTON/Raon-OpenTTS-Pool.texttext-to-speech100M<n<1B42 likes9k downloads4mo agoHugging Face14Tensor-Link /cascade-eval-pool cascade eval pool — lagged public reveal (exact bytes) Retired snapshots of the held-out evaluation pool used by the cascade subnet. Each folder is a byte-identical mirror of the pool/snapshots/block-<N>.tar that validators scored — downloaded from the private pool bucket, sha256-verified against the publisher index, and republished unmodified. A snapshot is revealed only after a newer snapshot has superseded it, so no revealed pool can be selected by a current or future round.… See the full description on the dataset page: https://huggingface.co/datasets/Tensor-Link/cascade-eval-pool.1 likes8.9k downloads5h agoHugging Face15placeholderlabs /locus-commit-pool-v1 Locus Commit Pool v1 Native Git history, preserved as replayable software changes Commit message · complete selected before-state · unified patches · native object IDs · provenance · experimental labels Locus Commit Pool v1 is a large evidence pool for studying and training on how real software changes. Each document represents one surviving single-parent, multi-file Git commit. It keeps the commit message, the selected files as they existed before the… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/locus-commit-pool-v1.text-generation0 likes6.6k downloads1mo agoHugging Face16Bkmz1124 /igris-pool0 likes4.9k downloads19d agoHugging Face17ZhejiangLab /CPT_Data_Pool CPT Data Pool This repository hosts a large-scale CPT corpus designed for continual pre-training of domain-specific large language models. It serves as a data component of the 👉F2D-LLM framework, an end-to-end pipeline for domain-specific LLM training. For detailed data processing, scoring methods, and sampling strategies, please refer to the official 👉GitHub repo. Overall, the dataset contains approximately 271B tokens and is stored as jsonl files with a unified schema… See the full description on the dataset page: https://huggingface.co/datasets/ZhejiangLab/CPT_Data_Pool.text100M<n<1B1 likes4.7k downloads3mo agoHugging Face18mlfoundations /dclm-pool-400m-1x3 likes4.6k downloads2y agoHugging Face19HCAI-Lab-GT /archive-dolma3-pool-150b-enriched archive-dolma3-pool-150b-enriched ARCHIVE (pre-6T era): WebOrganizer-enriched variant of the 150B pool sample. Provenance This dataset was renamed on 2026-05-25 as part of the HCAI-Lab HF naming convention cleanup (PR 3). See docs/HCAI_LAB_NAMING_CONVENTION.md in the project repo for the convention. Field Value Previous name HCAI-Lab/dolma3_pool_150B_enriched Renamed 2026-05-25 See docs/data_home/inventory.json for the full inventory including… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/archive-dolma3-pool-150b-enriched.0 likes3.1k downloads4mo agoHugging Face20poolside-laguna-hackathon /patchrecoverygym-laguna PatchRecoveryGym for Laguna Submitted by: Kannappan Sirchabesan (@kannappans) · Poolside Research Hackathon (Foundations track) A reproducible eval + RL environment that tests whether a coding agent can recover from a wrong first attempt — a real, under-measured agentic-coding weakness. Built for Poolside Laguna XS.2 on dependency-migration repair tasks. 📦 Installable Verifiers environment on the Prime Hub · 🎯 deterministic hidden-test reward · 🔁 144-candidate reranking… See the full description on the dataset page: https://huggingface.co/datasets/poolside-laguna-hackathon/patchrecoverygym-laguna.text-generationn<1K0 likes2.8k downloads4mo agoHugging Face21HCAI-Lab-GT /archive-dolma3-pool-150b archive-dolma3-pool-150b ARCHIVE (pre-6T era): the 150B-token Dolma3 pool sample. Kept for reproducibility of earlier work; current attribution work uses the 6T datasets. Provenance This dataset was renamed on 2026-05-25 as part of the HCAI-Lab HF naming convention cleanup (PR 3). See docs/HCAI_LAB_NAMING_CONVENTION.md in the project repo for the convention. Field Value Previous name HCAI-Lab/dolma3_pool_150B Renamed 2026-05-25 See… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/archive-dolma3-pool-150b.1 likes2.6k downloads4mo agoHugging Face22wAI-org /tmax-image-pool tmax image pool Every Apptainer/SIF image referenced by the three swerl-tmax-15k variants, stored once and shared between them. Read the layout from the repo, not from a prefix The plain images are split across four directories, not one. A consumer that filters on images/ silently gets 4,519 fewer files and then reports itself complete — the tasks that reference the missing blobs fail later, at sandbox start, one by one. directory files images/ 9,971… See the full description on the dataset page: https://huggingface.co/datasets/wAI-org/tmax-image-pool.3 likes2.2k downloads7d agoHugging Face23dataforge-labs /ethereum-attestation-pool Ethereum attestation pool Slot-level observations comparing attestations visible in a consensus client's pending pool with attestations subsequently included in blocks. The panel supports analysis of local pool coverage and inclusion counts. Contents Table Record ethereum_attestation_observations Attester-slot counts observed pending and included, with their difference and collection coverage Using the data attesters_seen_in_pool… See the full description on the dataset page: https://huggingface.co/datasets/dataforge-labs/ethereum-attestation-pool.tabulartime-series-forecasting10K<n<100K0 likes1.7k downloads4h agoHugging Face24dataforge-labs /bitcoin-mining-pool-templates Bitcoin mining pool templates Timestamped Stratum job messages collected directly from Bitcoin mining pool endpoints. The data records changes in the work each endpoint sends to miners, including the previous block hash, coinbase data and clean-jobs flag. Contents Table Record bitcoin_mining_pool_jobs A job received from a pool endpoint, with its observation time, nTime, coinbase, merkle branch count and clean-jobs flag Using the data… See the full description on the dataset page: https://huggingface.co/datasets/dataforge-labs/bitcoin-mining-pool-templates.tabulartime-series-forecasting100K<n<1M0 likes1.6k downloads4h agoHugging Face25spade-rl /SPADE-Environment-Pool-GPT5.5-Games SPARE GPT-5.5 Grounded Cognitive Multi-Turn Games This public dataset contains 7,872 validated Python game environments for actor-only SPARE training. Six cognitive skills, exactly 1,312 environments per skill Generated with GPT-5.5 and grounded by spice_megascience_15k.jsonl Grounding corpus SHA-256: a36a928b4940b5b5d9e3f4cb5804a94c69462360943adb3be14613c82f0f72c0 Maximum 25 turns and 32K generation context Every environment passes load, reset, step, and replay validation with… See the full description on the dataset page: https://huggingface.co/datasets/spade-rl/SPADE-Environment-Pool-GPT5.5-Games.reinforcement-learning1K<n<10K1 likes1.5k downloads27d agoHugging Face26FineEnvs /watercolour-reference-pool Watercolour reference pool The reference paintings that define the reward in the watercolour RL environment: an agent writes a p5.brush sketch, the sketch is rendered, and a vision judge compares the render against paintings sampled from this pool. What the pool contains is the reward function. Replace it and you have changed what the environment rewards, without touching a line of code. 178 paintings in two tiers, each with the JavaScript source that produced it. tier… See the full description on the dataset page: https://huggingface.co/datasets/FineEnvs/watercolour-reference-pool.imageimage-classificationn<1K2 likes1.3k downloads22d agoHugging Face27harpreetsahota /InsPLAD-workshop-pool Dataset Card for InsPLAD Workshop Pool This is a FiftyOne dataset with 1754 samples. Installation If you haven't already, install FiftyOne: pip install -U fiftyone Usage import fiftyone as fo from fiftyone.utils.huggingface import load_from_hub # Load the dataset # Note: other available arguments include 'max_samples', etc dataset = load_from_hub("harpreetsahota/InsPLAD-workshop-pool") # Launch the App session = fo.launch_app(dataset)… See the full description on the dataset page: https://huggingface.co/datasets/harpreetsahota/InsPLAD-workshop-pool.imageobject-detection1K<n<10K1 likes1.3k downloads1mo agoHugging Face28mlfoundations /dcvlm_pool_small_annotations DCVLM-Pool (small) — per-sample annotations Every filtering annotation we computed for the small data pool of our DataComp-VLM paper: image quality, image–text alignment, language ID, text-quality classifiers, multimodal perplexity, decontamination scores and more — up to 167 fields per sample (180 distinct fields overall), for all 120,940,134 samples across 166 source datasets. These are the raw annotations, not a filtered dataset. They are the inputs our curation pipeline… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dcvlm_pool_small_annotations.image-text-to-text100M<n<1B0 likes1.2k downloads2mo agoHugging Face29placeholderlabs /locus-repo-code-pool-v2 Locus repository code pool The Locus repository code pool contains curated repository-version documents for code-language-model research. Each row combines the useful files from one Git repository snapshot into a single deterministic text document while retaining source provenance, file order, language and test signals, health evidence, and stable content hashes. This release is a candidate acquisition pool, not a ready-made training split. Use a globally deduplicated manifest… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/locus-repo-code-pool-v2.text-generation1M<n<10M0 likes1k downloads1mo agoHugging Face30thaottn /datacomp-medium-pool-translatedimage100M<n<1B0 likes1k downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.