CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01bigscience /P3 Dataset Card for P3 Dataset Summary P3 (Public Pool of Prompts) is a collection of prompted English datasets covering a diverse set of NLP tasks. A prompt is the combination of an input template and a target template. The templates are functions mapping a data example into natural language for the input and target sequences. For example, in the case of an NLI dataset, the data example would include fields for Premise, Hypothesis, Label. An input template would be If… See the full description on the dataset page: https://huggingface.co/datasets/bigscience/P3.textother100M<n<1B235 likes241k downloads3y agoHugging Face02longisland3 /ptb-xl1 likes232k downloads2y agoHugging Face03atokforps /latent_worker_early3_20 likes190k downloads4y agoHugging Face04k9cli /video-vec2wav2-tokenizer-3 video-vec2wav2-tokenizer-3 Version 3 - continuation shard of the video-to-AI-dataset tokenizer project. Version 2 - continuation shard of the video-to-AI-dataset tokenizer project. Production-ready pipeline (Python package video_vec2wav2_tokenizer, CLI command video2dataset) that turns a folder of videos into clean AI training datasets for speech recognition (ASR) and text-to-speech (TTS). videos ──► audio (16 kHz mono PCM) ──► whisper transcript ──► clips ──►… See the full description on the dataset page: https://huggingface.co/datasets/k9cli/video-vec2wav2-tokenizer-3.1 likes186k downloads2mo agoHugging Face05HuggingFaceCode /stack-v3-train 🥞 The Stack v3 What is it? What is being released How to download and use it Dataset statistics Dataset structure Dataset creation Considerations for using the data Additional information What is it? The Stack v3 is the largest, most up-to-date open dataset of source code, crawled directly from GitHub and built to pre-train code LLMs with full-repository context. It is the successor to The Stack v2 and, like its predecessor, is released to make the training… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceCode/stack-v3-train.tabulartext-generation100M<n<1B382 likes184k downloads16h agoHugging Face06harborframework /terminal-bench-3.0 Terminal-Bench 3.0 The primary source is hosted on GitHub, please open issues and pull requests there, not here. The official published dataset is hosted on the Harbor Hub along with the official leaderboard. Usage e.g. harbor run -d terminal-bench/terminal-bench@3.0.0 This repo is a mirror of harbor-framework/terminal-bench at tag v3.0.0, laid out so it can be consumed directly by Harbor's git-repos dataset support. How to run via this Huggingface repo Always… See the full description on the dataset page: https://huggingface.co/datasets/harborframework/terminal-bench-3.0.5 likes171k downloads1mo agoHugging Face07GSMA /3GPP 3GPP specification mirror Part of the Open-Telco Telecom Standards Corpus.Sibling mirrors: 3GPP · ETSI · ITU-T · O-RAN · GSMA · TM Forum · CAMARA This dataset mirrors 3GPP's published specifications. Every source is kept twice: original/ is the document as 3GPP released it, and marked/ is that same document converted to Markdown for search and retrieval, one raw.md per document with any figures extracted beside it. The two trees share identical paths, so a file in original/… See the full description on the dataset page: https://huggingface.co/datasets/GSMA/3GPP.10K<n<100K5 likes104k downloads3mo agoHugging Face08YipengGao /3DCode Project page Paper Code 3dcodebench.com arXiv:2606.01057 gaoypeng/3dcodebench News [06/01/2026] Paper released on arXiv: 3DCodeBench: Benchmarking Agentic Procedural 3D Modeling Via Code. Note. This is an open-source reproduction of 3DCodeBench. ⚠️ Under final check. The 3DCodeData/ code is still undergoing final quality review and may contain occasional issues (non-executable scripts, mismatched captions/renders, or imperfect geometry). If you run… See the full description on the dataset page: https://huggingface.co/datasets/YipengGao/3DCode.3dtext-to-3d10K<n<100K25 likes83k downloads6d agoHugging Face09atokforps /latent_v1_fullrun_alpha3_040 likes82k downloads4y agoHugging Face10bigscience /xP3xP3 (Crosslingual Public Pool of Prompts) is a collection of prompts & datasets across 46 of languages & 16 NLP tasks. It is used for the training of BLOOMZ and mT0, multilingual language models capable of following human instructions in dozens of languages zero-shot.other100M<n<1B113 likes74k downloads3y agoHugging Face11allenai /dolma3_mix-6T Dolma 3 Mix (6T) The Dolma 3 Mix (6T) is the collection of data used during the pretraining stage to train the Olmo-3-1125-32B model. This dataset is made up of ~6 trillion tokens from a diverse mix of web content, academic publications, code, and more. The majority of this dataset comes from Common Crawl. For more information on Dolma, please see our original release here. Smaller Sample for Analysis Available! If you would like a smaller sample of this mix… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolma3_mix-6T.text-generation36 likes72k downloads8mo agoHugging Face12cadene /agibot_alpha_v30This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "AgiBot_A2D", "total_episodes": 28122, "total_frames": 47613574, "total_tasks": 30, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 500, "fps": 30, "splits": { "train": "0:28122" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/cadene/agibot_alpha_v30.tabularrobotics10M<n<100M2 likes69k downloads1y agoHugging Face13allenai /dolma3.5_pool⚠️ IMPORTANT NOTICE ⚠️ This is the Dolma 3.5 pool. It contains no quality upsampling or mixing. This is an updated version of the Dolma 3 pool with additional quality filtering and more data sources. If you are interested in the data used to train Olmo 3 7B and Olmo 3 32B, visit allenai/dolma3_mix-6T-1025. Dolma 3.5 Pool The Dolma 3.5 pool is a dataset of nearly 10 trillion tokens from a diverse mix of web content, academic publications, code, and more. For detailed… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolma3.5_pool.text-generation8 likes69k downloads3mo agoHugging Face14JIAQI-CHEN /3D-Front3 likes65k downloads2y agoHugging Face15allenai /tulu-3-sft-mixture Tulu 3 SFT Mixture Note that this collection is licensed under ODC-BY-1.0 license; different licenses apply to subsets of the data. Some portions of the dataset are non-commercial. We present the mixture as a research artifact. The Tulu 3 SFT mixture was used to train the Tulu 3 series of models. It contains 939,344 samples from the following sets: CoCoNot (ODC-BY-1.0), 10,983 prompts (Brahman et al., 2024) FLAN v2 via ai2-adapt-dev/flan_v2_converted, 89,982 prompts (Longpre et… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-3-sft-mixture.textother100K<n<1M265 likes61k downloads2y agoHugging Face16nc33 /multispan_quoreftext10K<n<100K0 likes59k downloads4y agoHugging Face17atom-in-the-universe /bild-deduped-30 likes52k downloads3y agoHugging Face18nkp37 /OpenVid-1M Summary This is the dataset proposed in our paper [ICLR 2025] OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation. OpenVid-1M is a high-quality text-to-video dataset designed for research institutions to enhance video quality, featuring high aesthetics, clarity, and resolution. It can be used for direct training or as a quality tuning complement to other video datasets. All videos in the OpenVid-1M dataset have resolutions of at least 512×512.… See the full description on the dataset page: https://huggingface.co/datasets/nkp37/OpenVid-1M.videotext-to-video1M<n<10M284 likes50k downloads6mo agoHugging Face19aliasfox /srtm30m-ozt2-v2 SRTM 30m OZT2 Elevation Tiles This dataset contains SRTM 30-meter resolution elevation data encoded in the OZT2 tile format. Format OZT2 is a high-performance elevation tile format: Compression: ~93% smaller than Terrarium PNG Prediction: Gradient-based prediction (left neighbor + vertical gradient) Quantization: Adaptive bit-depth (8/10/12/16-bit per channel) Codec: Zstd q3 (30× faster encode than Brotli, same decode speed) Each tile is 256×256 pixels in Web… See the full description on the dataset page: https://huggingface.co/datasets/aliasfox/srtm30m-ozt2-v2.n<1K0 likes46k downloads2d agoHugging Face20kelvin34501 /OakInk-v2 Dataset Card for OakInk-v2 Project: https://oakink.net/v2 Paper: https://arxiv.org/pdf/2403.19417.pdf image-to-3d1M<n<10M10 likes44k downloads1y agoHugging Face21brownu /deform360 Deform360: A Massive Multi-view Visuotactile Dataset for Deformable World Models Project Page | Paper | GitHub Repository Deform360 is a massive multi-view visuotactile dataset for deformable-object research, featuring 198 daily-life objects, 1,980 interaction sequences, and over 215 hours of observations from 41 surround-view cameras and bimanual tactile grippers to capture both global motion and contact-induced local deformations. Where the data lives… See the full description on the dataset page: https://huggingface.co/datasets/brownu/deform360.textrobotics13 likes44k downloads1mo agoHugging Face22nganvo31032 /nganvo310328 likes41k downloads21d agoHugging Face23atokforps /latent_v1_fullrun_alpha3_060 likes38k downloads4y agoHugging Face24allenai /dolma3_dolmino_pool⚠️ IMPORTANT NOTICE ⚠️ This is the Dolma 3 Dolmino pool; it hasn't been mixed. If you are interested in the data used to train: Olmo 3 7B: allenai/dolma3_dolmino_mix-100B-1025 Olmo 3 32B: allenai/dolma3_dolmino_mix-100B-1125 Dolma 3 Dolmino dataset pool for Olmo 3 stage 2 annealing training This dataset contains the high-quality pool of data considered for the second stage of Olmo 3 7B. Dataset Sources Source Category Tokens Documents TinyMATH Mind… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolma3_dolmino_pool.text-generation8 likes38k downloads9mo agoHugging Face25atokforps /latent_v1_fullrun_alpha3_130 likes37k downloads4y agoHugging Face26atokforps /latent_v1_fullrun_alpha3_010 likes33k downloads4y agoHugging Face27LindseyLarson3372 /imagesimage10K<n<100K6 likes33k downloads2mo agoHugging Face28atokforps /latent_v1_fullrun_alpha3_050 likes30k downloads4y agoHugging Face29atokforps /latent_v1_fullrun_alpha3_030 likes29k downloads4y agoHugging Face30allenai /dolma3_pool⚠️ IMPORTANT NOTICE ⚠️ This is the Dolma 3 pool, pre–quality upsampling and mixing. If you are interested in the data used to train Olmo 3 7B and Olmo 3 32B, visit allenai/dolma3_mix-6T-1025. Dolma 3 Pool The Dolma 3 pool is a dataset of over 9 trillion tokens from a diverse mix of web content, academic publications, code, and more. For detailed documenation on Dolma 3 processing and data, please see our Dolma 3 Github repository. For more information on Dolma in general… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolma3_pool.texttext-generation10B<n<100B41 likes29k downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.