CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01bigscience /P3 Dataset Card for P3 Dataset Summary P3 (Public Pool of Prompts) is a collection of prompted English datasets covering a diverse set of NLP tasks. A prompt is the combination of an input template and a target template. The templates are functions mapping a data example into natural language for the input and target sequences. For example, in the case of an NLI dataset, the data example would include fields for Premise, Hypothesis, Label. An input template would be If… See the full description on the dataset page: https://huggingface.co/datasets/bigscience/P3.textother100M<n<1B235 likes241k downloads3y agoHugging Face02HuggingFaceCode /stack-v3-train 🥞 The Stack v3 What is it? What is being released How to download and use it Dataset statistics Dataset structure Dataset creation Considerations for using the data Additional information What is it? The Stack v3 is the largest, most up-to-date open dataset of source code, crawled directly from GitHub and built to pre-train code LLMs with full-repository context. It is the successor to The Stack v2 and, like its predecessor, is released to make the training… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceCode/stack-v3-train.tabulartext-generation100M<n<1B382 likes183k downloads1d agoHugging Face03cadene /agibot_alpha_v30This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "AgiBot_A2D", "total_episodes": 28122, "total_frames": 47613574, "total_tasks": 30, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 500, "fps": 30, "splits": { "train": "0:28122" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/cadene/agibot_alpha_v30.tabularrobotics10M<n<100M2 likes69k downloads1y agoHugging Face04allenai /tulu-3-sft-mixture Tulu 3 SFT Mixture Note that this collection is licensed under ODC-BY-1.0 license; different licenses apply to subsets of the data. Some portions of the dataset are non-commercial. We present the mixture as a research artifact. The Tulu 3 SFT mixture was used to train the Tulu 3 series of models. It contains 939,344 samples from the following sets: CoCoNot (ODC-BY-1.0), 10,983 prompts (Brahman et al., 2024) FLAN v2 via ai2-adapt-dev/flan_v2_converted, 89,982 prompts (Longpre et… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-3-sft-mixture.textother100K<n<1M265 likes61k downloads2y agoHugging Face05nc33 /multispan_quoreftext10K<n<100K0 likes59k downloads4y agoHugging Face06YipengGao /3DCode Project page Paper Code 3dcodebench.com arXiv:2606.01057 gaoypeng/3dcodebench News [06/01/2026] Paper released on arXiv: 3DCodeBench: Benchmarking Agentic Procedural 3D Modeling Via Code. Note. This is an open-source reproduction of 3DCodeBench. ⚠️ Under final check. The 3DCodeData/ code is still undergoing final quality review and may contain occasional issues (non-executable scripts, mismatched captions/renders, or imperfect geometry). If you run… See the full description on the dataset page: https://huggingface.co/datasets/YipengGao/3DCode.3dtext-to-3d10K<n<100K25 likes48k downloads6d agoHugging Face07openbmb /Ultra-FineWeb-L3 Ultra-FineWeb-L3 📜 Ultra-FineWeb Technical Report | 📦 UltraData Collection | 🌐 UltraData | 🤗 MiniCPM5 Series English | 中文 📚 Introduction Ultra-FineWeb-L3 is the L3 refined data for general high-quality web data within UltraData's L0-L4 tiered data management framework. Moving beyond L2 quality selection, it transforms high-value web corpora into structured, high-learnability training data with clearer reasoning signals and richer educational… See the full description on the dataset page: https://huggingface.co/datasets/openbmb/Ultra-FineWeb-L3.texttext-generation1B<n<10B338 likes27k downloads1mo agoHugging Face08laion /BVD-I-300M-URLs LAION-BVD - 300M Video Frame URLs This repository contains the URLs for ~300 million keyframes extracted from publicly available web videos. No image data is included, only the source video URL and the frame timestamp needed to reproduce each frame. Frames were extracted from BVD-RAW and cover YouTube, Dailymotion, and Vimeo content. Dataset structure Column Type Description webpage_url string URL of the source video frame_pts_time float Presentation… See the full description on the dataset page: https://huggingface.co/datasets/laion/BVD-I-300M-URLs.textimage-text-to-text100M<n<1B2 likes25k downloads1mo agoHugging Face09IFM /TxT360-v2 TxT360-v2 Dataset Description Pre-training sources for the K2 Horizon training data release. This repository is part of the K2 Horizon collection. The repository is organized into multiple subsets. Every subset has a train split backed by Parquet shards. K2 Horizon Dataset Series Dataset repository Focus Subsets IFM/TxT360-v2 Web and question-answering text 3 IFM/Code-Reasoning Code reasoning and task synthesis 7 IFM/Math-Reasoning… See the full description on the dataset page: https://huggingface.co/datasets/IFM/TxT360-v2.tabulartext-generation1B<n<10B81 likes24k downloads4d agoHugging Face10open-thoughts /OpenThoughts3-1.2M paper | dataset | model [!NOTE] We have released a paper for OpenThoughts! See our paper here. OpenThoughts3-1.2M Open-source state-of-the-art reasoning dataset with 1.2M rows. 🚀 OpenThoughts3-1.2M is the third iteration in our line of OpenThoughts datasets, building on our previous OpenThoughts-114k and OpenThoughts2-1M. This time around, we scale even further and generate our dataset in a much more systematic way -- OpenThoughts3-1.2M is the result of a… See the full description on the dataset page: https://huggingface.co/datasets/open-thoughts/OpenThoughts3-1.2M.texttext-generation1M<n<10M264 likes24k downloads1y agoHugging Face11arsaporta /symile-m3 Dataset Card for Symile-M3 Symile-M3 is a multilingual dataset of (audio, image, text) samples. The dataset is specifically designed to test a model's ability to capture higher-order information between three distinct high-dimensional data types: by incorporating multiple languages, we construct a task where text and audio are both needed to predict the image, and where, importantly, neither text nor audio alone would suffice. Paper: https://arxiv.org/abs/2411.01053 GitHub:… See the full description on the dataset page: https://huggingface.co/datasets/arsaporta/symile-m3.audiozero-shot-classification10M<n<100M8 likes24k downloads2y agoHugging Face12allenai /tulu-3-sft-personas-instruction-following Dataset Descriptions This dataset contains 29980 examples and is synthetically created to enhance model's capabilities to follow instructions precisely and to satisfy user constraints. The constraints are borrowed from the taxonomy in IFEval dataset. To generate diverse instructions, we expand the methodology in Ge et al., 2024 by using personas. More details and exact prompts used to construct the dataset can be found in our paper. Curated by: Allen Institute for AI Paper: TBD… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-3-sft-personas-instruction-following.texttext-generation10K<n<100K68 likes16k downloads2y agoHugging Face13tanganke /sun397 SUN397 dataset The database contains 397 categories subset from the SUN dataset for Scene Recognition used in the following paper. The number of images varies across categories, but there are at least 100 images per category, and 108,754 images in total. All images are in jpg format. The images provided here are for research purposes only. The file ClassName.txt contains the name list for the 397 categories. Please cite the following paper if you use this dataset in your research.… See the full description on the dataset page: https://huggingface.co/datasets/tanganke/sun397.imageimage-classification10K<n<100K4 likes15k downloads2y agoHugging Face14simon3000 /genshin-voice Genshin Voice Genshin Voice is a dataset of voice lines from the popular game Genshin Impact. Hugging Face 🤗 Genshin-Voice ModelScope Genshin-Voice Per-speaker downloads are grouped by language and ZIP size. Browse every archive in the ZIP index. Last update at 2026-08-13 654252 wavs 7291 without speaker (1%) 52693 without transcription (8%) 1088 without inGameFilename (0%) Dataset Details Dataset Description The dataset contains voice lines… See the full description on the dataset page: https://huggingface.co/datasets/simon3000/genshin-voice.audioaudio-classification100K<n<1M271 likes15k downloads26d agoHugging Face15siyanzhao /Openthoughts_math_30k_opsdtext10K<n<100K10 likes13k downloads7mo agoHugging Face16ember-lab-berkeley /robocasa365-pretrain-mg Pretraining (MimicGen) — atomic MimicGen-generated rollouts across 60 atomic tasks (~10,000 demos/task). 1,615 hours total, generated by scripted augmentation from human demonstrations. Part of the RoboCasa365 collection. Flat LeRobot v3.0 mirror of RoboCasa365 — standard layout, drop-in loadable. Stats Episodes: 536,030 Frames: 116,246,439 (20 fps → 1615 h) Tasks: 720 (natural-language phrasings; underlying RoboCasa task classes: 60) Cameras: 3 × 256×256 h264… See the full description on the dataset page: https://huggingface.co/datasets/ember-lab-berkeley/robocasa365-pretrain-mg.tabularrobotics100M<n<1B2 likes11k downloads5mo agoHugging Face17oneHFR /3d-front-arimage10K<n<100K0 likes11k downloads1y agoHugging Face18lvogel123 /jailbreak-deepseek-v3.2-exptabular1K<n<10K1 likes10k downloads11mo agoHugging Face19stzhao /AnyWord-3MDataset from AnyText: Multilingual Visual Text Generation And Editing. Dataset description from Anytext Team: Currently, there is a relative scarcity of public datasets for text generation tasks, especially those involving non-Latin script languages. To address this, we introduce a large-scale multilingual dataset called AnyWord-3M. The images in this dataset are sourced from Noah-Wukong, LAION-400M, and OCR recognition datasets such as ArT, COCO-Text, RCTW, LSVT, MLT, MTWI, ReCTS, etc. These… See the full description on the dataset page: https://huggingface.co/datasets/stzhao/AnyWord-3M.imagetext-to-image1M<n<10M17 likes10k downloads2y agoHugging Face20CoRal-project /coral-v3gated CoRal: Danish Conversational and Read-aloud Dataset Version 3.0 Dataset Overview CoRal is a comprehensive Automatic Speech Recognition (ASR) dataset designed to capture the diversity of the Danish language across various dialects, accents, genders, and age groups. The primary goal of the CoRal dataset is to provide a robust resource for training and evaluating ASR models that can understand and transcribe spoken Danish in all its variations. Key Features… See the full description on the dataset page: https://huggingface.co/datasets/CoRal-project/coral-v3.audioautomatic-speech-recognition100K<n<1M6 likes9.9k downloads7mo agoHugging Face21gasstation /gs-images-v3tabular100K<n<1M0 likes9.3k downloads5mo agoHugging Face22Ta1k1 /HLE_200-3textn<1K0 likes9k downloads1y agoHugging Face23D4nt3 /esb-datasets-earnings22-validation-tiny-filteredA filtered (<=30s duration) slice (512 samples) of the Earnings22 dataset. def add_duration(sample): y, sr = sample['audio']["array"], sample['audio']["sampling_rate"] sample['duration_ms']=librosa.get_duration(y=y, sr=sr) * 1000 return sample tedlium = load_dataset("esb/datasets", "earnings22", split='validation', trust_remote_code=True) # compute duration to filter tedlium = tedlium.map(add_duration) tedlium = tedlium.select(range(512)) # Whisper max supported duration tedlium… See the full description on the dataset page: https://huggingface.co/datasets/D4nt3/esb-datasets-earnings22-validation-tiny-filtered.audion<1K0 likes8.7k downloads2y agoHugging Face24AnchorSR /TrainingData_Stage3 AnchorSR Stage3 · metric-v1.0 直接选择 Small / Large 配置 训练题数 用途 small 1,000,000 先验证答案监督/先验恢复,按新版 Large 联合分布抽样 large 89,801,853 筛选后的完整训练集合,包含 Small 全部样本 from datasets import load_dataset data = load_dataset('AnchorSR/TrainingData_Stage3', 'small', # 或 large revision='metric-v1.0', streaming=True) 这是对 scaling-v1.0 的语义筛选与统一任务分类,不是增加新数据源。 Large 从 89,828,269 题保留 89,801,853 题,隔离 26,416 题。 旧标签 scaling-v1.0 / video-v1.0 / large-v1.0… See the full description on the dataset page: https://huggingface.co/datasets/AnchorSR/TrainingData_Stage3.tabularvisual-question-answering100M<n<1B0 likes8.5k downloads5d agoHugging Face25VidaForge /VidaForge-3M 3.14 million scene-level video clips with multi-level captions, camera labels, semantic tags, quality signals, and duplicate groups. Paper · VidaForge Code · Project Blog · Source Dataset Overview VidaForge-3M is a large-scale video pretraining dataset produced with VidaForge, an open data pipeline for building and studying video foundation model pretraining data. The pipeline and dataset are described in the paper VidaForge: Open Research Infrastructure… See the full description on the dataset page: https://huggingface.co/datasets/VidaForge/VidaForge-3M.tabulartext-to-videon<1K13 likes7.7k downloads13d agoHugging Face26amithm3 /shrutilipiaudioautomatic-speech-recognition1M<n<10M6 likes7k downloads2y agoHugging Face27Salesforce /blip3-kale 🥬 BLIP3-KALE:Knowledge Augmented Large-scale Dense Captions BLIP3-KALE is an open-source dataset of 218 million image-text pairs, featuring knowledge-augmented dense captions combining web-scale knowledge with detailed image descriptions. Paper: [To be added] Uses BLIP3-KALE is designed to facilitate research in multimodal pretraining. The dataset can be used for training large multimodal models that require factually grounded, dense image captions. It has already been an… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/blip3-kale.imageimage-to-text100M<n<1B47 likes6.8k downloads2y agoHugging Face28scaleinvariant /paired-llama-3.2-1b-embeddings-lmsys-chat-1m Paired Llama 3.2 1B Token Embeddings (LMSYS-Chat-1M) This dataset contains paired activations corresponding to single token locations extracted from Meta's Llama 3.2 1B Instruct on conversations from LMSYS-Chat-1M. Embeddings are provided for layers 5 through 14, which capture the most interesting intermediate representations. This dataset was built to study things like: Learning different basis for activations at a given layer Studying if there are cases where position encodes… See the full description on the dataset page: https://huggingface.co/datasets/scaleinvariant/paired-llama-3.2-1b-embeddings-lmsys-chat-1m.tabularfeature-extraction100M<n<1B3 likes6.6k downloads7mo agoHugging Face29cc-clean /CC-MAIN-2025-33text100M<n<1B0 likes6.6k downloads1y agoHugging Face30simon3000 /zenless-voice Zenless Voice Zenless Voice is a dataset of voice lines from the popular game Zenless Zone Zero. Hugging Face 🤗 Zenless-Voice ModelScope Zenless-Voice Per-speaker downloads are grouped by language and WAV count. Browse every archive in the ZIP index. Last update at 2026-09-17, game version 3.2.0 406720 wavs 78785 without speaker (19%) 123429 without transcription (30%) 83509 without inGameFilename (21%) Speaker archives contain 327,935 WAVs in 4,322 ZIPs. The 78,785 rows… See the full description on the dataset page: https://huggingface.co/datasets/simon3000/zenless-voice.audioaudio-classification100K<n<1M5 likes6.4k downloads9d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.