CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mlfoundations /dclm-baseline-1.0 DCLM-baseline DCLM-baseline is a 4T token / 3B document pretraining dataset that achieves strong performance on language model benchmarks. Below are comparisions of model trained on DCLM-baseline with other models in the 7B regime. Model Params Tokens Open dataset? CORE MMLU EXTENDED Open weights, closed datasets Llama2 7B 2T ✗ 49.2 45.8 34.1 DeepSeek 7B 2T ✗ 50.7 48.5 35.3 Mistral-0.3 7B ? ✗ 57.0 62.7 45.1 QWEN-2 7B ? ✗ 57.5 71.9 50.5 Llama3 8B 15T ✗… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dclm-baseline-1.0.tabular1B<n<10B315 likes648k downloads2y agoHugging Face02cadene /droid_1.0.1This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "Franka", "total_episodes": 95600, "total_frames": 27612581, "total_tasks": 0, "total_videos": 286800, "total_chunks": 95, "chunks_size": 1000, "fps": 15, "splits": { "train": "0:95600" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/cadene/droid_1.0.1.robotics57 likes282k downloads2y agoHugging Face03EssentialAI /essential-web-v1.0 🌐 Essential-Web: Complete 24-Trillion Token Dataset 🏆 Website | 🖥️ Code | 📖 Paper | ☁️ AWS 📋 Dataset Description Essential-Web is a 24-trillion-token web dataset with document-level metadata designed for flexible dataset curation. The dataset provides metadata including subject matter classification, web page type, content complexity, and document quality scores for each of the 23.6 billion documents. Researchers can filter and curate specialized datasets… See the full description on the dataset page: https://huggingface.co/datasets/EssentialAI/essential-web-v1.0.10B<n<100B247 likes80k downloads1y agoHugging Face04mlfoundations /dclm-baseline-1.0-parquet DCLM-baseline Note: this is an identical copy of https://huggingface.co/datasets/mlfoundations/dclm-baseline-1.0, where all the files have been mapped to a parquet format. DCLM-baseline is a 4T token / 3B document pretraining dataset that achieves strong performance on language model benchmarks. Below are comparisions of model trained on DCLM-baseline with other models in the 7B regime. Model Params Tokens Open dataset? CORE MMLU EXTENDED Open weights, closed datasets… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dclm-baseline-1.0-parquet.tabular1B<n<10B56 likes19k downloads2y agoHugging Face05lerobot /droid_1.0.1This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "Franka", "total_episodes": 95658, "total_frames": 27630375, "total_tasks": 49630, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 500, "fps": 15, "splits": { "train": "0:95658" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/droid_1.0.1.tabularrobotics10M<n<100M24 likes16k downloads3mo agoHugging Face06parrotzone /sdxl-1.0 check sdxl.parrotzone.art for easy viewing ⋆。°✩ all images were made with SDXL 1.0 + the 0.9 VAE steps: 20 cfg scale: 7 no refiner random seeds image1K<n<10K14 likes16k downloads3y agoHugging Face07argilla /magpie-ultra-v1.0 Dataset Card for magpie-ultra-v1.0 This dataset has been created with distilabel. Dataset Summary magpie-ultra it's a synthetically generated dataset for supervised fine-tuning using the Llama 3.1 405B-Instruct model, together with other Llama models like Llama-Guard-3-8B and Llama-3.1-8B-Instruct. The dataset contains challenging instructions and responses for a wide variety of tasks, such as Coding & debugging, Math, Data analysis, Creative Writing… See the full description on the dataset page: https://huggingface.co/datasets/argilla/magpie-ultra-v1.0.tabular1M<n<10M52 likes8.7k downloads2y agoHugging Face08mulan-dataset /v1.0 MuLAn: : A Multi Layer Annotated Dataset for Controllable Text-to-Image Generation MuLAn is a novel dataset comprising over 44K MUlti-Layer ANnotations of RGB images as multilayer, instance-wise RGBA decompositions, and over 100K instance images. It is composed of MuLAn-COCO and MuLAn-LAION sub-datasets, which contain a variety of image decompositions in terms of style, composition and complexity. With MuLAn, we provide the first photorealistic resource providing instance… See the full description on the dataset page: https://huggingface.co/datasets/mulan-dataset/v1.0.imagetext-to-imagen<1K27 likes8.4k downloads2y agoHugging Face09LejuRobotics /LET-KUAVO-VLA-1.0-Datasetgated LET-KUAVO-VLA-1.0-Dataset videon<1K3 likes6.9k downloads17d agoHugging Face10IPEC-COMMUNITY /libero_spatial_no_noops_1.0.0_lerobottabular10K<n<100K5 likes6.7k downloads11mo agoHugging Face11IPEC-COMMUNITY /libero_object_no_noops_1.0.0_lerobottabular10K<n<100K1 likes6.2k downloads11mo agoHugging Face12IPEC-COMMUNITY /libero_10_no_noops_1.0.0_lerobottabular100K<n<1M3 likes6k downloads11mo agoHugging Face13IPEC-COMMUNITY /libero_goal_no_noops_1.0.0_lerobottabular10K<n<100K1 likes5.9k downloads11mo agoHugging Face14nvidia /Nemotron-AIQ-Agentic-Safety-Dataset-1.0 Nemotron-AIQ Agentic Safety Dataset Dataset Summary Nemotron-AIQ-Agentic-Safety-Dataset is a comprehensive dataset that captures a broad range of novel safety and security contextual risks that can emerge within agentic systems. It highlights the robustness of NVIDIA's open model, llama-3.3-nemotron-super-49b-v1, when deployed as a research assistant inside AIQ, demonstrating its ability to handle a diverse spectrum of agentic safety and security challenges. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-AIQ-Agentic-Safety-Dataset-1.0.texttext-generation10K<n<100K18 likes4.5k downloads10mo agoHugging Face15RealCADBench /RealCADBench-V1.0 RealCADBench-V1.0 RealCADBench-V1.0 contains design-intent inputs and ground-truth STL shapes for evaluating AI CAD models and agents on parts and assemblies. The test split is an evaluation collection, not a newly created train/test partition. The default all configuration combines every sample. The six other configurations provide the individual subsets without duplicating the Parquet files. Category Subset Samples part text 568 part 2d_drawing 236 part real_pic… See the full description on the dataset page: https://huggingface.co/datasets/RealCADBench/RealCADBench-V1.0.3d10K<n<100K16 likes4k downloads11d agoHugging Face16nvidia /Aegis-AI-Content-Safety-Dataset-1.0 🛡️ Nemotron Content Safety Dataset V1 Nemotron Content Safety Dataset V1, formerly known as Aegis AI Content Safety Dataset, is an open-source content safety dataset (CC-BY-4.0), which adheres to Nvidia's content safety taxonomy, covering 13 critical risk categories (see Dataset Description). Dataset Details Dataset Description Nemotron Content Safety Dataset V1 is comprised of approximately 11,000 manually annotated interactions between humans and LLMs, split… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Aegis-AI-Content-Safety-Dataset-1.0.texttext-classification10K<n<100K61 likes3.7k downloads1y agoHugging Face17occiglot /occiglot-fineweb-v1.0 Occiglot Fineweb v1.0 We present a more mature version of the multilingual Occiglot Fineweb corpus. In this early form, the dataset contains roughly 430M heavily cleaned documents from 10 languages. Occiglot Fineweb builds on our existing collection of curated datasets and pre-filtered web data. Subsequently, all documents were filtered with language-specific derivatives of the fine-web processing pipeline and different levels of depuplicated. We provide the data at 3 levels of… See the full description on the dataset page: https://huggingface.co/datasets/occiglot/occiglot-fineweb-v1.0.text-generation10B<n<100B3 likes3.5k downloads2y agoHugging Face18Last-Bullet /DOTAv1.0image1K<n<10K0 likes3.4k downloads2y agoHugging Face19VLABench /vlm_evaluation_v1.0 Datacard This dataset is the evaluation VLM dataset used in VLABench. It is designed to evaluate the planning capabilities of Vision-Language Models (VLMs) in embodied scenarios. Source Project Page: https://vlabench.github.io/ Arxiv Paper: https://arxiv.org/abs/2412.18194 Code: https://github.com/OpenMOSS/VLABench Uses The dataset structure is as follows: vlm_evaluation_v1.0/ ├── CommenSence/ ├── add_condiment_common_sense/ ├──… See the full description on the dataset page: https://huggingface.co/datasets/VLABench/vlm_evaluation_v1.0.image1K<n<10K0 likes3.2k downloads1y agoHugging Face20MLCommons /peoples_speech_v1.0 Dataset Card for People's Speech Dataset Summary The People's Speech Dataset is among the world's largest English speech recognition corpus today that is licensed for academic and commercial usage under CC-BY-SA and CC-BY 4.0. It includes 30,000+ hours of transcribed speech in English languages with a diverse set of speakers. This open dataset is large enough to train speech-to-text systems and crucially is available with a permissive license. Supported Tasks… See the full description on the dataset page: https://huggingface.co/datasets/MLCommons/peoples_speech_v1.0.automatic-speech-recognition8 likes2.7k downloads2y agoHugging Face21cadene /droid_1.0.1_v30This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "Franka", "total_episodes": 95584, "total_frames": 27607757, "total_tasks": 49596, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 500, "fps": 15, "splits": { "train": "0:95584" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/cadene/droid_1.0.1_v30.tabularrobotics10M<n<100M4 likes2.4k downloads1y agoHugging Face22villekuosmanen /dAgger_build_block_tower_1.0.0-advantages Advantage Values for villekuosmanen/dAgger_build_block_tower_1.0.0 Pre-computed advantage values for offline RL training. Source Dataset: villekuosmanen/dAgger_build_block_tower_1.0.0 Value Model: villekuosmanen/rewact_build_block_tower_all_3 N-step lookahead: 50 Files This dataset contains per-episode parquet files with advantage values for each frame. Usage from pathlib import Path import pandas as pd # Load advantages for a specific episode… See the full description on the dataset page: https://huggingface.co/datasets/villekuosmanen/dAgger_build_block_tower_1.0.0-advantages.tabularrobotics10K<n<100K0 likes2.4k downloads6mo agoHugging Face23aractingi /droid_1.0.1_testThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "Franka", "total_episodes": 95658, "total_frames": 27630375, "total_tasks": 49630, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 15, "splits": { "train": "0:95658" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/aractingi/droid_1.0.1_test.tabularrobotics10M<n<100M0 likes2.2k downloads10mo agoHugging Face24ekacare /spandan-1M-V1.0-raw Spandan A Large Photoplethysmography (PPG) Signal Dataset of 1 Million+ Indian Subjects In Sanskrit, "Spandan" (स्पन्दन - spandana) represents one of the most fundamental aspects of existence - the rhythmic pulsation that permeates all life. Derived from the root verb "spand" (स्पन्द), meaning "to throb" or "to pulsate," it beautifully captures the essence of the heartbeat. Dataset Overview Spandan is an extensive repository containing over 1 million… See the full description on the dataset page: https://huggingface.co/datasets/ekacare/spandan-1M-V1.0-raw.text100K<n<1M3 likes2.1k downloads2y agoHugging Face25birdsql /bird-critic-1.0-sqlite 📢 Update 2026-03-23 We release BIRD-Critic-SQLite, a dataset containing 500 high-quality user issues focused on real-world SQLite database applications. Along with the dataset, we also release three RL-trained models: BIRD-Talon-14B, BIRD-Talon-7B, and BIRD-Zeno-7B. The schema file is included in the code repository https://github.com/bird-bench/BIRD-CRITIC-1/blob/main/baseline/data/sqlite_schema.jsonl BIRD-CRITIC-1.0-SQLite BIRD-Critic is the first SQL debugging… See the full description on the dataset page: https://huggingface.co/datasets/birdsql/bird-critic-1.0-sqlite.textn<1K2 likes2k downloads6mo agoHugging Face26Reza8848 /AAAR-1.0 AAAR-1.0 Benchmark 📄 Paper: https://huggingface.co/papers/2410.22394 🌐 Website: https://renzelou.github.io/AAAR-1.0/ 🤗 This repository contains the AAAR-1.0 benchmark dataset. 🚨 Please DO NOT use the data for training! Get Performance on AAAR-1.0 Please refer to our code repository for detailed instructions on running various LLMs on the AAAR-1.0 benchmark, and report the performances. Data Details 1. Equation Inference 🌟:… See the full description on the dataset page: https://huggingface.co/datasets/Reza8848/AAAR-1.0.5 likes2k downloads2y agoHugging Face27changelinglab /cv-v1.0-segment CommonVoice v1 Phone-Segment Alignments Phone-level time alignments for 10 languages of Mozilla Common Voice, packaged in a canonical segmentation schema with embedded 16 kHz audio. The phone boundaries come from the charsiu/cv_ali release of MFA alignments; the audio and transcripts come from Common Voice Corpus 13.0 (2023-03-09). Dataset summary lang train rows train hrs val rows val hrs test rows test hrs en 1,008,669 1,354.0 3,537 4.9 1,285 1.7 rw… See the full description on the dataset page: https://huggingface.co/datasets/changelinglab/cv-v1.0-segment.audioautomatic-speech-recognition1M<n<10M3 likes2k downloads6mo agoHugging Face28chris745820 /more_peoples_speech_v1.0 Dataset Card for People's Speech Dataset Summary The People's Speech Dataset is among the world's largest English speech recognition corpus today that is licensed for academic and commercial usage under CC-BY-SA and CC-BY 4.0. It includes 30,000+ hours of transcribed speech in English languages with a diverse set of speakers. This open dataset is large enough to train speech-to-text systems and crucially is available with a permissive license. Supported Tasks… See the full description on the dataset page: https://huggingface.co/datasets/chris745820/more_peoples_speech_v1.0.automatic-speech-recognition0 likes2k downloads9mo agoHugging Face29LLMcompe-Team-Watanabe /hle_labeled-v1.0 HLE Labeled Dataset このデータセットは「Humanity’s Last Exam」ベンチマーク用のデータセットの categoryにsubcategoryを追加したものです。subcategoryの分類ラベルはqwen/qwen3-235b-a22bで生成しています。 モデルのカテゴリ別の評価に利用するのが目的です。 データ構造 変更点はもともとのHLEにsubcategoryフィールド追加したのみです。 id: レコードのユニークID question: 問題文(文字列) answer: 正解 answer_type: "exactMatch" などの解答形式 rationale: 解答手順・根拠 category: 大分類(例: "Math") subcategory: 小分類のリスト(例: ["Math/Number Theory","Math/Discrete Mathematics"]) image, image_preview, rationale_image:… See the full description on the dataset page: https://huggingface.co/datasets/LLMcompe-Team-Watanabe/hle_labeled-v1.0.text1K<n<10K0 likes1.8k downloads1y agoHugging Face30Muesli1 /dclm-baseline-1.0-llama3-tokenized-shuffled !! Note: this dataset is currently being uploaded and processed. The .bin files are intermediate files to allow shuffling. !! DCLM-Baseline Pretokenized (LLaMA 3.1, 8192 context) This dataset is a pretokenized and globally shuffled version of DCLM-Baseline (mlfoundations/dclm-baseline-1.0), prepared for large-scale language model pretraining. It is intended to be used as a direct drop-in pretraining corpus for LLaMA 3.1 style training pipelines. The original DCLM-Baseline… See the full description on the dataset page: https://huggingface.co/datasets/Muesli1/dclm-baseline-1.0-llama3-tokenized-shuffled.text0 likes1.8k downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.