CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mlfoundations /dclm-baseline-1.0 DCLM-baseline DCLM-baseline is a 4T token / 3B document pretraining dataset that achieves strong performance on language model benchmarks. Below are comparisions of model trained on DCLM-baseline with other models in the 7B regime. Model Params Tokens Open dataset? CORE MMLU EXTENDED Open weights, closed datasets Llama2 7B 2T ✗ 49.2 45.8 34.1 DeepSeek 7B 2T ✗ 50.7 48.5 35.3 Mistral-0.3 7B ? ✗ 57.0 62.7 45.1 QWEN-2 7B ? ✗ 57.5 71.9 50.5 Llama3 8B 15T ✗… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dclm-baseline-1.0.tabular1B<n<10B315 likes648k downloads2y agoHugging Face02mlfoundations /dclm-baseline-1.0-parquet DCLM-baseline Note: this is an identical copy of https://huggingface.co/datasets/mlfoundations/dclm-baseline-1.0, where all the files have been mapped to a parquet format. DCLM-baseline is a 4T token / 3B document pretraining dataset that achieves strong performance on language model benchmarks. Below are comparisions of model trained on DCLM-baseline with other models in the 7B regime. Model Params Tokens Open dataset? CORE MMLU EXTENDED Open weights, closed datasets… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dclm-baseline-1.0-parquet.tabular1B<n<10B56 likes19k downloads2y agoHugging Face03lerobot /droid_1.0.1This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "Franka", "total_episodes": 95658, "total_frames": 27630375, "total_tasks": 49630, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 500, "fps": 15, "splits": { "train": "0:95658" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/droid_1.0.1.tabularrobotics10M<n<100M24 likes16k downloads3mo agoHugging Face04argilla /magpie-ultra-v1.0 Dataset Card for magpie-ultra-v1.0 This dataset has been created with distilabel. Dataset Summary magpie-ultra it's a synthetically generated dataset for supervised fine-tuning using the Llama 3.1 405B-Instruct model, together with other Llama models like Llama-Guard-3-8B and Llama-3.1-8B-Instruct. The dataset contains challenging instructions and responses for a wide variety of tasks, such as Coding & debugging, Math, Data analysis, Creative Writing… See the full description on the dataset page: https://huggingface.co/datasets/argilla/magpie-ultra-v1.0.tabular1M<n<10M52 likes8.7k downloads2y agoHugging Face05IPEC-COMMUNITY /libero_spatial_no_noops_1.0.0_lerobottabular10K<n<100K5 likes6.7k downloads11mo agoHugging Face06IPEC-COMMUNITY /libero_object_no_noops_1.0.0_lerobottabular10K<n<100K1 likes6.2k downloads11mo agoHugging Face07IPEC-COMMUNITY /libero_10_no_noops_1.0.0_lerobottabular100K<n<1M3 likes6k downloads11mo agoHugging Face08IPEC-COMMUNITY /libero_goal_no_noops_1.0.0_lerobottabular10K<n<100K1 likes5.9k downloads11mo agoHugging Face09cadene /droid_1.0.1_v30This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "Franka", "total_episodes": 95584, "total_frames": 27607757, "total_tasks": 49596, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 500, "fps": 15, "splits": { "train": "0:95584" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/cadene/droid_1.0.1_v30.tabularrobotics10M<n<100M4 likes2.4k downloads1y agoHugging Face10villekuosmanen /dAgger_build_block_tower_1.0.0-advantages Advantage Values for villekuosmanen/dAgger_build_block_tower_1.0.0 Pre-computed advantage values for offline RL training. Source Dataset: villekuosmanen/dAgger_build_block_tower_1.0.0 Value Model: villekuosmanen/rewact_build_block_tower_all_3 N-step lookahead: 50 Files This dataset contains per-episode parquet files with advantage values for each frame. Usage from pathlib import Path import pandas as pd # Load advantages for a specific episode… See the full description on the dataset page: https://huggingface.co/datasets/villekuosmanen/dAgger_build_block_tower_1.0.0-advantages.tabularrobotics10K<n<100K0 likes2.4k downloads6mo agoHugging Face11aractingi /droid_1.0.1_testThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "Franka", "total_episodes": 95658, "total_frames": 27630375, "total_tasks": 49630, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 15, "splits": { "train": "0:95658" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/aractingi/droid_1.0.1_test.tabularrobotics10M<n<100M0 likes2.2k downloads10mo agoHugging Face12muacha /droid_1.0.1_v30This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "Franka", "total_episodes": 95658, "total_frames": 27630375, "total_tasks": 49630, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 500, "fps": 15, "splits": { "train": "0:95658" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/muacha/droid_1.0.1_v30.tabularrobotics10M<n<100M0 likes1.6k downloads29d agoHugging Face13william94 /chat1.0tabular100K<n<1M0 likes1.1k downloads8mo agoHugging Face14phail-anon /phail-v1.0 PhAIL: Real-Robot VLA Evaluation Benchmark (v1.0) This dataset accompanies an anonymous submission to the NeurIPS 2026 Evaluations and Datasets track. The paper, code, and dataset are all under double-blind review; identifying URLs and author information have been withheld. PhAIL is a real-robot evaluation benchmark for vision-language-action (VLA) policies. It contains synchronized exterior and wrist RGB video, end-effector and gripper telemetry, and per-rollout event… See the full description on the dataset page: https://huggingface.co/datasets/phail-anon/phail-v1.0.tabularrobotics10M<n<100M0 likes1.1k downloads5mo agoHugging Face15juiceb0xc0de /Ornith-1.0-9B-atlas juiceb0xc0de/Ornith-1.0-9B-atlas A brain atlas for deepreinforce-ai/Ornith-1.0-9B, the 9B agentic-coding model that reports SOTA results on Terminal-Bench, SWE-Bench, and other agentic coding benchmarks. This is not a chat dataset or a benchmark — it is an internal-mechanics map built by running activations through a corpus of prompts and scoring what each layer, component, head, and feature direction is doing. If you want to know why this model survives surgical edits, where… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/Ornith-1.0-9B-atlas.image1M<n<10M1 likes1k downloads10d agoHugging Face16aractingi /droid_1.0.1This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "Franka", "total_episodes": 95658, "total_frames": 27630375, "total_tasks": 49630, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 500, "fps": 15, "splits": { "train": "0:95658" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/aractingi/droid_1.0.1.tabularrobotics10M<n<100M5 likes967 downloads10mo agoHugging Face17ASTRAI-labs /Pluto-Nano-1.0-Pretrain-v2 ASTRAI Pluto Nano 1.0 — Pretrain Mix (v2) Curated multilingual pretraining corpus (~50 GB parquet, ~12 B tokens after tokenization) used for ASTRAI Pluto Nano 1.0, a 1 B-total / 50 M-active MoE model with 64 k vocabulary and 5 target languages (EN, PT, ES, ZH, HI). v2 additions vs v1: OpenThoughts3 (CoT reasoning), openstax textbooks + peS2o (science), and reweighting for better balance. NOTE: factsense (openbmb) was used at training time but is not redistributed here due to its… See the full description on the dataset page: https://huggingface.co/datasets/ASTRAI-labs/Pluto-Nano-1.0-Pretrain-v2.tabulartext-generation10M<n<100M2 likes820 downloads3mo agoHugging Face18FT-LLM-2026-RAMEN /droid_1.0.1tabular10M<n<100M0 likes778 downloads8mo agoHugging Face19nyu-dice-lab /lm-eval-results-shyamieee-Padma-SLM-7b-v1.0-private Dataset Card for Evaluation run of shyamieee/Padma-SLM-7b-v1.0 Dataset automatically created during the evaluation run of model shyamieee/Padma-SLM-7b-v1.0 The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-shyamieee-Padma-SLM-7b-v1.0-private.tabular100K<n<1M0 likes515 downloads2y agoHugging Face20ygtxr1997 /droid_1.0.1This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "Franka", "total_episodes": 95617, "total_frames": 27618651, "total_tasks": 49611, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 500, "fps": 15, "splits": { "train": "0:95617" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ygtxr1997/droid_1.0.1.tabularrobotics10M<n<100M0 likes482 downloads8mo agoHugging Face21dgrachev /droid_1.0.1 DROID 1.0.1 for pi05 training: successful, non-idle segments lerobot/droid_1.0.1 rewritten twice, videos untouched. First augment_droid_delta_ee_gripper_events.py (lerobot_policy_framepick) added action.delta_ee, observation.extrinsics.static1/static2/wrist1 and the observation.gripper.time_* columns; the flat action is the delta-EE command and the flat observation.state the EE state. Then benchmarks/robolab/augment_droid_dataset.py (this revision) mirrored the data pipeline of… See the full description on the dataset page: https://huggingface.co/datasets/dgrachev/droid_1.0.1.tabularrobotics10M<n<100M0 likes454 downloads15d agoHugging Face22kothasuhas /dclm-baseline-1.0_subset_30Mtabular10M<n<100M0 likes369 downloads2y agoHugging Face23SM-Bello /PHI-SPIKE-C172x-Community-Dataset-v1.0 PHI-SPIKE C172X Community Dataset v1.0 Dataset Summary PHI-SPIKE C172X Community Dataset v1.0 is a simulation-based aerospace Prognostics and Health Management (PHM) dataset and training-artifact release developed from the PHI-SPIKE C172X research campaign. The release provides: JSBSim C172X reference telemetry; benchmark metadata; training histories; trained PyTorch model checkpoints; per-run evaluation metrics; and five-seed campaign summaries. The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/SM-Bello/PHI-SPIKE-C172x-Community-Dataset-v1.0.tabulartime-series-forecasting1K<n<10K1 likes349 downloads10d agoHugging Face24JackHsieh /4B-predict.rule-r-1.0-k-256.L-1024.statml-arxivtabular1M<n<10M0 likes314 downloads4mo agoHugging Face25Leyo /droid_1.0.1_v30_merged Leyo/droid_1.0.1_v30_merged In-place merged per chunk. One row per episode in each data/chunk-XXX/file-YYY.parquet. Timestamp-level columns are lists; episode-level fields are meta__*. Join key: episode_index. Generated with Polars (lazy/streaming), parallelized across files. Parquet compression: zstd. tabular100K<n<1M0 likes278 downloads1y agoHugging Face26Wanfq /3_4_fusechat_v1_openchat-3.5_mixtral-8x7b-instruct-v0.1_solar-10.7b-instruct-v1.0_representationtabular10K<n<100K0 likes261 downloads2y agoHugging Face27wge118 /Fnii-VLA-Kinova-1.0tabular100K<n<1M1 likes245 downloads1y agoHugging Face28cadene /droid_1.0.1_v30_compact_5This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "Franka", "total_episodes": 95658, "total_frames": 27630375, "total_tasks": 49630, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 500, "fps": 15, "splits": { "train": "0:95658" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/cadene/droid_1.0.1_v30_compact_5.tabularrobotics10M<n<100M1 likes236 downloads1y agoHugging Face29taejoon89 /Ko-Agent-Trajectories-1.0 Ko-Agent-Trajectories-1.0 Dataset card v1.1.1 (2026-09-22). The pipeline code is now released in this repository under pipeline/, together with the API catalogue, the scenario templates and the complete prompt set. The card reports the completed human review study and the v1.1 artefacts (behaviour DPO config, per-item validation scores, manifest, filter asset). Korean edition: README.ko.md. TL;DR A Korean multi-turn agent ↔ tool trajectory corpus synthesized… See the full description on the dataset page: https://huggingface.co/datasets/taejoon89/Ko-Agent-Trajectories-1.0.tabulartext-generation100K<n<1M0 likes234 downloads2d agoHugging Face30OmniAICreator /Qiita-1.07MThis dataset contains 1,074,174 articles published on Qiita. tabulartext-classification1M<n<10M2 likes232 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.