CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01natolambert /GeneralThought-430K-filteredData from https://huggingface.co/datasets/GeneralReasoning/GeneralThought-430K, removed prompts with non commercial data tabular100K<n<1M35 likes3k downloads2y agoHugging Face02RJT1990 /GeneralThoughtArchive GeneralThought-430K Thought wants to be free Open reasoning data for March 14 2025. This dataset was part of a side-project in the weeks following the R1 release by Chengxi and Ross - we are no longer maintaining this dataset but are archiving it here. The dataset contains questions, reference answers, reasoning traces, final answers and other metadata from several popular reasoning models including DeepSeek-R1, DeepSeek-R1-Zero, OpenThoughts-32B, LIMO… See the full description on the dataset page: https://huggingface.co/datasets/RJT1990/GeneralThoughtArchive.tabular100K<n<1M79 likes2.1k downloads1y agoHugging Face03OpenOneRec /OpenOneRec-General-Pretrain 通用文本数据集 本目录包含 OpenOneRec 项目使用的通用文本数据集信息。这些数据集均来自 HuggingFace,经过清洗处理和对齐到项目统一的数据格式,并转换为 Parquet 格式用于训练。 数据格式说明 所有数据集均已转换为统一的 Parquet 格式,符合项目的数据格式规范(参考 ../README.md)。数据格式支持: Segments 格式:用于普通文本数据,使用 segments 字段存储文本段落列表 Chat 格式:用于对话数据,使用 messages 字段存储对话消息列表 每个 Parquet 文件包含以下核心字段: uuid: 唯一标识符 source: 数据来源标识 metadata: JSON 格式的元数据字典 segments 或 messages: 文本内容(根据数据类型选择) 详细的数据格式规范请参考 ../README.md。 数据集列表 数据集名称 样本数量 HuggingFace 仓库 reasoning_v1_20m 1,666… See the full description on the dataset page: https://huggingface.co/datasets/OpenOneRec/OpenOneRec-General-Pretrain.tabular1M<n<10M3 likes1.1k downloads9mo agoHugging Face04cjfcsjt /AITW_Generaltabular100K<n<1M2 likes1.1k downloads2y agoHugging Face05Brainquiver /general-master-en-202608 General · Master · English · 2026-08 English pretraining text, assembled from three public sources, cleaned with one character-level cleaner, and filtered for repetition. 109,337,531 documents and 468,064,046,462 characters. Composition Config Documents Characters What it is fineweb-edu-dedup 65,010,430 297,544,916,118 Web text an educational classifier kept cosmopedia-v2 38,591,146 144,011,993,012 Synthetic prose from a seeded generator… See the full description on the dataset page: https://huggingface.co/datasets/Brainquiver/general-master-en-202608.tabulartext-generation100M<n<1B1 likes983 downloads27d agoHugging Face06Brainquiver /general-web-it-202608 General · Web · Italian · 2026-08 Italian pretraining text, built from the Italian portion of EPFL's FineWeb2-HQ, which is the high quality slice of FineWeb-2. Every document passes one character-level cleaner and a repetition filter. 21,065,052 documents and 66,158,573,443 characters of Italian prose. Contents Config Documents Characters Upstream fineweb2-hq-ita_Latn 21,065,052 66,158,573,443 epfml/FineWeb2-HQ, ita_Latn The character count is… See the full description on the dataset page: https://huggingface.co/datasets/Brainquiver/general-web-it-202608.tabulartext-generation10M<n<100M0 likes561 downloads27d agoHugging Face07Brainquiver /general-web-fr-202608 General · Web · French · 2026-08 French pretraining text, built from the French portion of EPFL's FineWeb2-HQ, which is the high quality slice of FineWeb-2. Every document passes one character-level cleaner and a repetition filter. 31,999,309 documents and 118,346,333,763 characters of French prose. Contents Config Documents Characters Upstream fineweb2-hq-fra_Latn 31,999,309 118,346,333,763 epfml/FineWeb2-HQ, fra_Latn The character count is exact.… See the full description on the dataset page: https://huggingface.co/datasets/Brainquiver/general-web-fr-202608.tabulartext-generation10M<n<100M0 likes510 downloads27d agoHugging Face08yufan /SFT_Chinese_Generaltabular1M<n<10M7 likes443 downloads2y agoHugging Face09Hkang /moda-general-capability-rollouts MODA General Capability Retention Rollouts This dataset contains the raw model generations and evaluation results for the MODA general-capability retention experiments. It covers 16 models, seven benchmarks, 260,592 prompt records, and 2,605,920 stored generations. The evaluation code is pinned to source commit 12ea99b2a57a354f2b7d6792f62a3d9313192fa7. Evaluation protocol Benchmarks: GSM8K, MMLU abstract_algebra, GPQA Diamond, BoolQ, HellaSwag, TruthfulQA, and… See the full description on the dataset page: https://huggingface.co/datasets/Hkang/moda-general-capability-rollouts.tabulartext-generation1M<n<10M0 likes430 downloads2mo agoHugging Face10lszoszk /treaty-bodies-general-comments Treaty Bodies General Comments A paragraph-level dataset of General Comments and General Recommendations adopted by the nine UN human-rights Treaty Bodies, with concerned-group labels and document metadata. Companion to the UNHRD search interface. Licence The curated dataset (paragraph segmentation, label annotation, document metadata enrichment, footnote and section work) is released under Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International… See the full description on the dataset page: https://huggingface.co/datasets/lszoszk/treaty-bodies-general-comments.tabulartext-classification1K<n<10K0 likes394 downloads22d agoHugging Face11girardijp /sam_general_smolvlaThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "bi_sam_follower", "total_episodes": 5, "total_frames": 5282, "total_tasks": 1, "total_videos": 20, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:5" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/girardijp/sam_general_smolvla.tabularrobotics100K<n<1M0 likes384 downloads1y agoHugging Face12hivemind-research /general-layerC-200klanguage: en license: apache-2.0 size_categories: 100K<n<1M pretty_name: General LayerC — training-ready (200,000 samples) tags: task_categories:text-generation task_categories:text2text-generation language:en license:apache-2.0 domain:general reasoning synthetic sft general-knowledge science creative-writing glaive deepseek-r1-distill distillation layer-c layer-training-ready General LayerC (training-ready) — 200k Built from glaiveai/reasoning-v1-20m (spread-sampled across… See the full description on the dataset page: https://huggingface.co/datasets/hivemind-research/general-layerC-200k.tabular100K<n<1M0 likes332 downloads2mo agoHugging Face13masato-ka /SO100_evaluate_generalize_pick_posThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so100_follower", "total_episodes": 90, "total_frames": 33529, "total_tasks": 1, "total_videos": 180, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:90" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/masato-ka/SO100_evaluate_generalize_pick_pos.tabularrobotics10K<n<100K0 likes191 downloads1y agoHugging Face14akabhinav32 /kikobot_final_50_generalThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "kikobot", "total_episodes": 83, "total_frames": 45277, "total_tasks": 1, "total_videos": 166, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:83" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/akabhinav32/kikobot_final_50_general.tabularrobotics10K<n<100K0 likes142 downloads1y agoHugging Face15YangZhoumill /infini_gsm_4k_noise_generaltabular10K<n<100K0 likes117 downloads2y agoHugging Face16masato-ka /SO100_evaluate_generalize_pick_pos_extendThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so100_follower", "total_episodes": 30, "total_frames": 11184, "total_tasks": 1, "total_videos": 60, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:30" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/masato-ka/SO100_evaluate_generalize_pick_pos_extend.tabularrobotics10K<n<100K0 likes111 downloads1y agoHugging Face17YangZhoumill /infini_gsm_8k_noise_generaltabular10K<n<100K0 likes102 downloads2y agoHugging Face18roboticshack /team15-groot-two-task-generalization mvp This dataset was generated using a phospho starter pack. This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot and RLDS. tabularrobotics10K<n<100K1 likes92 downloads1y agoHugging Face19ar0s /centrifuge-generalization-v1This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "ee_x.pos", "ee_y.pos", "ee_z.pos", "ee_wx.pos", "ee_wy.pos", "ee_wz.pos", "gripper.pos" ], "shape": [ 7… See the full description on the dataset page: https://huggingface.co/datasets/ar0s/centrifuge-generalization-v1.tabularrobotics10K<n<100K0 likes82 downloads5d agoHugging Face20harness-generalization /dataset Harness Generalization Rollouts Evolution-run rollouts for Qwen3-4B-Instruct-2507, Qwen2.5-3B-Instruct, gpt-oss-120b, and gpt-oss-20b. One Parquet file per run is stored at data/<task>/<model>/<configuration>/<timestamp>.parquet. tabular1M<n<10M0 likes78 downloads1mo agoHugging Face21robot-learning-team43 /so101_generalThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "shape": [ 6 ], "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/robot-learning-team43/so101_general.tabularrobotics100K<n<1M0 likes75 downloads4mo agoHugging Face22allday-technology /general_object_pickupThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "trossen_subversion": "v1.0", "robot_type": "trossen_ai_stationary", "total_episodes": 6, "total_frames": 1706, "total_tasks": 1, "total_videos": 18, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:6" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/allday-technology/general_object_pickup.tabularrobotics10K<n<100K0 likes73 downloads9mo agoHugging Face23GitBag /llama3-uf-meta-general-tokenizedtabular10K<n<100K0 likes70 downloads2y agoHugging Face24Cartinoe5930 /GeneralThought-Sciencetabular10K<n<100K0 likes68 downloads2y agoHugging Face25kshitijthakkar /nemotron-sft-general-focused-stage1-2-ChatML-V2 Nemotron SFT Dataset (Chat Template Formatted) Overview This dataset is a curated supervised fine-tuning (SFT) dataset built from NVIDIA's Nemotron-Cascade-SFT-Stage-1 and Stage-2 datasets. Important: This dataset uses the tokenizer's apply_chat_template() method to properly format conversations from the original messages/conversations fields. Statistics Total Samples: 100,000 Total Tokens: 222,951,633 Average Tokens per Sample: 2229.5 Tokenizer:… See the full description on the dataset page: https://huggingface.co/datasets/kshitijthakkar/nemotron-sft-general-focused-stage1-2-ChatML-V2.tabular100K<n<1M0 likes64 downloads8mo agoHugging Face26hivemind-research /general-layerB-200klanguage: en license: apache-2.0 size_categories: 100K<n<1M pretty_name: General LayerB — teacher-supervision (200,000 samples) tags: task_categories:text-generation task_categories:text2text-generation language:en license:apache-2.0 domain:general reasoning synthetic sft general-knowledge science creative-writing glaive deepseek-r1-distill distillation layer-b layer-teacher-supervision teacher-supervision General LayerB (teacher-supervision) — 200k Built from… See the full description on the dataset page: https://huggingface.co/datasets/hivemind-research/general-layerB-200k.tabular100K<n<1M0 likes63 downloads2mo agoHugging Face27GitBag /llama3-uf-meta-general-tokenized-1024-2048tabular10K<n<100K0 likes61 downloads2y agoHugging Face28eorikun /eval_smolvla_generalThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "so_follower", "total_episodes": 19, "total_frames": 2949, "total_tasks": 8, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 30, "splits": { "train": "0:19" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/eorikun/eval_smolvla_general.tabularrobotics1K<n<10K0 likes57 downloads6mo agoHugging Face29strakammm /generals_io_replays ⚔️ Generals.io High-Rank Replay Dataset 🌟 Overview This dataset contains a curated collection of 1v1 game replays from the online strategy game generals.io, specifically designed for training high-level reinforcement learning agents 🤖. 🏆 High-Quality Matches: Includes games where at least one participant had a star rating of 70 or higher 📈, ensuring a baseline of quality and strategic depth. ✅ Clean Data: Carefully filtered to remove outliers, games with AFK players… See the full description on the dataset page: https://huggingface.co/datasets/strakammm/generals_io_replays.tabular10K<n<100K5 likes54 downloads1y agoHugging Face30jiosephlee /assay-transfer-record-level-v27-carcinogens-general-intern Carcinogens record-level V27-general This level-agnostic release uses current V10 physical evidence and Gold-v1 parent-disjoint validation/test queries. Training pairs stay inside an intrinsic V10 assay bucket and cross normalized parents. Gold level is optional provenance, not a pairing or calibration boundary. support gate: at least 10 records and 4 parents variance ratio: undefined allowed; finite values above 0.75 excluded train record-role degree cap: 24 ID/OOD validation… See the full description on the dataset page: https://huggingface.co/datasets/jiosephlee/assay-transfer-record-level-v27-carcinogens-general-intern.tabular100K<n<1M0 likes52 downloads4d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.