CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Shofo /shofo-tiktok-general-small Shofo TikTok General (Small) Overview Shofo TikTok General (Small) is a dataset containing 50,000 TikTok videos with comprehensive metadata, transcripts, comments, and engagement metrics. This is a curated subset of Shofo's larger TikTok index, which contains hundreds of millions of indexed videos. Size: ~50K videos (~500GB) Modality: Video + Audio + Text (transcripts, comments, captions) Source: TikTok Schema Column Type Description file_name… See the full description on the dataset page: https://huggingface.co/datasets/Shofo/shofo-tiktok-general-small.tabularvideo-classification10K<n<100K22 likes7.3k downloads7mo agoHugging Face02natolambert /GeneralThought-430K-filteredData from https://huggingface.co/datasets/GeneralReasoning/GeneralThought-430K, removed prompts with non commercial data tabular100K<n<1M35 likes3.4k downloads1y agoHugging Face03General-Medical-AI /GMAI-VL-5.5M GMAI-VL-5.5M Dataset GMAI-VL-5.5M is a comprehensive, large-scale medical General Medical AI Vision-Language (GMAI-VL) dataset built specifically for training multimodal foundation models in the medical domain. It contains an extraordinary scale of high-quality instructions encompassing over 5.5 million multimodal question-answering pairs, carefully constructed based on hundreds of medical classification, segmentation, and detection datasets. This repository… See the full description on the dataset page: https://huggingface.co/datasets/General-Medical-AI/GMAI-VL-5.5M.imagevisual-question-answering1M<n<10M6 likes2.9k downloads5mo agoHugging Face04RJT1990 /GeneralThoughtArchive GeneralThought-430K Thought wants to be free Open reasoning data for March 14 2025. This dataset was part of a side-project in the weeks following the R1 release by Chengxi and Ross - we are no longer maintaining this dataset but are archiving it here. The dataset contains questions, reference answers, reasoning traces, final answers and other metadata from several popular reasoning models including DeepSeek-R1, DeepSeek-R1-Zero, OpenThoughts-32B, LIMO… See the full description on the dataset page: https://huggingface.co/datasets/RJT1990/GeneralThoughtArchive.tabular100K<n<1M79 likes2.7k downloads1y agoHugging Face05OpenOneRec /OpenOneRec-General-Pretrain 通用文本数据集 本目录包含 OpenOneRec 项目使用的通用文本数据集信息。这些数据集均来自 HuggingFace,经过清洗处理和对齐到项目统一的数据格式,并转换为 Parquet 格式用于训练。 数据格式说明 所有数据集均已转换为统一的 Parquet 格式,符合项目的数据格式规范(参考 ../README.md)。数据格式支持: Segments 格式:用于普通文本数据,使用 segments 字段存储文本段落列表 Chat 格式:用于对话数据,使用 messages 字段存储对话消息列表 每个 Parquet 文件包含以下核心字段: uuid: 唯一标识符 source: 数据来源标识 metadata: JSON 格式的元数据字典 segments 或 messages: 文本内容(根据数据类型选择) 详细的数据格式规范请参考 ../README.md。 数据集列表 数据集名称 样本数量 HuggingFace 仓库 reasoning_v1_20m 1,666… See the full description on the dataset page: https://huggingface.co/datasets/OpenOneRec/OpenOneRec-General-Pretrain.tabular1M<n<10M3 likes1.3k downloads9mo agoHugging Face06BlidReview /steady-rans-generalization Steady-RANS cross-family generalization dataset Data for the paper "Towards generalized flow field prediction: one model across unseen object families" (under double blind review; this account is anonymous for that reason). Trained checkpoints and evaluation code are in the companion model repo: steady-rans-surrogates. Steady incompressible k-omega SST (OpenFOAM simpleFoam) external flow around 855 distinct shapes (17 scripted parametric families plus 40 ModelNet object… See the full description on the dataset page: https://huggingface.co/datasets/BlidReview/steady-rans-generalization.3d1K<n<10K0 likes1.3k downloads1mo agoHugging Face07cjfcsjt /AITW_Generaltabular100K<n<1M2 likes1.1k downloads2y agoHugging Face08Brainquiver /general-master-en-202608 General · Master · English · 2026-08 English pretraining text, assembled from three public sources, cleaned with one character-level cleaner, and filtered for repetition. 109,337,531 documents and 468,064,046,462 characters. Composition Config Documents Characters What it is fineweb-edu-dedup 65,010,430 297,544,916,118 Web text an educational classifier kept cosmopedia-v2 38,591,146 144,011,993,012 Synthetic prose from a seeded generator… See the full description on the dataset page: https://huggingface.co/datasets/Brainquiver/general-master-en-202608.tabulartext-generation100M<n<1B1 likes893 downloads23d agoHugging Face09Hkang /moda-general-capability-rollouts MODA General Capability Retention Rollouts This dataset contains the raw model generations and evaluation results for the MODA general-capability retention experiments. It covers 16 models, seven benchmarks, 260,592 prompt records, and 2,605,920 stored generations. The evaluation code is pinned to source commit 12ea99b2a57a354f2b7d6792f62a3d9313192fa7. Evaluation protocol Benchmarks: GSM8K, MMLU abstract_algebra, GPQA Diamond, BoolQ, HellaSwag, TruthfulQA, and… See the full description on the dataset page: https://huggingface.co/datasets/Hkang/moda-general-capability-rollouts.tabulartext-generation1M<n<10M0 likes582 downloads1mo agoHugging Face10Brainquiver /general-web-it-202608 General · Web · Italian · 2026-08 Italian pretraining text, built from the Italian portion of EPFL's FineWeb2-HQ, which is the high quality slice of FineWeb-2. Every document passes one character-level cleaner and a repetition filter. 21,065,052 documents and 66,158,573,443 characters of Italian prose. Contents Config Documents Characters Upstream fineweb2-hq-ita_Latn 21,065,052 66,158,573,443 epfml/FineWeb2-HQ, ita_Latn The character count is… See the full description on the dataset page: https://huggingface.co/datasets/Brainquiver/general-web-it-202608.tabulartext-generation10M<n<100M0 likes522 downloads23d agoHugging Face11Brainquiver /general-web-fr-202608 General · Web · French · 2026-08 French pretraining text, built from the French portion of EPFL's FineWeb2-HQ, which is the high quality slice of FineWeb-2. Every document passes one character-level cleaner and a repetition filter. 31,999,309 documents and 118,346,333,763 characters of French prose. Contents Config Documents Characters Upstream fineweb2-hq-fra_Latn 31,999,309 118,346,333,763 epfml/FineWeb2-HQ, fra_Latn The character count is exact.… See the full description on the dataset page: https://huggingface.co/datasets/Brainquiver/general-web-fr-202608.tabulartext-generation10M<n<100M0 likes476 downloads23d agoHugging Face12yufan /SFT_Chinese_Generaltabular1M<n<10M7 likes441 downloads2y agoHugging Face13mariiakoroliuk /generalization-science-datadocumentn<1K0 likes404 downloads1d agoHugging Face14girardijp /sam_general_smolvlaThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "bi_sam_follower", "total_episodes": 5, "total_frames": 5282, "total_tasks": 1, "total_videos": 20, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:5" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/girardijp/sam_general_smolvla.tabularrobotics100K<n<1M0 likes374 downloads1y agoHugging Face15arcadia-impact /reward-projection-goal-generalisation-vlmtabular1K<n<10K0 likes358 downloads2mo agoHugging Face16hivemind-research /general-layerC-200klanguage: en license: apache-2.0 size_categories: 100K<n<1M pretty_name: General LayerC — training-ready (200,000 samples) tags: task_categories:text-generation task_categories:text2text-generation language:en license:apache-2.0 domain:general reasoning synthetic sft general-knowledge science creative-writing glaive deepseek-r1-distill distillation layer-c layer-training-ready General LayerC (training-ready) — 200k Built from glaiveai/reasoning-v1-20m (spread-sampled across… See the full description on the dataset page: https://huggingface.co/datasets/hivemind-research/general-layerC-200k.tabular100K<n<1M0 likes273 downloads2mo agoHugging Face17masato-ka /SO100_evaluate_generalize_pick_posThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so100_follower", "total_episodes": 90, "total_frames": 33529, "total_tasks": 1, "total_videos": 180, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:90"}, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/masato-ka/SO100_evaluate_generalize_pick_pos.tabularrobotics10K<n<100K0 likes183 downloads1y agoHugging Face18roboticshack /team15-groot-two-task-generalization mvp This dataset was generated using a phospho starter pack. This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot and RLDS. tabularrobotics10K<n<100K1 likes145 downloads1y agoHugging Face19akabhinav32 /kikobot_final_50_generalThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "kikobot", "total_episodes": 83, "total_frames": 45277, "total_tasks": 1, "total_videos": 166, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:83" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/akabhinav32/kikobot_final_50_general.tabularrobotics10K<n<100K0 likes136 downloads1y agoHugging Face20kshitijthakkar /nemotron-sft-general-focused-stage1-2-ChatML-V3 Nemotron SFT Dataset (Chat Template Formatted) Overview This dataset is a curated supervised fine-tuning (SFT) dataset built from NVIDIA's Nemotron-Cascade-SFT-Stage-1 and Stage-2 datasets. Important: This dataset uses the tokenizer's apply_chat_template() method to properly format conversations from the original messages/conversations fields. Statistics Total Samples: 496,385 Total Tokens: 1,114,218,401 Average Tokens per Sample: 2244.7 Tokenizer:… See the full description on the dataset page: https://huggingface.co/datasets/kshitijthakkar/nemotron-sft-general-focused-stage1-2-ChatML-V3.tabular100K<n<1M0 likes132 downloads8mo agoHugging Face21hieupham14022003 /general_speech_distortedtabular100K<n<1M0 likes128 downloads1y agoHugging Face22luyu1021 /seedance_general_all_dance_scm_latent_lmdb Seedance General-All + Dance SCM Latent LMDB This dataset stores precomputed SCM latents used for TurboT2AV training. Source mapping: seedance_general_all_dance_mapping.csv Successful latent samples: 44,305 Shards: 8 LMDB shards under scm_latent_lmdb/shard_00000 ... shard_00007 Video latent shape per sample: (1, 16, 128, 16, 24) Audio latent shape per sample: (1, 127, 128) The source mapping combines Seedance general-all data with a dance subset. The mapping contains 44,504… See the full description on the dataset page: https://huggingface.co/datasets/luyu1021/seedance_general_all_dance_scm_latent_lmdb.tabulartext-to-video10K<n<100K0 likes118 downloads3mo agoHugging Face23masato-ka /SO100_evaluate_generalize_pick_pos_extendThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so100_follower", "total_episodes": 30, "total_frames": 11184, "total_tasks": 1, "total_videos": 60, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:30"}, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/masato-ka/SO100_evaluate_generalize_pick_pos_extend.tabularrobotics10K<n<100K0 likes117 downloads1y agoHugging Face24generalagents /showdown-clicks showdown-clicks General Agents 🤗 Dataset | GitHub showdown is a suite of offline and online benchmarks for computer-use agents. showdown-clicks is a collection of 5,679 left clicks of humans performing various tasks in a macOS desktop environment. It is intended to evaluate instruction-following and low-level control capabilities of computer-use agents. As of March 2025, we are releasing a subset of the full set, showdown-clicks-dev, containing 557 clicks. All examples are… See the full description on the dataset page: https://huggingface.co/datasets/generalagents/showdown-clicks.imageimage-to-textn<1K16 likes98 downloads1y agoHugging Face25letrinhan /vn-provinces-ethnic-minority-general-school-pupils Vietnam ethnic-minority general school pupils by level Vietnam ethnic-minority general school pupils by level. Geographic labels are English (UN/GSO style ASCII romanization). Tables cover provinces, regions and national total where present. Province names follow ar_core.vn_geo (historical 63-province system). Figures Hero Hero (continued) Comparison Color key Files provinces (1152 rows) data/provinces.csv data/provinces.dta data/provinces.xlsx… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-ethnic-minority-general-school-pupils.tabular1K<n<10K0 likes91 downloads1d agoHugging Face26YangZhoumill /infini_gsm_4k_noise_generaltabular10K<n<100K0 likes88 downloads2y agoHugging Face27agentjudge-anon /GeneralAgentBench GeneralAgentBench GeneralAgentBench is a 1,400+ task benchmark for evaluating whether general-purpose AI agents have genuinely completed a task, spanning Mobile / Browser / Desktop environments. It is the evaluation resource accompanying an anonymous NeurIPS 2026 Evaluations & Datasets Track submission (Submission 173, AgentJudge). This release is fully anonymized for double-blind review. Each task provides a natural-language instruction plus a list of verification checkpoints.… See the full description on the dataset page: https://huggingface.co/datasets/agentjudge-anon/GeneralAgentBench.documentother1K<n<10K0 likes87 downloads2mo agoHugging Face28letrinhan /vn-provinces-general-schools Vietnam provinces general schools by school type Vietnam provinces general schools by school type. Geographic labels are English (UN/GSO style ASCII romanization). Tables cover provinces, regions and national total where present. Province names follow ar_core.vn_geo (historical 63-province system). Figures Hero Hero (continued) Comparison Color key Files provinces (1260 rows) data/provinces.csv data/provinces.dta data/provinces.xlsx regions… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-general-schools.tabular1K<n<10K0 likes85 downloads1d agoHugging Face29harness-generalization /dataset Harness Generalization Rollouts Evolution-run rollouts for Qwen3-4B-Instruct-2507, Qwen2.5-3B-Instruct, gpt-oss-120b, and gpt-oss-20b. One Parquet file per run is stored at data/<task>/<model>/<configuration>/<timestamp>.parquet. tabular1M<n<10M0 likes81 downloads1mo agoHugging Face30letrinhan /vn-provinces-ethnic-minority-general-school-teachers Vietnam ethnic-minority general school teachers (selected provinces) Vietnam ethnic-minority general school teachers (selected provinces). Geographic labels are English (UN/GSO style ASCII romanization). Tables cover provinces, regions and national total where present. Province names follow ar_core.vn_geo (historical 63-province system). Figures Hero Hero (continued) Comparison Color key Files provinces (702 rows) data/provinces.csv… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-ethnic-minority-general-school-teachers.tabularn<1K0 likes81 downloads1d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.