CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01simple-world-lab /HiFi-UMI-2K HiFi-UMI-2K: High-Fidelity Robot-Free Manipulation Data 2,000 hours released · 6 synchronized camera views · 480+ scenes · 3 mm pose accuracy · <40 µs synchronization 🌐 Project Website | 📦 Dataset | 📄 Paper: arXiv:2607.25895 Examples from the HiFi-UMI corpus. Click the image to play the video. 📚 Introduction HiFi-UMI is a portable, high-fidelity bimanual capture system for collecting robot-free manipulation demonstrations.… See the full description on the dataset page: https://huggingface.co/datasets/simple-world-lab/HiFi-UMI-2K.tabularrobotics100M<n<1B55 likes113k downloads2mo agoHugging Face02Stage-jh-monitor /appworld-qwen35-4b-manysource-2k-newprompt-solvability-junhee-epoch8-tmp01-reeval1 appworld-qwen35-4b-manysource-2k-newprompt-solvability-junhee-epoch8-tmp01-reeval1 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.40546875 Action score: 0.475 Valid samples: 320/320 tabularn<1K0 likes4.5k downloads17d agoHugging Face03Stage-jh-monitor /appworld-qwen35-4b-manysource-2k-newprompt-solvability-junhee-epoch8-reeval1 appworld-qwen35-4b-manysource-2k-newprompt-solvability-junhee-epoch8-reeval1 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.4046875 Action score: 0.4703125 Valid samples: 320/320 tabularn<1K0 likes4.5k downloads17d agoHugging Face04Stage-jh-monitor /appworld-qwen35-4b-manysource-2k-newprompt-solvability-junhee-epoch8 appworld-qwen35-4b-manysource-2k-newprompt-solvability-junhee-epoch8 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.39921875 Action score: 0.44375 Valid samples: 320/320 tabularn<1K0 likes4.5k downloads17d agoHugging Face05Stage-jh-monitor /appworld-qwen35-4b-manysource-2k-newprompt-solvability-junhee-epoch8-t01 appworld-qwen35-4b-manysource-2k-newprompt-solvability-junhee-epoch8-t01 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.38359375 Action score: 0.4703125 Valid samples: 320/320 tabularn<1K0 likes4.5k downloads17d agoHugging Face06Eedi /Question-Anchored-Tutoring-Dialogues-2k Question-Anchored-Tutoring-Dialogues-2k This dataset contains dialogues from math tutoring interventions recorded on Eedi. Dataset Details Dataset Description Each dialogue represents a chat-based conversation between a tutor and a student prompted by the student requesting assistance while working on a lesson. Dialogues are accompanied with 2 sources of meta-data: DQ-Question-Metadata: The question the student was working on that prompted the tutoring… See the full description on the dataset page: https://huggingface.co/datasets/Eedi/Question-Anchored-Tutoring-Dialogues-2k.tabulartext-generation10K<n<100K10 likes488 downloads7mo agoHugging Face07a3124371940 /radeonvla_reflex_physical_2k RadeonVLA-Reflex Physical-2K Physical-2K contains 2,000 strictly validated successful Genesis episodes for language-conditioned Franka fruit sorting. Coverage is exactly five fruits × four bowl positions × 100 episodes = 2,000 episodes. Verified release facts Item Value Episodes 2,000 Frames 468,889 Registered task variations 20 Episodes per task variation 100 Control frequency 20 Hz Strict success certificates 2,000 Sampled image frames 1… See the full description on the dataset page: https://huggingface.co/datasets/a3124371940/radeonvla_reflex_physical_2k.tabularroboticsn<1K0 likes277 downloads2mo agoHugging Face08Satgoy152 /Muse-Glimmer-SWE-Gym-2k Muse-Glimmer-SWE-Gym-2k Agentic coding traces from meta-models/Muse-Glimmer-30B, recorded for training a speculative-decoding drafter. 1,981 mini-swe-agent trajectories over SWE-Gym and SWE-bench-extra instances, and the 159,999 individual chat-completion calls behind them. Configs Config Rows Size What it is train 1,981 57 MB One row per trajectory: the full conversation as messages. raw 159,999 2.7 GB One row per recorded API call: request and… See the full description on the dataset page: https://huggingface.co/datasets/Satgoy152/Muse-Glimmer-SWE-Gym-2k.tabulartext-generation100K<n<1M2 likes254 downloads23d agoHugging Face09data-is-better-together /fineweb2-2k-samplestabular100K<n<1M0 likes216 downloads2y agoHugging Face10ESGBERT /environmental_2ktabular1K<n<10K2 likes137 downloads3y agoHugging Face11ESGBERT /governance_2ktabular1K<n<10K0 likes132 downloads3y agoHugging Face12ESGBERT /social_2ktabular1K<n<10K1 likes122 downloads3y agoHugging Face13mlfoundations-dev /sci_question_exp__scp_116k__training_2k_for_GPQA_eval_03-11-25_17-16-22_f912 mlfoundations-dev/sci_question_exp__scp_116k__training_2k_for_GPQA_eval_03-11-25_17-16-22_f912 Precomputed model outputs for evaluation. Evaluation Results GPQADiamond Average Accuracy: 26.94% ± 4.54% Number of Runs: 3 Run Accuracy Questions Solved Total Questions 1 19.70% 39 198 2 23.23% 46 198 3 37.88% 75 198 tabularn<1K1 likes116 downloads2y agoHugging Face14SWE-Router /v3-2k-traj-claude-opus-4.7tabularn<1K2 likes75 downloads5mo agoHugging Face15rasdani /github-patches-genesys-swe-prompt-2k-context-1k-difftabular1K<n<10K0 likes72 downloads1y agoHugging Face16bonnieliu2002 /eval_act_collect_empty_bottle_black_white_wrist_2k_bs8_k20_testingThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "lekiwi_client", "total_episodes": 1, "total_frames": 1336, "total_tasks": 1, "total_videos": 1, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:1" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/bonnieliu2002/eval_act_collect_empty_bottle_black_white_wrist_2k_bs8_k20_testing.tabularrobotics1K<n<10K0 likes68 downloads1y agoHugging Face17mbitai /secops-2k SecOps-2k 2,000 synthetic security log lines with template labels in LogHub format. Parser papers mostly test on system logs (HDFS, BGL, Apache). No comparable set existed for security telemetry, so we built one: sshd, sudo, firewall, and audit lines, all invented. No real hosts, users, or IPs. Authors: TMFNK and MbitAI. Generator code: TMFNK/LogParser-Dataset. Archived release: doi:10.5281/zenodo.22341506. What is inside One host (secops-01), one day (14 Jun), 2… See the full description on the dataset page: https://huggingface.co/datasets/mbitai/secops-2k.tabularother1K<n<10K0 likes67 downloads20d agoHugging Face18rasdani /github-patches-genesys-2k-context-1k-difftabular1K<n<10K0 likes57 downloads1y agoHugging Face19rasdani /github-patches-genesys-agentless-prompt-2k-context-1k-difftabular10K<n<100K0 likes57 downloads1y agoHugging Face20prestonfu /OpenMathReasoning-subset30kfiltered-Qwen3-1.7B-2k-concise-with-answertabular10K<n<100K0 likes52 downloads10mo agoHugging Face21rasdani /github-patches-genesys-agentless-prompt-2k-context-1k-diff-2k-examplestabular1K<n<10K0 likes50 downloads1y agoHugging Face22suehyunpark /gsm_infinite_symbolic_2ktabular10K<n<100K0 likes48 downloads1y agoHugging Face23cfierro /c4-en-2k-tos-game-replay Fixed English C4 replay subset A subset of allenai/c4, English configuration, training split. C4 is derived from Common Crawl; see the upstream card for provenance and licensing. Sized against cfierro/tos_game_synthetic_docs, split train, using raw text tokens without special tokens or truncation. All 6,219 documents are in train, with 2,982,687 raw tokens. Whole documents are kept until the target is reached; exact duplicate texts are skipped. id is SHA-256 of the original… See the full description on the dataset page: https://huggingface.co/datasets/cfierro/c4-en-2k-tos-game-replay.tabular1K<n<10K0 likes47 downloads19d agoHugging Face24JackHsieh /statML-arxiv-RL-2k-docsTrain-only prefix of JackHsieh/statML-arxiv-RL-4k-docs's train split, drawn from JackHsieh/statML-arxiv. Each row is one randomly sampled contiguous window of exactly 4_096 Qwen3 tokens (Qwen/Qwen3-4B-Instruct-2507) from a distinct paper. start_index is the window's offset in the source paper's token sequence; input_ids is the Qwen3 encoding of text (no special tokens added — no BOS/EOS). Same schema and recipe as JackHsieh/statML-arxiv-40M-20M. Nesting: these are the first 2_048 rows of… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/statML-arxiv-RL-2k-docs.tabular1K<n<10K0 likes45 downloads17d agoHugging Face25prestonfu /OpenMathReasoning-subset30kfiltered-Qwen3-1.7B-2k-v5tabular10K<n<100K0 likes43 downloads9mo agoHugging Face26bonnieliu2002 /eval_act_collect_empty_bottle_black_white_wrist_2k_bs8_k50_tec05_testing1This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "lekiwi_client", "total_episodes": 1, "total_frames": 1140, "total_tasks": 1, "total_videos": 1, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:1" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/bonnieliu2002/eval_act_collect_empty_bottle_black_white_wrist_2k_bs8_k50_tec05_testing1.tabularrobotics1K<n<10K0 likes42 downloads1y agoHugging Face27prestonfu /OpenMathReasoning-subset30kfiltered-Qwen3-1.7B-2k-concisetabular10K<n<100K0 likes42 downloads10mo agoHugging Face28prestonfu /OpenMathReasoning-subset30k-Qwen3-1.7B-2k-concise-with-answertabular10K<n<100K0 likes40 downloads10mo agoHugging Face29prestonfu /OpenMathReasoning-subset30k-Qwen3-1.7B-2k-concisetabular10K<n<100K0 likes39 downloads10mo agoHugging Face30JWei05 /swe_smith_py_trajs_2k_incompletetabular1K<n<10K0 likes39 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.