CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01suchirsalhan /beetle-phase-transition Induction-head phase transition in BeetleLM Per-checkpoint mechanistic metrics for Beetle language models, tracking how the induction circuit forms during training. The files here have six different schemas, so they are exposed as separate configs. Loading the directory as a single table fails with a cast error — that is why the configs above exist, not a bug. from datasets import load_dataset traj = load_dataset("suchirsalhan/beetle-phase-transition", "trajectories") abl =… See the full description on the dataset page: https://huggingface.co/datasets/suchirsalhan/beetle-phase-transition.tabularother1K<n<10K0 likes10k downloads26m agoHugging Face02asapp /slue-phase-2 Licensing Information SLUE-HVB SLUE-HVB dataset contains a subset of the Gridspace-Stanford Harper Valley speech dataset and the copyright of this subset remains the same with the original license, CC-BY-4.0. See also original license notice (https://github.com/cricketclub/gridspace-stanford-harper-valley/blob/master/LICENSE) Additionally, we provide dialog act classification annotation and it is covered with the same license as CC-BY-4.0. SLUE-SQA-5… See the full description on the dataset page: https://huggingface.co/datasets/asapp/slue-phase-2.audio10K<n<100K11 likes2.4k downloads3y agoHugging Face03gurukondaveeti /indic-lma-corpus-phase1imagen<1K0 likes1.8k downloads9d agoHugging Face04team-hatakeyama-phase2 /ndlj_tosho_1 国会図書館に収蔵される著作権切れのデータです textn<1K0 likes701 downloads2y agoHugging Face05AbstractPhil /sdxl-qwen-phase0 SDXL–Qwen Phase-0 dataset Purpose-built training set for AbstractPhil/geolip-sdxl-aleph. Each row pairs a Qwen-Image-Lightning render with the caption that produced it and an encoder-invariant geometric "aleph" address derived from the caption's bytes. It exists to retrain SDXL (which stays the base model) around a new text encoder (Qwen in place of CLIP-G) under a rectified-flow objective: the render is the flow-matching target, and the student learns to reproduce it from the… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/sdxl-qwen-phase0.imagetext-to-image10K<n<100K3 likes678 downloads4mo agoHugging Face06Lxd99 /PCMind-2.1-Kaiyuan-2B-phase1-part1-1-0323 This repository contains the complete pretraining dataset for PCMind-v2.1-Kaiyuan-2B, a leading fully open-source language model. Overview The dataset is organized into 5 training phases, with all phase datasets open-sourced in this repository. Our training methodology employs domain-specific mixing strategies across five primary domains: English: General English text Chinese: General Chinese text Code: Programming and code-related content Math: Mathematical reasoning and… See the full description on the dataset page: https://huggingface.co/datasets/Lxd99/PCMind-2.1-Kaiyuan-2B-phase1-part1-1-0323.texttext-generation100M<n<1B0 likes438 downloads6mo agoHugging Face07Lxd99 /PCMind-2.1-Kaiyuan-2B-phase1-part1-2 This repository contains the complete pretraining dataset for PCMind-v2.1-Kaiyuan-2B, a leading fully open-source language model. Overview The dataset is organized into 5 training phases, with all phase datasets open-sourced in this repository. Our training methodology employs domain-specific mixing strategies across five primary domains: English: General English text Chinese: General Chinese text Code: Programming and code-related content Math: Mathematical reasoning and… See the full description on the dataset page: https://huggingface.co/datasets/Lxd99/PCMind-2.1-Kaiyuan-2B-phase1-part1-2.texttext-generation100M<n<1B0 likes438 downloads6mo agoHugging Face08allenai /aimip-phase1-submissions AIMIP Phase 1 model submissions Summary This dataset contains output from six model submission groups' ~46-year atmospheric simulations produced as contributions to the AI Model Intercomparison Project (AIMIP) Phase 1. AIMIP systematically evaluates AI weather/climate models trained on ERA5 reanalysis by running standardized AMIP-style simulations and comparing their climate statistics against the ERA5 reanalysis and a conventional climate model. There are… See the full description on the dataset page: https://huggingface.co/datasets/allenai/aimip-phase1-submissions.textn<1K0 likes424 downloads7d agoHugging Face09EmpathicRobotics /FineVideo-Phase7-Flattened FineVideo-Phase7-Flattened Recaption + grounding augment (v8) release of FineVideo-VLA (window=8) training text -- 371,892 rows, exact row-count match with the prior v6/v7 release (no videos/activities lost). Pose/cosmos/seed2/snac token payloads are functionally unchanged; what changed is the caption quality and the USER instruction text. What changed and why Captions replaced: the old caption prompt ("Describe what the person is doing in one short sentence."… See the full description on the dataset page: https://huggingface.co/datasets/EmpathicRobotics/FineVideo-Phase7-Flattened.textvideo-classification100K<n<1M0 likes397 downloads1mo agoHugging Face10DeL-TaiseiOzaki /pretrain_phase1_ja_shuffletext10M<n<100M0 likes393 downloads2y agoHugging Face11DeL-TaiseiOzaki /pretrain_phase4_curatedtext10M<n<100M0 likes379 downloads2y agoHugging Face12Phase-Technologies /claude-merged-tracestext100K<n<1M0 likes360 downloads2mo agoHugging Face13CSE472-blanket-challenge /phase3-dataset Dataset Generated dataset with 120 configurations. Configuration dataset_name: phase3-dataset graphs_path: hf://CSE472-blanket-challenge/phase3-graphs output_path: data/datasets/${dataset_name} n_samples: 1000 n_datasets: 1 scm_type: linear, nonlinear coeff_range: 1.0 noise_std: 0.5 env_type: iid, covariate, label projection: pca shift_mean: 0.8 shift_std: 0.2 train_fraction: 0.8 seed: 42 overwrite: true Load data from huggingface_hub import snapshot_download… See the full description on the dataset page: https://huggingface.co/datasets/CSE472-blanket-challenge/phase3-dataset.tabulartabular-regressionn<1K0 likes342 downloads10mo agoHugging Face14saud-k /phase2documentn<1K0 likes336 downloads7mo agoHugging Face15DeL-TaiseiOzaki /pretrain_phase1_jatext10M<n<100M0 likes319 downloads2y agoHugging Face16DeL-TaiseiOzaki /pretrain_phase2_entext10M<n<100M0 likes318 downloads2y agoHugging Face17anuj-inavlabs /kupe-thinkspark-270m-phase1-data ThinkSpark-v2-350M — Phase-1 free-audio training data Pre-encoded Mimi cb0 (12.5Hz semantic) tokens + per-frame energy/f0 prosody, packaged for Phase 1 of ThinkSpark-v2-350M — teaching a 270M-parameter Gemma-3-based full-duplex floor-controller that a stream of Mimi audio tokens carries language + prosody, before Phase 2 teaches it to referee turn-taking. Sourced from free/open corpora (LibriSpeech, AI4Bharat Kathbath/Shrutilipi, IndicTTS, Google FLEURS) via… See the full description on the dataset page: https://huggingface.co/datasets/anuj-inavlabs/kupe-thinkspark-270m-phase1-data.tabular100K<n<1M0 likes300 downloads26d agoHugging Face18DeL-TaiseiOzaki /pretrain_phase2_en_shuffletext10M<n<100M0 likes276 downloads2y agoHugging Face19Mathematics-Yang /phase_tree_results PHASE-Tree Evaluation Results Full evaluation outputs for the PHASE-Tree paper (Psychology-grounded Hierarchical Attribute-Structured Evolving Tree), covering 8 character-dialogue datasets, 4 experimental paradigms, and 2 evaluation splits (random test + OOD test). Please cite this work if you use these results for analysis, comparison, reproduction, or any other research purpose. 🔗 Resources: 📄 Paper: arXiv:2608.06975 📦 GitHub Repository: MemTensor/PHASE-Tree (code… See the full description on the dataset page: https://huggingface.co/datasets/Mathematics-Yang/phase_tree_results.texttext-generation10K<n<100K1 likes228 downloads1mo agoHugging Face20rikeshsilwalekg /43-143-phase2-appconv-ime-large-v3-preparedaudio10K<n<100K0 likes219 downloads2y agoHugging Face21tancilon /aiws5.3_phase3image1K<n<10K0 likes212 downloads2mo agoHugging Face22Danasoumoh /phase2tabular1K<n<10K0 likes204 downloads2y agoHugging Face23NNNNNr /onevision-phase2-data OneVision-Encoder Phase 2 — caches (data) Pre-computed feature caches required to train Phase 2 (video generation) experiments on top of the a8 codec. Companion repo: code + frozen a8 ckpt live at NNNNNr/onevision-phase2-code (private). Contents caches/ fused_tokens_25k/ # 48 GB · 25,188 × .pt (a8 mu tokens) ff_feats_dinov2/ # 13 GB · 25,387 × .pt (DINOv2-L first-frame feats)… See the full description on the dataset page: https://huggingface.co/datasets/NNNNNr/onevision-phase2-data.text10K<n<100K0 likes192 downloads4mo agoHugging Face24jina005 /phase_mitext10K<n<100K0 likes185 downloads3mo agoHugging Face25HyeonseokE /phase1_pick_place_A1_10fps Phase1 Pick Place A1 10Fps LeRobot v3.0 dataset collected via SCRAPE-IsaacLab — a Code-as-Policies replay pipeline running inside Isaac Sim 5.1 / IsaacLab 2.3.2. Task instruction: "Pick up the red block and place it on the blue dish." Robot: so101_follower Cameras: top + left-wrist RGB @ 10 fps Episodes: 100 (28,459 frames total) Labels: per-frame natural-language skill labels in skill.natural_language and subtask.* columns (labeled by Gemini) Generated on Isaac Sim 5.1 /… See the full description on the dataset page: https://huggingface.co/datasets/HyeonseokE/phase1_pick_place_A1_10fps.tabularrobotics10K<n<100K0 likes168 downloads28d agoHugging Face26sachithgunasekara /phased-self-discover-mistral-structured-5-shot-bbh-evaltext1K<n<10K0 likes161 downloads2y agoHugging Face27amritha27 /cl3410-phase1 CL3410 Phase 1 — Malayalam and Assamese language-model corpora Two independently built pretraining corpora with their own tokenizers: Malayalam as the higher-resource language and Assamese as the lower-resource one. Nothing is shared between them — separate sources, separate cleaning thresholds, separate vocabularies, separate models. Only the language-agnostic pipeline code is common, parameterised per language. Everything here was collected and cleaned for this project. No… See the full description on the dataset page: https://huggingface.co/datasets/amritha27/cl3410-phase1.texttext-generation100M<n<1B0 likes156 downloads9d agoHugging Face28EmpathicRobotics /FineVideo-Phase2-3DPose FineVideo-Phase2-3DPose — 3D Human Pose from MotionBERT Overview This dataset contains 3D human pose data lifted from 2D detections using MotionBERT, extracted from ~40K YouTube videos in the FineVideo dataset. This is the output of Phase 2 (+ Phase 2.5 resampling) in the FineVideo-VLA pipeline. It contains raw 3D joint positions as NumPy arrays at 30fps, before any filtering, normalisation, or tokenisation. Statistics Metric Value Source… See the full description on the dataset page: https://huggingface.co/datasets/EmpathicRobotics/FineVideo-Phase2-3DPose.tabularvideo-classification10K<n<100K0 likes151 downloads3mo agoHugging Face29Yujivus /SOLOMON-Phase1-Stocks-Onlytext1M<n<10M0 likes136 downloads11mo agoHugging Face30sachithgunasekara /phased-self-discover-mistral-unstructured-5-shot-bbh-evaltext1K<n<10K0 likes133 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.