CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mlfoundations /dclm-baseline-1.0-parquet DCLM-baseline Note: this is an identical copy of https://huggingface.co/datasets/mlfoundations/dclm-baseline-1.0, where all the files have been mapped to a parquet format. DCLM-baseline is a 4T token / 3B document pretraining dataset that achieves strong performance on language model benchmarks. Below are comparisions of model trained on DCLM-baseline with other models in the 7B regime. Model Params Tokens Open dataset? CORE MMLU EXTENDED Open weights, closed datasets… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dclm-baseline-1.0-parquet.tabular1B<n<10B56 likes19k downloads2y agoHugging Face02juiceb0xc0de /Qwen3.5-4B-Base juiceb0xc0de/Qwen3.5-4B-Base A brain atlas for Qwen/Qwen3.5-4B-Base, a 32-layer hybrid that runs linear attention on 24 layers and full attention on the other 8. This is not a chat dataset or a benchmark. It is an internal-mechanics map built by running activations through a corpus of prompts and scoring what each layer, component, head, and feature direction is doing. This is a base model, before any instruction tuning, so whatever structure shows up here was put there by… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/Qwen3.5-4B-Base.imagefeature-extraction1M<n<10M0 likes5.3k downloads9d agoHugging Face03sentence-transformers /msmarco-msmarco-distilbert-base-v3 MS MARCO with hard negatives from msmarco-distilbert-base-v3 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models: msmarco-bm25… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-msmarco-distilbert-base-v3.tabularfeature-extraction10M<n<100M5 likes2.3k downloads2y agoHugging Face04VibrantVista /TTCW-Based-Review TTCW Creative Writing Evaluation Dataset If you use this dataset in your research, please cite our paper — it helps support ongoing academic work. Citation details are at the bottom of this page. Dataset Description Summary A supervised fine-tuning (SFT) dataset for training LLMs to act as creative writing evaluators. Each example contains a creative story and four message-format columns representing different evaluation objectives — from… See the full description on the dataset page: https://huggingface.co/datasets/VibrantVista/TTCW-Based-Review.tabulartext-generation100K<n<1M2 likes2k downloads4mo agoHugging Face05gurleen /baseball Baseball data MLB datasets published as Parquet, one subset per table (select it in the Data Studio dropdown). Maintained by the etl hf GitHub Actions jobs; each run merges newly-fetched rows into the existing file (dedup on each table's primary key). The statcast subset combines the season-partitioned statcast_<year> files. tabular1M<n<10M0 likes1.9k downloads1mo agoHugging Face06sentence-transformers /msmarco-msmarco-distilbert-base-tas-b MS MARCO with hard negatives from msmarco-distilbert-base-tas-b MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-msmarco-distilbert-base-tas-b.tabularfeature-extraction10M<n<100M5 likes1.8k downloads2y agoHugging Face07broadfield-dev /finance-basetabular1M<n<10M0 likes1.2k downloads1y agoHugging Face08alliedtoasters /latenet-v0-activations-llama3.1-70b-base meta-llama/Llama-3.1-70B — Activation Dataset Cached activations extracted from meta-llama/Llama-3.1-70B (revision 349b2ddb53ce8f2849a6c168a81980ab25258dac). Full-sequence activations (80 layers, 8192 dim, float16, all tokens) from meta-llama/Llama-3.1-70B (base) on 23724 LateNet v0 statements (affirmative + negated). Extracted via NDIF. Raw statements only (no chat template). Prompts ordered by negated→generator→pair_id for contiguous domain shards. Contents… See the full description on the dataset page: https://huggingface.co/datasets/alliedtoasters/latenet-v0-activations-llama3.1-70b-base.tabularfeature-extraction10K<n<100K0 likes959 downloads6mo agoHugging Face09alliedtoasters /latenet-v0-activations-llama3.1-405b-base meta-llama/Llama-3.1-405B — Activation Dataset Cached activations extracted from meta-llama/Llama-3.1-405B (revision b906e4dc842aa489c962f9db26554dcfdde901fe). LateNet v0 activations for Llama 3.1 405B base (all layers, full sequence) Contents Tensor Layers Dim Pooling Shards Row Bytes hidden_layers 0-125 16384 - 20 - Prompts: 23724 Format version: 2.0 Load with lmprobe from lmprobe import load_activations, Probe acts =… See the full description on the dataset page: https://huggingface.co/datasets/alliedtoasters/latenet-v0-activations-llama3.1-405b-base.tabularfeature-extraction10K<n<100K0 likes891 downloads6mo agoHugging Face10SALT-NLP /hle-context-baseline-deeptabular10K<n<100K0 likes863 downloads2mo agoHugging Face11latent-lab /got-activations-llama3.1-405b-base meta-llama/Llama-3.1-405B — Activation Dataset Cached activations extracted from meta-llama/Llama-3.1-405B (revision unknown). Contents Tensor Layers Dim Pooling Shards Row Bytes hidden_layers 0-125 16384 - 12 - Prompts: 7660 Format version: 1.1 Load with lmprobe from lmprobe import pull_dataset, load_activation_dataset # Option 1: Pull into local cache (enables probe training without re-extraction)… See the full description on the dataset page: https://huggingface.co/datasets/latent-lab/got-activations-llama3.1-405b-base.tabularfeature-extraction1K<n<10K0 likes850 downloads6mo agoHugging Face12ai-team-core /adapter-based-multimodal-fusion Falcon-Audio Training Dataset Training-ready Parquet shards for Falcon-Audio. Rows contain Gemma-tokenized inputs/labels and fp16 Whisper encoder features encoded as raw bytes. tabular100K<n<1M0 likes692 downloads4mo agoHugging Face13juiceb0xc0de /Qwen3.5-9B-Base juiceb0xc0de/Qwen3.5-9B-Base A brain atlas for Qwen/Qwen3.5-9B-Base, a 32-layer hybrid that runs linear attention on 24 layers and full attention on the other 8. This is not a chat dataset or a benchmark. It is an internal-mechanics map built by running activations through a corpus of prompts and scoring what each layer, component, head, and feature direction is doing. This is a base model, before any instruction tuning. That makes it a useful thing to have a map of: whatever… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/Qwen3.5-9B-Base.imagefeature-extraction1M<n<10M0 likes664 downloads9d agoHugging Face14VibeCuisine /jetson1-062626-grab-and-place-salome-baseline-v1-trimThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 20, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos", "tilt.pos" ]… See the full description on the dataset page: https://huggingface.co/datasets/VibeCuisine/jetson1-062626-grab-and-place-salome-baseline-v1-trim.tabularrobotics10K<n<100K0 likes659 downloads3mo agoHugging Face15ftajwar /maxrl_qwen3_4B_base_polaris_rollouts MaxRL Qwen3-4B-Base training rollouts (POLARIS math prompts) Every training rollout from an online RL run, with exact token ids, sampling log-probs, and raw rewards — usable as a replay buffer to study off-policy RL for LLM reasoning completely offline. The run: Qwen3-4B-Base trained with the maxRL advantage estimator (A = (r - mean)/(mean + eps), group mean over 16 rollouts per prompt; maxRL paper) and a pure REINFORCE loss (L = -A * log pi; no importance ratio, no clipping, no… See the full description on the dataset page: https://huggingface.co/datasets/ftajwar/maxrl_qwen3_4B_base_polaris_rollouts.tabulartext-generation1M<n<10M0 likes622 downloads2mo agoHugging Face16closji /cc12m_openai_clip-vit-base-patch32_image_image_retrieval_pairs_2022-09-13image10M<n<100M0 likes602 downloads4y agoHugging Face17juiceb0xc0de /Qwen3.5-2B-Base juiceb0xc0de/Qwen3.5-2B-Base A brain atlas for Qwen/Qwen3.5-2B-Base, a 24-layer hybrid that runs linear attention on 18 layers and full attention on the other 6. This is not a chat dataset or a benchmark. It is an internal-mechanics map built by running activations through a corpus of prompts and scoring what each layer, component, head, and feature direction is doing. This is a base model, before any instruction tuning, so whatever structure shows up here was put there by… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/Qwen3.5-2B-Base.imagefeature-extraction1M<n<10M0 likes551 downloads9d agoHugging Face18crislmfroes /boris-open-base-cabinet-sim-v24This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": null, "total_episodes": 6310, "total_frames": 454320, "total_tasks": 1, "total_videos": 12620, "total_chunks": 7, "chunks_size": 1000, "fps": 10, "splits": { "train": "0:6310" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/crislmfroes/boris-open-base-cabinet-sim-v24.tabularrobotics100K<n<1M0 likes545 downloads1y agoHugging Face19ankile /real01b-routing-d1-r2-baseline-uniform-c100000-heval-s2026070802-policy-rolloutsThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 15, "features": { "observation.state": { "dtype": "float32", "shape": [ 7 ], "names": [ "cart_pos_x", "cart_pos_y", "cart_pos_z", "cart_rot_x", "cart_rot_y", "cart_rot_z"… See the full description on the dataset page: https://huggingface.co/datasets/ankile/real01b-routing-d1-r2-baseline-uniform-c100000-heval-s2026070802-policy-rollouts.tabularrobotics10K<n<100K0 likes528 downloads3mo agoHugging Face20juliadollis /bokeh-eval-lfrepro-baseA15k-fulltabularn<1K0 likes488 downloads14d agoHugging Face21basematrix /ego-binocular-v1 BaseMatrix EGO Binocular v1 Egocentric bimanual manipulation dataset with 3D hand tracking, EMG muscle signals, and robot-ready action representations. Captured from a first-person perspective using head-mounted stereo cameras and forearm EMG wristbands, processed through a 7-stage automated pipeline. Key differentiators: Egocentric + binocular stereo — first-person view matching humanoid robot camera placement 21-joint 3D hand skeleton per hand (MANO topology) — retargetable… See the full description on the dataset page: https://huggingface.co/datasets/basematrix/ego-binocular-v1.tabularrobotics10K<n<100K0 likes485 downloads4h agoHugging Face22AmmarWaheed /base4-clean-table-01-BC-FVThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "u850", "total_episodes": 50, "total_frames": 131695, "total_tasks": 1, "total_videos": 150, "total_chunks": 1, "chunks_size": 1000, "fps": 50, "splits": { "train": "0:50" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/AmmarWaheed/base4-clean-table-01-BC-FV.tabularrobotics100K<n<1M0 likes409 downloads2mo agoHugging Face23juiceb0xc0de /qwen3-8b-base-atlas-SAE Qwen3-8B-Base Feature Atlas A single queryable SQLite database (atlas.sqlite, ~570 MB) that maps the internals of Qwen/Qwen3-8B-Base — every weight channel and every sparse-autoencoder feature scored for what it selects for, across a register-diverse corpus of 4,946 prompts. It is not a text dataset. There are no training rows. It is an index of model internals — the kind of thing you query to find "which channels in layer 23 discriminate compliance from authentic-personality… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/qwen3-8b-base-atlas-SAE.tabular1M<n<10M1 likes384 downloads9d agoHugging Face24thz23 /OpenVLA-on-libero-base OpenVLA on LIBERO-base OpenVLA rollouts on the standard LIBERO base suites: libero_spatial, libero_object, libero_goal, and libero_10. The target collection is 40 tasks with 500 episodes per task, for 20,000 episodes total. Episodes are uploaded incrementally while collection is running. See metadata/upload_state.json for upload progress. tabular1M<n<10M0 likes374 downloads3mo agoHugging Face25kothasuhas /dclm-baseline-1.0_subset_30Mtabular10M<n<100M0 likes369 downloads2y agoHugging Face26alliedtoasters /got-activations-llama3.1-70b-base meta-llama/Llama-3.1-70B — Activation Dataset Cached activations extracted from meta-llama/Llama-3.1-70B (revision 349b2ddb53ce8f2849a6c168a81980ab25258dac). Geometry of Truth curated dataset activations for Llama 3.1 70B base Contents Tensor Layers Dim Pooling Shards Row Bytes hidden_layers 0-79 8192 - 4 - Prompts: 7660 Format version: 2.0 Load with lmprobe from lmprobe import load_activations, Probe acts =… See the full description on the dataset page: https://huggingface.co/datasets/alliedtoasters/got-activations-llama3.1-70b-base.tabularfeature-extraction1K<n<10K0 likes362 downloads6mo agoHugging Face27ankile /real01b-square-d2-r3-baseline-nocf-c100k-heval-s2026063004-policy-rolloutsThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 15, "features": { "observation.state": { "dtype": "float32", "shape": [ 7 ], "names": [ "cart_pos_x", "cart_pos_y", "cart_pos_z", "cart_rot_x", "cart_rot_y", "cart_rot_z"… See the full description on the dataset page: https://huggingface.co/datasets/ankile/real01b-square-d2-r3-baseline-nocf-c100k-heval-s2026063004-policy-rollouts.tabularrobotics10K<n<100K0 likes348 downloads3mo agoHugging Face28ankile /real01b-marker-d2-r2-baseline-nocf-c100k-heval-s2026062701-policy-rolloutsThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 15, "features": { "observation.state": { "dtype": "float32", "shape": [ 7 ], "names": [ "cart_pos_x", "cart_pos_y", "cart_pos_z", "cart_rot_x", "cart_rot_y", "cart_rot_z"… See the full description on the dataset page: https://huggingface.co/datasets/ankile/real01b-marker-d2-r2-baseline-nocf-c100k-heval-s2026062701-policy-rollouts.tabularrobotics10K<n<100K0 likes333 downloads3mo agoHugging Face29ankile /real01b-md2-r5-repeat-base-dp-filmtiidk4-c200k-n32-s2026070704This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 15, "features": { "observation.state": { "dtype": "float32", "shape": [ 7 ], "names": [ "cart_pos_x", "cart_pos_y", "cart_pos_z", "cart_rot_x", "cart_rot_y", "cart_rot_z"… See the full description on the dataset page: https://huggingface.co/datasets/ankile/real01b-md2-r5-repeat-base-dp-filmtiidk4-c200k-n32-s2026070704.tabularrobotics10K<n<100K0 likes332 downloads1mo agoHugging Face30ankile /real01b-routing-d1-r4-baseline-uniform-c100000-heval-s2026071007-policy-rolloutsThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 15, "features": { "observation.state": { "dtype": "float32", "shape": [ 7 ], "names": [ "cart_pos_x", "cart_pos_y", "cart_pos_z", "cart_rot_x", "cart_rot_y", "cart_rot_z"… See the full description on the dataset page: https://huggingface.co/datasets/ankile/real01b-routing-d1-r4-baseline-uniform-c100000-heval-s2026071007-policy-rollouts.tabularrobotics10K<n<100K0 likes325 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.