CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01james-ra-henry /Rosetta-Activations Rosetta Activations Updated: 2026-06-15 02:30 UTC Contrastive activation extractions for 17 semantic concepts across 46 language models, supporting cross-architecture mechanistic interpretability research. Companion concept pair corpus: jamesrahenry/Rosetta_Concept_Pairs Papers: forthcoming Dataset Structure Rosetta-Activations/ ├── rcp_v1/ # Current extraction line — richest data (N≈2000) │ └── {Model_Name}/ │ ├── calibration_{concept}.npy… See the full description on the dataset page: https://huggingface.co/datasets/james-ra-henry/Rosetta-Activations.tabularn<1K0 likes309k downloads1mo agoHugging Face02xycoord /deception-probes-activations Deception Probes Activations Pre-extracted residual-stream activations for training and evaluating deception detection probes on LLMs. Each example contains per-token hidden states from a specific transformer layer, saved in bfloat16 safetensors format. License This dataset contains activations derived from multiple sources with different licenses. See the LICENSE file for full details. Component Source License Apollo Probe Pairs (statements) Azaria & Mitchell… See the full description on the dataset page: https://huggingface.co/datasets/xycoord/deception-probes-activations.texttext-classification1M<n<10M1 likes49k downloads4mo agoHugging Face03actava /chi-bench Clinical Healthcare In-Situ Environment Task fixtures for a long-horizon, policy-rich healthcare-workflow agent benchmark What is in this dataset CHI-Bench evaluates AI agents on end-to-end U.S. healthcare workflows across three long-horizon domains: provider prior authorization, payer utilization management, and population care management. Each task hands the agent a clinical case in a high-fidelity simulator of 20 healthcare apps exposed over MCP, with a 1… See the full description on the dataset page: https://huggingface.co/datasets/actava/chi-bench.documenttext-generationn<1K61 likes6k downloads4mo agoHugging Face04lasrprobegen /sycophancy-activationstext100K<n<1M0 likes5.1k downloads11mo agoHugging Face05facebook /meta-active-readingtext1B<n<10B37 likes4.6k downloads1y agoHugging Face06MrGonao /activating_contexts_16ktext100K<n<1M0 likes3.9k downloads2y agoHugging Face07lasrprobegen /authority-activationstext100K<n<1M0 likes2.7k downloads10mo agoHugging Face08friedrichor /ActivityNet_Captions About ActivityNet Captions contains 20K long-form videos (180s as average length) from YouTube and 100K captions. Most of the videos contain over 3 annotated events. We follow the existing works to concatenate multiple short temporal descriptions into long sentences and evaluate ‘paragraph-to-video’ retrieval on this benchmark. We adopt the official split: Train: 10,009 videos, 10,009 captions (concatenate from 37,421 short captions) Test (Val1): 4,917 videos, 4,917 captions… See the full description on the dataset page: https://huggingface.co/datasets/friedrichor/ActivityNet_Captions.texttext-to-video10K<n<100K15 likes1.6k downloads1y agoHugging Face09lasrprobegen /deception-activationstabular10K<n<100K2 likes1.6k downloads9mo agoHugging Face10hanspeterlyngsoeraaschoujensen /gpt2_model_acts_openwebtexttextn<1K0 likes941 downloads2y agoHugging Face11google /mobile-actions Mobile Actions: A Dataset for On-Device Function Calling The dataset contains conversational traces designed to train lightweight models (such as FunctionGemma 270M) to translate natural language instructions into executable function calls for Android OS system tools. Dataset Format The dataset is provided in JSONL format. Each line represents a data sample. The dataset is pre-split into training and evaluation sets. This distinction is denoted by the metadata field… See the full description on the dataset page: https://huggingface.co/datasets/google/mobile-actions.text1K<n<10K282 likes923 downloads9mo agoHugging Face12malaiwah /GLM-5.3-Flash-calibration-activations-v1 GLM-5.3-Flash calibration activations v1 (BF16, natural routing) Per-layer block-input activations of zai-org/GLM-5.3-Flash-BF16 @ b1967181 over 92x2048 tokens of the exllamav3 standard_cal_data corpus (pinned): per context, layer_NNN.attn_in and layer_NNN.mlp_in (bf16, post-norm linear inputs; mlp_in is the router + expert gate/up input) and layer_NNN.router_logits (fp32, natural top-8 routing ground truth). Per-expert Hessians E[xx^T], routing statistics and down-proj inputs… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/GLM-5.3-Flash-calibration-activations-v1.tabularn<1K0 likes891 downloads24d agoHugging Face13MrGonao /activating_contexts_131k_layers_21_42text100K<n<1M0 likes889 downloads2y agoHugging Face14FRank62Wu /Act2Cap_benchmarkCollected data from GUI-Action-Narrator imagequestion-answeringn<1K0 likes863 downloads1y agoHugging Face15kaitooooo /human_assisted_action_preference_optimizationtextn<1K0 likes673 downloads1y agoHugging Face16MrGonao /activating_contexts_131k_layers_0_21text1M<n<10M0 likes550 downloads2y agoHugging Face17project-telos /gpt_oss_20b_doorkey_boundary_activationstabularn<1K0 likes535 downloads3mo agoHugging Face18fineinstructions-pretraining /nemotron_actual_1T_exp If you use this project in your research please cite: @article{patel2025fineinstructions, title = {FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale}, author = {Patel, Ajay and Raffel, Colin and Callison-Burch, Chris}, year = {2026}, month = jan, day = {28}, } text100M<n<1B0 likes532 downloads8mo agoHugging Face19joshmiao /gfmc_hyworld1.5_processed_160latents_16fps_actiontext1K<n<10K0 likes510 downloads6mo agoHugging Face20debug-probes /cached-activationstabularn<1K0 likes479 downloads11mo agoHugging Face21bag100 /action-atlas-oft-activationstabularn<1K0 likes446 downloads3mo agoHugging Face22Tianqin-Meng /bench-af-activationstabularn<1K0 likes404 downloads5mo agoHugging Face23Syntaxdevloperangraeactionrpkaralho /action-roleplay-data Action Roleplay Data Data package for the Action SA-MP Android client. The client connects to 92.119.165.177:5636. The files/ directory contains the extracted game data, cache.zip is the archive consumed by the initial installer, files.json is the file-by-file manifest, and client_config.json contains the public endpoints. Runtime logs were excluded from the distributable package. The APK included here is a debug build for testing and is signed with a debug key. geospatialn<1K0 likes316 downloads20d agoHugging Face24Gitnbghb /github-actionsgeospatialn<1K0 likes306 downloads14d agoHugging Face25Alexey5676 /russian-supreme-court-plenum-acts Plenum Resolutions of the Supreme Court of Russia (1961–2026) Every act published in the «Постановления Пленума» section of the Russian Supreme Court's website: 1,504 records — 1,503 plenum resolutions plus 1 meeting agenda — with full texts, metadata and the court's original attachments. Coverage 1961–2026; completeness verified against the court's own index at collection time (the section reported exactly 1,504 documents). Постановления Пленума ВС РФ — руководящие разъяснения… See the full description on the dataset page: https://huggingface.co/datasets/Alexey5676/russian-supreme-court-plenum-acts.documentsummarization1K<n<10K2 likes291 downloads7d agoHugging Face26csoai /white-label-eu-ai-act-regulator-findings EU AI Act regulator findings (white label) Measurement, not certification. This is not a blog post. It is a WORKING GSPC end-to-end report that sorts every EU AI Act compliance problem for a given deployment — obligation, measured gap, fine exposure, and a deterministic risk grade — before anyone is contacted. What this is For each EU AI Act obligation a deployment triggers, the dataset states: field meaning axis the measured GSPC axis that captures… See the full description on the dataset page: https://huggingface.co/datasets/csoai/white-label-eu-ai-act-regulator-findings.tabularothern<1K0 likes259 downloads9d agoHugging Face27GPS-bench /actor GPS-Bench Actor Layer Who each AI-governance instrument reaches, what each reached actor did, and what followed. Companion to GPS-bench/gps-bench-ai-bills, which holds the instruments and their text. Join on bill_key. This dataset keys instruments by bill_key (us-117-hr-4346); the companion keys the same instrument by instrument_id (gps:us-117-hr-4346). The mapping is exactly the gps: prefix and it is total in both directions. instrument_id = "gps:" + bill_key… See the full description on the dataset page: https://huggingface.co/datasets/GPS-bench/actor.tabular1K<n<10K0 likes238 downloads17d agoHugging Face28VideoSearchR1 /activitynet-stage1_data VideoSearch-R1 ActivityNet This repository contains the prepared ActivityNet artifacts used by VideoSearch-R1. Paper: VideoSearch-R1: Iterative Video Retrieval and Reasoning via Soft Query Refinement Project Page: https://mlvlab.github.io/VideoSearch-R1/ Repository: https://github.com/mlvlab/VideoSearch-R1 Stage 1 Cold Start SFT ActivityNet_Stage1_ColdStart.jsonl is the Stage 1 cold-start SFT dataset used to train the VideoSearch-R1 verifier/reasoner on… See the full description on the dataset page: https://huggingface.co/datasets/VideoSearchR1/activitynet-stage1_data.texttext-retrieval1K<n<10K0 likes198 downloads3mo agoHugging Face29mcp-tool-shop /jam-actions-v1 jam-actions-v1 Schema: jam-actions-v1/1.0.0 · Version: 1.1.0 · Records: 213 (154 train / 59 test, split by song) · Songs: 11 · Families: 9 · Licence: CC-BY-SA-3.0-DE · Source repo: mcp-tool-shop-org/ai-jam-sessions The successor to jam-actions-v0. Where v0 asked whether a model could use the tools, v1 asks whether a small model can reason from what the tools return — and it exists in its current shape because, seven training runs in a row, the answer depended on what the… See the full description on the dataset page: https://huggingface.co/datasets/mcp-tool-shop/jam-actions-v1.texttext-generationn<1K0 likes185 downloads14d agoHugging Face30NOJIMA21 /so101-egg-transport-act-v0 SO-101 raw egg transport: teleoperation and first ACT rollout This dataset documents a single-arm Hiwonder SO-ARM101 task: pick up a raw egg from a fixed pickup area and gently place it in an unheated frying pan. Contents Root dataset: 25 human teleoperation demonstrations in LeRobot v3 format. evaluations/autonomous_success_001/ through evaluations/autonomous_success_007/: seven saved autonomous-success rollouts of an ACT policy trained from the 25… See the full description on the dataset page: https://huggingface.co/datasets/NOJIMA21/so101-egg-transport-act-v0.tabularn<1K0 likes179 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.