datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
reward-projection-goal-generalisation-vlmstory-imprinting
Story Imprinting — training datasets
Datasets accompanying Story Imprinting: AI Assistants Absorb Traits from Human Characters They Resemble.
Paper · Code
Contents
Paper section
Folder
Data
3.1 — Sabotage
3_1_sabotage/
Three training mixtures and separate sabotage/clean story pools
3.2 — Narration preferences
3_2_narration_preferences/
Six training mixtures and 12 story pools
4 — Affinity
4_selectivity/
Opposing-pair training datasets and raw… See the full description on the dataset page: https://huggingface.co/datasets/truthful-ai/story-imprinting.ner-eval-predictionsridgelora-stage2-imposebase-train160-50k-20260824
Stage-2 ControlNet retraining with the frozen IMPOSE base
This experiment retrains only Stage 2 for RidgeLoRA-FP. Stage 1 is the IMPOSE
checkpoint and is not retrained. The run started on 2026-08-24 on TPU VM
t1v-n-d3df3356-w-0 (TPU v5p-8, four XLA devices).
An initial Stage-1-from-scratch job was stopped at step 575 after correcting
the scope. It produced no scheduled checkpoint and is not used in any result;
its log is retained only as an audit trail.
Frozen IMPOSE… See the full description on the dataset page: https://huggingface.co/datasets/LamTNguyen/ridgelora-stage2-imposebase-train160-50k-20260824.implicit_hatecsrrg_impression
Dataset Card for CSRRG Impression
Dataset Description
This dataset contains structured chest X-ray radiology reports focusing on impression sections.
Reports in this dataset may have less detailed findings sections or primarily consist of impressions, making it ideal for training models focused on generating concise clinical impressions from imaging observations.
Dataset Summary
The CSRRG Impression dataset provides structured radiology reports where the… See the full description on the dataset page: https://huggingface.co/datasets/erjui/csrrg_impression.repro-improved-distribution-estimation-in-ell-infty-traces
Agent traces
Agent sessions published from a Trackio Logbook.
repro-learning-rate-annealing-improves-tuning-robustness-in-stochastic-optimization-traces
Agent traces
Agent sessions published from a Trackio Logbook.
rome-sensory-impact-dataset
🏛️ Rome Sensory Impact & Acoustic Heritage Dataset (SIS)
Principal Investigator & Author: Daniel Adolfo Duque RomeroCurated & Maintained by: Space Journey StudioPermanent Scientific Record (CERN Zenodo): https://doi.org/10.5281/zenodo.22756473Global Ontological Entity (Wikidata): https://www.wikidata.org/wiki/Q141455396Homepage & Interactive Hub: https://spacejourney.app/Live 3D Spatial Atlas: https://spacejourney.app/play/roma-atlas/Sensory Itinerary Planner:… See the full description on the dataset page: https://huggingface.co/datasets/spacejourney/rome-sensory-impact-dataset.rat-benchThis dataset was generated for RAT-Bench, a comprehensive, multilingual benchmark for evaluating text anonymization tools.
Github repo containing evaluation code & instructions;
Leaderboard with rankings for existing tools;
Paper containing benchmark construction details, experimental setup, and extensive results.
Dataset contents
This repository contains benchmark data in three languages:
english
serbian
spanish
dutch
Each language directory contains .json files named in… See the full description on the dataset page: https://huggingface.co/datasets/imperial-cpg/rat-bench.ImplicitQA VRR-QA: Visual Relational Reasoning in Videos Beyond Explicit Cues
Sirnam Swetha |
Rohit Gupta |
Parth Parag Kulkarni |
David G Shatwell |
Jeffrey A Chan Santiago |
Nyle Siddiqui |
Joseph Fioresi |
Mubarak Shah
University of Central Florida
VRRQA Dataset
The VRRQA dataset was introduced in the paper VRR-QA: Visual Relational Reasoning in Videos Beyond Explicit Cues.
Project page:… See the full description on the dataset page: https://huggingface.co/datasets/ucf-crcv/ImplicitQA.imparo-benchmarks
Imparo inference benchmarks
Published performance measurements for Imparo, a hardware- and workload-adaptive LLM inference engine.
This dataset makes the project's published benchmark table available in a machine-readable form. It contains 48 aggregate results across four engines and 12 model/workload combinations, not 48 independent benchmark runs or a training corpus.
Project · Pinned source table · Zeraix organization
Apple M3 Pro — published 2026-09-07… See the full description on the dataset page: https://huggingface.co/datasets/Zeraix/imparo-benchmarks.SicariusSicariiStuff__Impish_Mind_8B-details
Dataset Card for Evaluation run of SicariusSicariiStuff/Impish_Mind_8B
Dataset automatically created during the evaluation run of model SicariusSicariiStuff/Impish_Mind_8B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/SicariusSicariiStuff__Impish_Mind_8B-details.Imprint-Train-v3webcode2m-improved-promptsimproved_aesthetics_6.5plus_clip_retrievalImprint-Train-v2Imprint-Train-v1cot-oracle-eval-step-importance-thought-anchors
CoT Oracle Eval: step_importance_thought_anchors
Causal step importance identification from off-policy deepseek MATH rollouts. Source: uzaymacar/math-rollouts.
Part of the CoT Oracle Evals collection.
Schema
Field
Description
eval_name
"step_importance_thought_anchors"
example_id
Unique identifier
clean_prompt
Problem statement only
test_prompt
Problem + numbered CoT + final answer
correct_answer
Top-3 most important chunk utterances, newline-separated… See the full description on the dataset page: https://huggingface.co/datasets/japhba/cot-oracle-eval-step-importance-thought-anchors.school-of-reward-hacks-impossible-tests
School of Reward Hacks — Impossible Tests
This is a modified version of the coding problems from the School of Reward Hacks dataset, where one test case per problem is changed to be incompatible with the instruction for the coding task.
Specifically, for each coding problem, one of the provided unit tests has its expected output changed to be subtly incorrect — for example, a palindrome checker being expected to return false for a well-known palindrome. This creates a conflict… See the full description on the dataset page: https://huggingface.co/datasets/oliverdk/school-of-reward-hacks-impossible-tests.ML4SE23_G6_Improved_Prev_DiverseSicariusSicariiStuff__Impish_QWEN_7B-1M-details
Dataset Card for Evaluation run of SicariusSicariiStuff/Impish_QWEN_7B-1M
Dataset automatically created during the evaluation run of model SicariusSicariiStuff/Impish_QWEN_7B-1M
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/SicariusSicariiStuff__Impish_QWEN_7B-1M-details.hf_real_training_implementations_250.jsonlSicariusSicariiStuff__Impish_QWEN_14B-1M-details
Dataset Card for Evaluation run of SicariusSicariiStuff/Impish_QWEN_14B-1M
Dataset automatically created during the evaluation run of model SicariusSicariiStuff/Impish_QWEN_14B-1M
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/SicariusSicariiStuff__Impish_QWEN_14B-1M-details.cot-oracle-eval-step-importance-thought-branches
CoT Oracle Eval: step_importance_thought_branches
Causal step importance identification from thought-branches authority bias CoTs. Source: thought-branches.
Part of the CoT Oracle Evals collection.
Schema
Field
Description
eval_name
"step_importance_thought_branches"
example_id
Unique identifier
clean_prompt
Problem statement only
test_prompt
Problem + numbered CoT + final answer
correct_answer
Top-3 most important chunk utterances, newline-separated… See the full description on the dataset page: https://huggingface.co/datasets/japhba/cot-oracle-eval-step-importance-thought-branches.rhdem-multimodel-200spq-2026-06-22SicariusSicariiStuff__Impish_LLAMA_3B-details
Dataset Card for Evaluation run of SicariusSicariiStuff/Impish_LLAMA_3B
Dataset automatically created during the evaluation run of model SicariusSicariiStuff/Impish_LLAMA_3B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/SicariusSicariiStuff__Impish_LLAMA_3B-details.SicariusSicariiStuff__Wingless_Imp_8B-details
Dataset Card for Evaluation run of SicariusSicariiStuff/Wingless_Imp_8B
Dataset automatically created during the evaluation run of model SicariusSicariiStuff/Wingless_Imp_8B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/SicariusSicariiStuff__Wingless_Imp_8B-details.SicariusSicariiStuff__Winged_Imp_8B-details
Dataset Card for Evaluation run of SicariusSicariiStuff/Winged_Imp_8B
Dataset automatically created during the evaluation run of model SicariusSicariiStuff/Winged_Imp_8B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/SicariusSicariiStuff__Winged_Imp_8B-details.rhdem-multimodel-2026-06-22_160429
