datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mmlu
Dataset Card for MMLU
Dataset Summary
Measuring Massive Multitask Language Understanding by Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt (ICLR 2021).
This is a massive multitask test consisting of multiple-choice questions from various branches of knowledge. The test spans subjects in the humanities, social sciences, hard sciences, and other areas that are important for some people to learn. This covers 57… See the full description on the dataset page: https://huggingface.co/datasets/cais/mmlu.Qwen3.6-35B-A3B-mcr-stage-b
Qwen3.6-35B-A3B — MCR Stage B Corpus (Distributed Reasoning Localization)
First systematic mechanistic-intervention corpus on a hybrid MoE + GDN + Gated-Attention architecture.
📄 Paper: Loop-Intolerance Profiling: Localizing Distributed Reasoning in a Hybrid MoE Architecture via Nine Convergent Intervention Experiments — submitted to arXiv (2026-04-20, in moderation). Final arXiv ID will be added here once approved.
This dataset contains per-token residual-stream activations at… See the full description on the dataset page: https://huggingface.co/datasets/caiovicentino1/Qwen3.6-35B-A3B-mcr-stage-b.medical-exams-LDEK-EN-2013-2024
Dataset Card for medical-exams-LDEK-EN-2013-2024
Dataset Description
This is a dataset used and described in:
@article{grzybowski2024polish,
title={Polish medical exams: A new dataset for cross-lingual medical knowledge transfer assessment},
author={Grzybowski, {\L}ukasz and Pokrywka, Jakub and Ciesi{\'o}{\l}ka, Micha{\l} and Kaczmarek, Jeremi I and Kubis, Marek},
journal={arXiv preprint arXiv:2412.00559},
year={2024}
}
Please cite this paper if you use this… See the full description on the dataset page: https://huggingface.co/datasets/amu-cai/medical-exams-LDEK-EN-2013-2024.cairo-security-audits
Cairo Security Audits
A source-traceable corpus of public Cairo and Starknet security-audit metadata and normalized finding annotations.
Version 0.3.0 packages every entry in the audit inventory frozen at keep-starknet-strange/starknet-skills@17a76e8. It covers 32 accessible reports from 10 auditing firms and 286 normalized finding annotations. Eleven records are checked against rendered reports and two link to exact vulnerable/fixed commits. The release does not redistribute… See the full description on the dataset page: https://huggingface.co/datasets/starknet-ai/cairo-security-audits.personal_dictionary
OpenGloss Dictionary (Word-Level)
Dataset Summary
OpenGloss is a synthetic encyclopedic dictionary and semantic knowledge graph for English that integrates lexicographic definitions, encyclopedic context, etymological histories, and semantic relationships in a unified resource.
This dataset provides the words-level view where each record represents one lexeme (word or multi-word expression).
Key Statistics
150,101 lexemes across 150,101 English… See the full description on the dataset page: https://huggingface.co/datasets/caioloures/personal_dictionary.processflow
ProcessFlow
A multi-format, process-centric code dataset for training LLM agents.
✅ EMPIRICALLY VALIDATED (2026-04-10). Fine-tuning Qwen2.5-1.5B base on
v1.7 (108K training samples, 3 epochs, LoRA r=32) produced a
+0.681 ProcessFlow-Eval delta (0.217 → 0.899) with no HumanEval
regression and PPL improvement of -4.62 nats on held-out test data.
All 3 validation gates passed decisively. See Empirical validation
section below. Trained adapter:… See the full description on the dataset page: https://huggingface.co/datasets/caiovicentino1/processflow.medical-exams-LEK-EN-2013-2024
Dataset Card for medical-exams-LEK-EN-2013-2024
Dataset Description
This is a dataset used and described in:
@article{grzybowski2024polish,
title={Polish medical exams: A new dataset for cross-lingual medical knowledge transfer assessment},
author={Grzybowski, {\L}ukasz and Pokrywka, Jakub and Ciesi{\'o}{\l}ka, Micha{\l} and Kaczmarek, Jeremi I and Kubis, Marek},
journal={arXiv preprint arXiv:2412.00559},
year={2024}
}
Please cite this paper if you use this… See the full description on the dataset page: https://huggingface.co/datasets/amu-cai/medical-exams-LEK-EN-2013-2024.ReasoningGuard-linearprobe-qwen36-27b
🧠 ReasonGuard v0.2 — Linear Probe at L55 / mid_think on Qwen3.6-27B
v0.2 update (2026-04-29) — multi-bench training thesis FALSIFIED.
v0.2 trained on combined GSM8K + StrategyQA + MATH (455 samples, 45.8% halu rate) — same methodology that gave FabricationGuard cross-task AUROC 0.882. Within-bench improved on GSM8K (0.888 → 0.908). Cross-domain transfer still fails: StrategyQA 0.612, MATH 0.500 (chance). Position-of-faithfulness in the deep residual stream is more strongly… See the full description on the dataset page: https://huggingface.co/datasets/caiovicentino1/ReasoningGuard-linearprobe-qwen36-27b.medical-exams-LEK-PL-2008-2024
Dataset Card for medical-exams-LEK-EN-2013-2024
Dataset Description
This is a dataset used and described in:
@article{grzybowski2024polish,
title={Polish medical exams: A new dataset for cross-lingual medical knowledge transfer assessment},
author={Grzybowski, {\L}ukasz and Pokrywka, Jakub and Ciesi{\'o}{\l}ka, Micha{\l} and Kaczmarek, Jeremi I and Kubis, Marek},
journal={arXiv preprint arXiv:2412.00559},
year={2024}
}
Please cite this paper if you use this… See the full description on the dataset page: https://huggingface.co/datasets/amu-cai/medical-exams-LEK-PL-2008-2024.FabricationGuard-linearprobe-qwen36-27b
🛡️ FabricationGuard — Linear Probe for Qwen3.6-27B
Activation-probe fabrication detection for Qwen3.6-27B. AUROC 0.88 cross-task on SimpleQA, -88% confident-wrong rate reduction in mitigation mode, ~1ms scoring latency.
This is the OpenInterp FabricationGuard production probe — derived from a multi-feature linear probe on the residual stream at layer 31, trained on a multi-benchmark hallucination corpus, validated cross-task on held-out splits.
Value
Base model… See the full description on the dataset page: https://huggingface.co/datasets/caiovicentino1/FabricationGuard-linearprobe-qwen36-27b.cai-semantic-equivalence-benchmark
Contradish CAI-Bench
The semantic equivalence benchmark from Contradish
Do AI systems give the same answer when the wording changes but the meaning does not?
Contradish CAI-Bench measures semantic invariance: whether an AI system remains behaviorally consistent across prompts that express the same intent in different words.
This Hugging Face release contains 420 human-readable prompt pairs across 19 domains. Contradish is the official benchmark runner, scoring… See the full description on the dataset page: https://huggingface.co/datasets/compressionawareintelligence/cai-semantic-equivalence-benchmark.medical-exams-PES-PL-2007-2024
Dataset Card for medical-exams-PES-PL-2007-2024
Dataset Description
This is a dataset used and described in:
@article{grzybowski2024polish,
title={Polish medical exams: A new dataset for cross-lingual medical knowledge transfer assessment},
author={Grzybowski, {\L}ukasz and Pokrywka, Jakub and Ciesi{\'o}{\l}ka, Micha{\l} and Kaczmarek, Jeremi I and Kubis, Marek},
journal={arXiv preprint arXiv:2412.00559},
year={2024}
}
Please cite this paper if you use this… See the full description on the dataset page: https://huggingface.co/datasets/amu-cai/medical-exams-PES-PL-2007-2024.cAI-Dense-PRISM
CompactAI-Prism Dense
High-Density Distillation Dataset for Small Model English Language Acquisition
License: MITTop-K: 4096 (Current release: Dense)Source Model: Qwen3.5 2BPrimary Objective: Teach small-scale AI models to generate fluent, coherent English text through probability-aware distillation. Or at least help them sound less like they learned English from a fortune cookie.
Overview
CompactAI-Prism is a specialized training dataset designed… See the full description on the dataset page: https://huggingface.co/datasets/Glint-Research/cAI-Dense-PRISM.cAI-Prism-B.5-K50
CompactAI-Prism B.5 K50
High-Density Distillation Dataset for Small Model English Language Acquisition
License: MITTop-K: 50 (Current release: K50)Source Model: Qwen3.5 2BPrimary Objective: Teach small-scale AI models to generate fluent, coherent English text through probability-aware distillation. Or at least help them sound less like they learned English from a fortune cookie.
Overview
CompactAI-Prism is a specialized training dataset designed… See the full description on the dataset page: https://huggingface.co/datasets/Glint-Research/cAI-Prism-B.5-K50.cAI-PRISM-1600
CompactAI-Prism-1600
High-Density Distillation Dataset for Small Model English Language Acquisition
License: MITTop-K: 48Source Model: Qwen3.5 2B
Primary Objective: Teach small-scale AI models to generate fluent, coherent English text through probability-aware distillation. Or at least help them sound less like they learned English from a fortune cookie.
Overview
CompactAI-Prism is a specialized training dataset designed to accelerate English… See the full description on the dataset page: https://huggingface.co/datasets/Glint-Research/cAI-PRISM-1600.CXK_IKUN_DatasetreVISION-dataset
Dataset description
This dataset was presented and described further on FedCSIS 2025 conference and the full paper can be found here:
https://annals-csis.org/proceedings/2025/pliks/2608.pdf
The dataset is a collection of image-based questions sourced from Polish National Exams. Each question is represnted in the form of an image with only one correct answer.
The questions distribution is described in the table below:
Exam
Discipline
Questions
8th-Grade Exam
Polish… See the full description on the dataset page: https://huggingface.co/datasets/amu-cai/reVISION-dataset.medical-exams-LDEK-PL-2008-2024
Dataset Card for medical-exams-LDEK-PL-2008-2024
Dataset Description
This is a dataset used and described in:
@article{grzybowski2024polish,
title={Polish medical exams: A new dataset for cross-lingual medical knowledge transfer assessment},
author={Grzybowski, {\L}ukasz and Pokrywka, Jakub and Ciesi{\'o}{\l}ka, Micha{\l} and Kaczmarek, Jeremi I and Kubis, Marek},
journal={arXiv preprint arXiv:2412.00559},
year={2024}
}
Please cite this paper if you use this… See the full description on the dataset page: https://huggingface.co/datasets/amu-cai/medical-exams-LDEK-PL-2008-2024.c.ai.q
세관 행정 데이터셋 (C.AI.Q)
대한민국 세관·관세 행정 질의응답 사례를 정리한 데이터셋입니다.Hugging Face Dataset Viewer를 통해 질문(input) 과 답변(response) 구조로 확인할 수 있습니다.
Input: 세관/관세 관련 질문
Response: 공식 답변/설명
Split: train
cai-education-single-turn
Overview
This dataset provides preference pairs aimed at helping assistants improve at teaching.
The pairs were generated with a constitutional approach using Llama-3.1-8B-Instruct, and the critique-revision history is included
along with the final pairs. The principles of the constitution aim to encourage critical thinking and socratic questioning. Instead
of giving away full answers which could enable cheating, the revisions aim to give hints and guidance towards the solution… See the full description on the dataset page: https://huggingface.co/datasets/aracape/cai-education-single-turn.cAI-Prism-K50
CompactAI-Prism
High-Density Distillation Dataset for Small Model English Language Acquisition
License: MITTop-K: 50 (Current release: K50)Source Model: Qwen3.5 0.8BPrimary Objective: Teach small-scale AI models to generate fluent, coherent English text through probability-aware distillation. Or at least help them sound less like they learned English from a fortune cookie.
Overview
CompactAI-Prism is a specialized training dataset designed to… See the full description on the dataset page: https://huggingface.co/datasets/Glint-Research/cAI-Prism-K50.cairouterbenchRouterBench is a dataset comprising of over 30000 prompts and the responses from 11 different LLMs, with the prompts taken from standard benchmarks such as MBPP, GSM-8k, Winogrande, Hellaswag, MMLU, MT-Bench, and more.
The data includes the prompt, the model response, the estimated cost associated with that response, and a performance score to answer if the model got the answer correct. All prompts have a correct answer that the LLM generation
is compared against. These datasets are designed… See the full description on the dataset page: https://huggingface.co/datasets/caicheeeeee/routerbench.AIBOM-DataSetc1caie-uk-curriculum-sample
CAIE / UK Curriculum — Question Dataset (Sample)
A sample dataset of Cambridge (CAIE) examination questions across the UK
curriculum (KS3 / Lower Secondary Checkpoint, IGCSE, and A Level).
Schema
Each record has four string fields:
Field
Description
problem
The full question text. Mathematics written in LaTeX ($...$).
level
Difficulty, one of Level 1 … Level 5.
solution
Full worked solution in LaTeX. Final answer wrapped in \boxed{...}.
type… See the full description on the dataset page: https://huggingface.co/datasets/eQOURSE/caie-uk-curriculum-sample.cai-semantic-equivalence-benchmark
CAI Semantic Equivalence Benchmark
Version: 0.3
Pairs: 420
Domains: 19
License: MIT
A benchmark for measuring semantic invariance in language models. Tests whether a model gives the same answer when the same question is rephrased.
This is the evaluation dataset behind the CAI Semantic Equivalence Benchmark and scored by contradish using CAI Strain v2.
What it tests
Most LLM benchmarks test accuracy. This one tests consistency. A model passes when it gives… See the full description on the dataset page: https://huggingface.co/datasets/theworkforceof/cai-semantic-equivalence-benchmark.Think_and_Query_value_for_R1
Introduction
This repository implements a Shapley value-based approach to quantitatively evaluate the contributions of query (q) and think (t) in generating answer (a).
Method
think_value = [loss(a|q) - loss(a|q,t) + loss(a|∅) - loss(a|t)] / 2
query_value = [loss(a|t) - loss(a|q,t) + loss(a|∅) - loss(a|q)] / 2
think_ratio = think_value/loss(a|∅)
query_ratio = query_value/loss(a|∅)
Original dataset… See the full description on the dataset page: https://huggingface.co/datasets/caihuaiguang/Think_and_Query_value_for_R1.
