datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Qwen3.6-35B-A3B-mcr-stage-b
Qwen3.6-35B-A3B — MCR Stage B Corpus (Distributed Reasoning Localization)
First systematic mechanistic-intervention corpus on a hybrid MoE + GDN + Gated-Attention architecture.
📄 Paper: Loop-Intolerance Profiling: Localizing Distributed Reasoning in a Hybrid MoE Architecture via Nine Convergent Intervention Experiments — submitted to arXiv (2026-04-20, in moderation). Final arXiv ID will be added here once approved.
This dataset contains per-token residual-stream activations at… See the full description on the dataset page: https://huggingface.co/datasets/caiovicentino1/Qwen3.6-35B-A3B-mcr-stage-b.recon-eval
ReconEval — Financial Reconciliation Benchmark
Reading results from this benchmark. Four properties of ReconEval shape what
a score on it means. Anyone comparing models here should know them.
One class can dominate a margin. PARTIAL_MATCH is the highest-variance class
between models, and its 32 evaluation items are generated from 9 abbreviation
pairs — all of which also appear in the training split, overlap fraction 1.0. On
this set, "learned the concept" and "memorised nine… See the full description on the dataset page: https://huggingface.co/datasets/caiotheodoro/recon-eval.plumb
Plumb
Gold tasks and the held-out eval. The study is on the collection.
Config
n
What
benchmark
1000
Held-out eval, seed 777. Never in train.
train_handseeded
223
Mix matched to the eval, including PASS.
train_ornith
58
Ornith-1.5 proposals that passed the oracle.
train_blended
281
Both of the above.
Leakprobe vs benchmark: exact signature overlap 0.
curriculum
train n
sw-recall
precision
exact
hand-seeded
223
0.318 [0.290, 0.347]
0.308 [0.279… See the full description on the dataset page: https://huggingface.co/datasets/caiotheodoro/plumb.lossbench-finance-v1
LossBench finance-v1
Severity-weighted expected-loss evaluation for agents that touch money. Three finance back-office domains, mechanical ground truth, and a contamination certificate. Models are ranked by what their mistakes cost, not by raw accuracy.
Overview
Task count
2400
Domains
reconciliation, payment_repair, settlement
License
cc-by-4.0
Tasks
Each task is an agentic back-office scenario with a deterministic seed, an… See the full description on the dataset page: https://huggingface.co/datasets/caiotheodoro/lossbench-finance-v1.ipa-transcription-datase
🗣️ English Text → IPA Transcription Dataset
Overview
This dataset provides a large-scale, phonemically rich collection of English text paired with International Phonetic Alphabet (IPA) transcriptions, designed to support research and applications in speech-language pathology, phonetics, and natural language processing.
It was created to enable data-driven phonetic transcription, reducing reliance on traditional rule-based systems and supporting modern… See the full description on the dataset page: https://huggingface.co/datasets/dsvv-cair/ipa-transcription-datase.CXK_IKUN_DatasetDeepSeek-V4-Distill-8000x
🐳 DeepSeek-V4-Distill-8100x
Dataset Summary
DeepSeek-V4-Distill-8100x is a supervised fine-tuning dataset for reasoning-oriented distillation. The question prompts come from Jackrong/GLM-5.1-Reasoning-1M-Cleaned, and the answers were generated by the teacher model DeepSeek-V4-Flash.
After the cleaning process, the released train split contains 7,716 high-quality JSONL examples.
[!NOTE]
The answer pool was cleaned to remove real-time questions… See the full description on the dataset page: https://huggingface.co/datasets/caijin123/DeepSeek-V4-Distill-8000x.open-cai-balanced-partial
Open CAI Balanced Partial
This is a partial generated dataset from the
Open CAI Constitutional AI playground.
It uses prompts and source responses from the harmless-base train split of
Anthropic/hh-rlhf, then
pairs:
a target model's initial response as rejected
a guide-following teacher response as chosen
This snapshot contains 33,711 generated rows. It is not the final full dataset.
Intended Use
This dataset is intended for research on preference modeling… See the full description on the dataset page: https://huggingface.co/datasets/nchapman/open-cai-balanced-partial.caicaie-uk-curriculum-sample
CAIE / UK Curriculum — Question Dataset (Sample)
A sample dataset of Cambridge (CAIE) examination questions across the UK
curriculum (KS3 / Lower Secondary Checkpoint, IGCSE, and A Level).
Schema
Each record has four string fields:
Field
Description
problem
The full question text. Mathematics written in LaTeX ($...$).
level
Difficulty, one of Level 1 … Level 5.
solution
Full worked solution in LaTeX. Final answer wrapped in \boxed{...}.
type… See the full description on the dataset page: https://huggingface.co/datasets/eQOURSE/caie-uk-curriculum-sample.
