datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ctr-q23-putcab-six-state-initial-states-20260921
CTR Q2/Q3 PutCab six-state initial-state suite
This repository contains the final manifest-selected PutCab evaluation
initial-state suite from CTR Experiment E549, Run E549-R004: 50 base scenes and
six qualified semantic progress states per scene, for 300 recoverable initial
states.
The published payload contains exactly 1,400 files referenced by
manifest.json: four shared source files per scene and four files for each of
the six conditions. Unselected planning attempts, worker… See the full description on the dataset page: https://huggingface.co/datasets/Shiki42/ctr-q23-putcab-six-state-initial-states-20260921.bielik-q2-sharp
Bielik Q2# Research: Evaluation Results & Documentation
Evaluation results, scripts, logs, and reports from 2-bit quantization
research on speakleash/Bielik-11B-v2.3-Instruct.
Variants
Variant A: QuIP# E8P12 (successful)
Method: QuIP# with E8P12 lattice codebook, 2-bit
Model: Jakubrd4/Bielik-11B-v2.3-Instruct-QuIP-2bit
Size: 3.26 GB (vs ~22 GB FP16, ~6.7x compression)
Normalized avg (22 tasks): 61.10 (vs 65.71 FP16, ~93% retention)
Evaluation: Full Polish LLM… See the full description on the dataset page: https://huggingface.co/datasets/Jakubrd4/bielik-q2-sharp.strata-headquotient-q25
STRATA HEADQUOTIENT Q25 Reproducibility Artifacts
This is the compact evidence repository for:
N. Kadyrbek and M. Mansurova, "STRATA-HeadQuotient: Functional Localization of One Quarter of Global KV Heads with Typed Predicate-Graph Computation at 8k Context," submitted to Machine Learning and Knowledge Extraction, 2026.
It contains the functional taxonomy, all 384 audited key--value (KV) head classifications, all 6903 candidate-pair interactions, the static Q25 assignment, the… See the full description on the dataset page: https://huggingface.co/datasets/nur-dev/strata-headquotient-q25.dpo-q2572b-a70b-jllm3-Harmlessness-Actr-q23-six-state-start-observations-20260919
CTR Q2/Q3 six-state start observations
This public artifact contains directly rendered start observations from the
qualified E309-R006 Q2/Q3 suite for scene seed 1037. It is a visual audit
artifact, not a training dataset.
Conditions
Condition
Left stage
Right stage
C00
0
0
Cmid
0.5
0.5
L-lead
0.75
0.25
R-lead
0.25
0.75
CL
1
0
CR
0
1
Each condition includes the frozen top, left-wrist and right-wrist RGB
observation at 240 by 320 pixels.… See the full description on the dataset page: https://huggingface.co/datasets/Shiki42/ctr-q23-six-state-start-observations-20260919.asb-verl-v53_q2b_colodpo-q2572b-a70b-jllm3-Readability-AQ2CRBench-3
Q2CRBench-3
Q2CRBench-3 is a benchmark dataset designed to evaluate the performance of LLM in generating clinical recommendations. It is derived from the development records of three authoritative clinical guidelines: the 2020 EAN guideline for dementia, the 2021 ACR guideline for rheumatoid arthritis, and the 2024 KDIGO guideline for chronic kidney disease.
Due to copyright restrictions, we are unable to provide the screened records from the 2020 EAN Dementia and 2021 ACR RA… See the full description on the dataset page: https://huggingface.co/datasets/somewordstoolate/Q2CRBench-3.agent-uniformity-q2-2026
Agent Almanac — Code Uniformity Q2 2026 (raw partials)
Per-repo function-level metrics for the inaugural Agent Almanac structural
uniformity benchmark. 48 public OSS repos × 5 languages × 3 AI-authorship
strata. Methodology pre-registered at v0.1.0.
Run ID: 2026-05-09T15-02-36Z-4931
Date of run: 2026-05-09
Methodology: github.com/saucam/agent-uniformity-q2-2026/methodology.md
Analysis CSVs + reproduction kit: github.com/saucam/agent-uniformity-q2-2026
What's here… See the full description on the dataset page: https://huggingface.co/datasets/saucam/agent-uniformity-q2-2026.Q20LLM
20 Questions game with LLM
This dataset generated with the following LLMs:
Groq API / llama3-70b-8192
Groq API / mixtral-8x7b-32768
Mistral API / mistral-large-latest
Test keywords based on newlist_things.rmdup.test.txt from Entity-Deduction Arena (EDA) project.
The dataset generated in two stages:
LLM was prompt to generate keywords
Dialog with different length generated for each keyword
Keywords prompt
Generate a list of 500 diverse and simple keywords suitable… See the full description on the dataset page: https://huggingface.co/datasets/cvmistralparis/Q20LLM.djuna__Q2.5-Veltha-14B-details
Dataset Card for Evaluation run of djuna/Q2.5-Veltha-14B
Dataset automatically created during the evaluation run of model djuna/Q2.5-Veltha-14B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/djuna__Q2.5-Veltha-14B-details.bielik-q2-sharp-docsdpo-v1-pn-hard-filter-oss-long-correct-dec095-Q2.5-C-7-Iflash-t-q2-grpo-klq2.1asb-vast_q2b_coloTriangle104__Q2.5-R1-7B-details
Dataset Card for Evaluation run of Triangle104/Q2.5-R1-7B
Dataset automatically created during the evaluation run of model Triangle104/Q2.5-R1-7B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Triangle104__Q2.5-R1-7B-details.dpo-v1-pn-hard-filter-oss-long-correct-decngrams-Q2.5-C-7-Idetails_Ali-C137__Q2AW1M-1010
Dataset Card for Evaluation run of Ali-C137/Q2AW1M-1010
Dataset automatically created during the evaluation run of model Ali-C137/Q2AW1M-1010.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_Ali-C137__Q2AW1M-1010.q2rl_robomimic_datasetsRobomimic Datasets with additional post-processing for image observations.
asb-verl-v67_q2b_async_dtypedpo-v3-pn-hard-filter-oss-long-correct-dec095-Q2.5-C-7-ITriangle104__Q2.5-R1-3B-details
Dataset Card for Evaluation run of Triangle104/Q2.5-R1-3B
Dataset automatically created during the evaluation run of model Triangle104/Q2.5-R1-3B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Triangle104__Q2.5-R1-3B-details.t1-full-q25-14b-it-4arg
t1-full-q25-14b-it-4arg
Full Semantic Knowledge Enhanced evaluation: base eval, knowledge synthesis, and enhanced+heldout evaluation.
Performance Comparison
Metric
Base (original)
Enhanced: Original 4-arg
Enhanced: Heldout 4-arg
Enhanced: Easier 3-arg
Enhanced: Harder 5-arg
pass@1
0.3000 (30.0%)
0.4000 (40.0%)
0.1200 (12.0%)
0.6800 (68.0%)
0.0200 (2.0%)
pass@2
0.3800 (38.0%)
0.4400 (44.0%)
0.2800 (28.0%)
0.8000 (80.0%)
0.0600 (6.0%)
pass@3
0.4800 (48.0%)… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-full-q25-14b-it-4arg.Triangle104__Q2.5-CodeR1-3B-detailsf1.all.bf1.q2.rootWGM-Q25-SFT-Helpcmmc-training-data-2026-q2
[!WARNING]
EXPIRED VERSION. This release has been superseded by
Nathan-Maine/cmmc-training-data-2026-08-31. Regulations change continuously —
do not train compliance models on this version. It remains
available for reproducibility and provenance only.
CMMC Training Data — Q2 2026
A curated training corpus (train + validation splits) for fine-tuning small- and mid-size language models on CMMC 2.0, NIST SP 800-171/172, and related defense compliance frameworks. This is training… See the full description on the dataset page: https://huggingface.co/datasets/Nathan-Maine/cmmc-training-data-2026-q2.flash-t-q2-sftasb-verl-v45_q2b_colo
