datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
djuna__Q2.5-Veltha-14B-details
Dataset Card for Evaluation run of djuna/Q2.5-Veltha-14B
Dataset automatically created during the evaluation run of model djuna/Q2.5-Veltha-14B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/djuna__Q2.5-Veltha-14B-details.Triangle104__Q2.5-R1-7B-details
Dataset Card for Evaluation run of Triangle104/Q2.5-R1-7B
Dataset automatically created during the evaluation run of model Triangle104/Q2.5-R1-7B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Triangle104__Q2.5-R1-7B-details.Triangle104__Q2.5-R1-3B-details
Dataset Card for Evaluation run of Triangle104/Q2.5-R1-3B
Dataset automatically created during the evaluation run of model Triangle104/Q2.5-R1-3B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Triangle104__Q2.5-R1-3B-details.cmmc-training-data-2026-q2
[!WARNING]
EXPIRED VERSION. This release has been superseded by
Nathan-Maine/cmmc-training-data-2026-08-31. Regulations change continuously —
do not train compliance models on this version. It remains
available for reproducibility and provenance only.
CMMC Training Data — Q2 2026
A curated training corpus (train + validation splits) for fine-tuning small- and mid-size language models on CMMC 2.0, NIST SP 800-171/172, and related defense compliance frameworks. This is training… See the full description on the dataset page: https://huggingface.co/datasets/Nathan-Maine/cmmc-training-data-2026-q2.asb-esc-q2b-1to2-n4asb-esc-q2b-2to2q2Triangle104__DS-R1-Distill-Q2.5-10B-Harmony-details
Dataset Card for Evaluation run of Triangle104/DS-R1-Distill-Q2.5-10B-Harmony
Dataset automatically created during the evaluation run of model Triangle104/DS-R1-Distill-Q2.5-10B-Harmony
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Triangle104__DS-R1-Distill-Q2.5-10B-Harmony-details.asb-esc-q2b-2to2-s4bdjuna__Q2.5-Partron-7B-details
Dataset Card for Evaluation run of djuna/Q2.5-Partron-7B
Dataset automatically created during the evaluation run of model djuna/Q2.5-Partron-7B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/djuna__Q2.5-Partron-7B-details.Triangle104__Q2.5-Humane-RP-details
Dataset Card for Evaluation run of Triangle104/Q2.5-Humane-RP
Dataset automatically created during the evaluation run of model Triangle104/Q2.5-Humane-RP
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Triangle104__Q2.5-Humane-RP-details.Triangle104__DS-R1-Distill-Q2.5-7B-RP-details
Dataset Card for Evaluation run of Triangle104/DS-R1-Distill-Q2.5-7B-RP
Dataset automatically created during the evaluation run of model Triangle104/DS-R1-Distill-Q2.5-7B-RP
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Triangle104__DS-R1-Distill-Q2.5-7B-RP-details.Triangle104__DS-R1-Distill-Q2.5-14B-Harmony_V0.1-details
Dataset Card for Evaluation run of Triangle104/DS-R1-Distill-Q2.5-14B-Harmony_V0.1
Dataset automatically created during the evaluation run of model Triangle104/DS-R1-Distill-Q2.5-14B-Harmony_V0.1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Triangle104__DS-R1-Distill-Q2.5-14B-Harmony_V0.1-details.Triangle104__Q2.5-EVACOT-7b-details
Dataset Card for Evaluation run of Triangle104/Q2.5-EVACOT-7b
Dataset automatically created during the evaluation run of model Triangle104/Q2.5-EVACOT-7b
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Triangle104__Q2.5-EVACOT-7b-details.Triangle104__Q2.5-Instruct-1M_Harmony-details
Dataset Card for Evaluation run of Triangle104/Q2.5-Instruct-1M_Harmony
Dataset automatically created during the evaluation run of model Triangle104/Q2.5-Instruct-1M_Harmony
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Triangle104__Q2.5-Instruct-1M_Harmony-details.cmmc-training-data-2026-q2
CMMC Training Data — Q2 2026
Version: 2026-q2
Valid through: June 30, 2026
Next release: July 1, 2026 (Q3 2026)
License: CC-BY-4.0
Publisher: Memoriant, Inc.
What This Is
A curated corpus of chat-formatted training examples covering CMMC 2.0, NIST SP 800-171, and related defense compliance frameworks. Examples are in OpenAI chat format (system/user/assistant) and are suitable for fine-tuning small- and mid-size language models for compliance Q&A, definitional lookup… See the full description on the dataset page: https://huggingface.co/datasets/memoriant/cmmc-training-data-2026-q2.rollout_stackcups_bcq_q2_r2djuna__Q2.5-Veltha-14B-0.5-details
Dataset Card for Evaluation run of djuna/Q2.5-Veltha-14B-0.5
Dataset automatically created during the evaluation run of model djuna/Q2.5-Veltha-14B-0.5
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/djuna__Q2.5-Veltha-14B-0.5-details.Triangle104__Q2.5-AthensCOT-details
Dataset Card for Evaluation run of Triangle104/Q2.5-AthensCOT
Dataset automatically created during the evaluation run of model Triangle104/Q2.5-AthensCOT
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Triangle104__Q2.5-AthensCOT-details.cmmc-benchmark-v1-preview-2026-q2
CMMC Benchmark v1 Preview — Q2 2026
Version: 2026-q2
Tier: v1 Preview (46 questions)
Purpose: Methodology sample — NOT a real evaluation
Valid through: June 30, 2026
Next release: July 1, 2026 (Q3 2026)
License: CC-BY-4.0
Author: Nathan Maine
What This Is
A 46-question preview of the CMMC compliance-AI benchmark methodology. This is NOT a full evaluation — it is a small sample intended to show how this benchmark structures compliance AI testing.
v1 is the… See the full description on the dataset page: https://huggingface.co/datasets/Nathan-Maine/cmmc-benchmark-v1-preview-2026-q2.cmmc-benchmark-v2-spotcheck-2026-q2
CMMC Benchmark v2 Spot Check — Q2 2026
Version: 2026-q2
Tier: v2 Spot Check (454 questions)
Purpose: Triage tool — catch obvious failures before committing to full evaluation
Valid through: June 30, 2026
Next release: July 1, 2026 (Q3 2026)
License: CC-BY-4.0
Author: Nathan Maine
What This Is
A 454-question spot check for evaluating compliance AI systems against CMMC 2.0, NIST SP 800-171/172, and DFARS knowledge. "Spot check" is deliberate terminology — v2 is… See the full description on the dataset page: https://huggingface.co/datasets/Nathan-Maine/cmmc-benchmark-v2-spotcheck-2026-q2.Triangle104__Q2.5-EvaHumane-RP-details
Dataset Card for Evaluation run of Triangle104/Q2.5-EvaHumane-RP
Dataset automatically created during the evaluation run of model Triangle104/Q2.5-EvaHumane-RP
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Triangle104__Q2.5-EvaHumane-RP-details.cmmc-benchmark-v3-comprehensive-2026-q2
CMMC Benchmark v3 Comprehensive — Q2 2026
Version: 2026-q2
Tier: v3 Comprehensive (1,273 questions, 15 evaluation dimensions)
Purpose: The full, authoritative evaluation for compliance AI
Valid through: June 30, 2026
Next release: July 1, 2026 (Q3 2026)
License: CC-BY-4.0
Author: Nathan Maine
What This Is
This is the comprehensive tier of the CMMC Compliance Benchmark suite: 1,273 questions across 15 evaluation dimensions, covering the full scope of CMMC 2.0 /… See the full description on the dataset page: https://huggingface.co/datasets/Nathan-Maine/cmmc-benchmark-v3-comprehensive-2026-q2.asb-esc-q2b-12Triangle104__Q2.5-14B-Instruct-1M-Harmony-details
Dataset Card for Evaluation run of Triangle104/Q2.5-14B-Instruct-1M-Harmony
Dataset automatically created during the evaluation run of model Triangle104/Q2.5-14B-Instruct-1M-Harmony
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Triangle104__Q2.5-14B-Instruct-1M-Harmony-details.math-18kcmmc-benchmark-v1-preview-2026-q2
CMMC Benchmark v1 Preview — Q2 2026
Version: 2026-q2
Tier: v1 Preview (46 questions)
Purpose: Methodology sample — NOT a real evaluation
Valid through: June 30, 2026
Next release: July 1, 2026 (Q3 2026)
License: CC-BY-4.0
Publisher: Memoriant, Inc.
What This Is
A 46-question preview of the Memoriant Industrial Benchmark methodology. This is NOT a full evaluation — it's a small sample intended to show you how Memoriant structures compliance AI testing.
v1 is the… See the full description on the dataset page: https://huggingface.co/datasets/memoriant/cmmc-benchmark-v1-preview-2026-q2.cmmc-benchmark-v2-spotcheck-2026-q2
CMMC Benchmark v2 Spot Check — Q2 2026
Version: 2026-q2
Tier: v2 Spot Check (454 questions)
Purpose: Triage tool — catch obvious failures before committing to full evaluation
Valid through: June 30, 2026
Next release: July 1, 2026 (Q3 2026)
License: CC-BY-4.0
Publisher: Memoriant, Inc.
What This Is
A 454-question spot check of compliance AI systems. "Spot check" is deliberate terminology — v2 is targeted sampling, not comprehensive evaluation.
v2 is a triage tool, not… See the full description on the dataset page: https://huggingface.co/datasets/memoriant/cmmc-benchmark-v2-spotcheck-2026-q2.asb-esc-q2b-colotest_wiki_rewire_q2
