datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
black-box-api-challenges
Dataset Card
Paper: On the Challenges of Using Black-Box APIs for Toxicity Evaluation in Research
Abstract: Perception of toxicity evolves over time and often differs between geographies and cultural backgrounds. Similarly, black-box commercially available APIs for detecting toxicity, such as the Perspective API, are not static, but frequently retrained to address any unattended weaknesses and biases. We evaluate the implications of these changes on the reproducibility of findings… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/black-box-api-challenges.test
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/boxin-wbx/test.Neurips2026_evaluating_accuracy_KV-cache_reuse_techniques
BoxOffice Verified Seeds
This dataset contains the released BoxOffice seed datasets used in the
benchmark pipeline described in the accompanying paper. The release includes
ten verified seeds:
7
11
13
17
19
23
29
31
47
73
For each seed, we provide:
a full JSONL file containing warmup rows plus evaluation rows
an eval JSONL file containing only the evaluation rows
a manifest JSON file
a validation JSON file with directional warmup counts
Layout
viewer/
normalized… See the full description on the dataset page: https://huggingface.co/datasets/Boxoffice1280/Neurips2026_evaluating_accuracy_KV-cache_reuse_techniques.open-thoughts-4-5k-math-qwen3-32b-only-235b-has-boxed
OpenThoughts4 5K Math - Qwen3-32B Only 235B Has Boxed Answer
Overview
This dataset contains 5,441 samples from the OpenThoughts4 math dataset where only Qwen3-235B-A22B produced a valid \boxed{} answer (Qwen3-32B did not). This dataset contains the Qwen3-32B reasoning traces (which lack valid boxed answers).
Relationship to Other Datasets
This is one of 10 child datasets derived from two parent datasets:
Parent datasets (29,963 samples each):… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-5k-math-qwen3-32b-only-235b-has-boxed.open-thoughts-4-1k-math-qwen3-235b-a22b-only-32b-has-boxed
OpenThoughts4 1K Math - Qwen3-235B-A22B Only 32B Has Boxed Answer
Overview
This dataset contains 1,477 samples from the OpenThoughts4 math dataset where only Qwen3-32B produced a valid \boxed{} answer (Qwen3-235B-A22B did not). This dataset contains the Qwen3-235B-A22B reasoning traces (which lack valid boxed answers).
Relationship to Other Datasets
This is one of 10 child datasets derived from two parent datasets:
Parent datasets (29,963 samples each):… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-1k-math-qwen3-235b-a22b-only-32b-has-boxed.open-thoughts-4-6k-math-qwen3-32b-neither-has-boxed
OpenThoughts4 6K Math - Qwen3-32B Neither Has Boxed Answer
Overview
This dataset contains 6,110 samples from the OpenThoughts4 math dataset where neither Qwen3-32B nor Qwen3-235B-A22B produced a valid \boxed{} answer. This dataset contains the Qwen3-32B reasoning traces.
Relationship to Other Datasets
This is one of 10 child datasets derived from two parent datasets:
Parent datasets (29,963 samples each):… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-6k-math-qwen3-32b-neither-has-boxed.open-thoughts-4-6k-math-qwen3-235b-a22b-neither-has-boxed
OpenThoughts4 6K Math - Qwen3-235B-A22B Neither Has Boxed Answer
Overview
This dataset contains 6,110 samples from the OpenThoughts4 math dataset where neither Qwen3-32B nor Qwen3-235B-A22B produced a valid \boxed{} answer. This dataset contains the Qwen3-235B-A22B reasoning traces.
Relationship to Other Datasets
This is one of 10 child datasets derived from two parent datasets:
Parent datasets (29,963 samples each):… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-6k-math-qwen3-235b-a22b-neither-has-boxed.Boxing-Champions
World Heavyweight Boxing Champions Dataset
This dataset contains information about world heavyweight boxing champions extracted from the Wikipedia page. It includes details such as champion names, reign periods, and title defenses.
Dataset Structure
Columns
Column Name
Description
No
The ordinal number of the champion.
Champion
Name of the heavyweight boxing champion.
Recognition
The organization or title under which the reign was recognized.… See the full description on the dataset page: https://huggingface.co/datasets/iamramzan/Boxing-Champions.open-thoughts-4-5k-math-qwen3-235b-a22b-only-235b-has-boxed
OpenThoughts4 5K Math - Qwen3-235B-A22B Only 235B Has Boxed Answer
Overview
This dataset contains 5,441 samples from the OpenThoughts4 math dataset where only Qwen3-235B-A22B produced a valid \boxed{} answer (Qwen3-32B did not). This dataset contains the Qwen3-235B-A22B reasoning traces.
Relationship to Other Datasets
This is one of 10 child datasets derived from two parent datasets:
Parent datasets (29,963 samples each):… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-5k-math-qwen3-235b-a22b-only-235b-has-boxed.open-thoughts-4-1k-math-qwen3-32b-only-32b-has-boxed
OpenThoughts4 1K Math - Qwen3-32B Only 32B Has Boxed Answer
Overview
This dataset contains 1,477 samples from the OpenThoughts4 math dataset where only Qwen3-32B produced a valid \boxed{} answer (Qwen3-235B-A22B did not). This dataset contains the Qwen3-32B reasoning traces.
Relationship to Other Datasets
This is one of 10 child datasets derived from two parent datasets:
Parent datasets (29,963 samples each):… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-1k-math-qwen3-32b-only-32b-has-boxed.UltraData-Math-boxed
UltraData-Math-boxed
High-confidence subset of UltraData/UltraData-Math containing only problems with \boxed{} answers.
Original dataset: UltraData/UltraData-Math (53M samples across 2 configs)
Variants
Config
Samples
Description
default
390,082
Raw \boxed{} extractions, minimal cleanup
normalized
217,049
Filtered + validated: no units, no variables, no multi-value, sympy LaTeX verified
Why boxed-only?
UltraData-Math is synthetic textbook… See the full description on the dataset page: https://huggingface.co/datasets/apat1n/UltraData-Math-boxed.boxoffice-verified-seeds
BoxOffice Verified Seeds
This dataset contains the released BoxOffice seed datasets used in the
benchmark pipeline described in the accompanying paper. The release includes
ten verified seeds:
7
11
13
17
19
23
29
31
47
73
For each seed, we provide:
a full JSONL file containing warmup rows plus evaluation rows
an eval JSONL file containing only the evaluation rows
a manifest JSON file
a validation JSON file with directional warmup counts
Layout
seeds/… See the full description on the dataset page: https://huggingface.co/datasets/littlechest/boxoffice-verified-seeds.boxoffice1280
BoxOffice Verified Seeds
This dataset contains the released BoxOffice seed datasets used in the
benchmark pipeline described in the accompanying paper. The release includes
ten verified seeds:
7
11
13
17
19
23
29
31
47
73
For each seed, we provide:
a full JSONL file containing warmup rows plus evaluation rows
an eval JSONL file containing only the evaluation rows
a manifest JSON file
a validation JSON file with directional warmup counts
Layout
viewer/
normalized… See the full description on the dataset page: https://huggingface.co/datasets/Boxoffice1280/boxoffice1280.OpenMathInstruct-2-Boxed-8k
alex-chiu/OpenMathInstruct-2-Boxed-8k
A deterministic 8K subset of nvidia/OpenMathInstruct-2 for short non-thinking math SFT before RL.
Splits
train: 8,192 rows
validation: 256 rows
Construction
Every generated_solution contains a complete final \boxed{...}
Duplicate problems and explicit <think> / </think> outputs are excluded
No answer verification against expected_answer is performed
The original generated_solution is preserved verbatim… See the full description on the dataset page: https://huggingface.co/datasets/alex-chiu/OpenMathInstruct-2-Boxed-8k.potable
Potable Dataset
Expert-curated fine-tuning data for drinking water treatment operations
Maintained by Operational Inference | Keith Wilkinson, T5 Certified Water Treatment Operator
Dataset description
The Potable Dataset is a fine-tuning dataset for language models in the drinking water treatment domain. Every example is authored or reviewed by a licensed water treatment operator with Class T5 certification — the highest treatment license issued in California — with over… See the full description on the dataset page: https://huggingface.co/datasets/boxwrench/potable.
