datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
glm-5.3-flash-mathnet-bon
glm-5.3-flash-mathnet-bon
This is the continuation and the final set of ox-alpha-mathnet-bon.
Verified chain-of-thought reasoning traces for competition mathematics, generated with GLM-5.3-Flash via best-of-N rejection sampling against the ShadenA/MathNet dataset (ICLR 2026).
Statistics (this split)
Metric
Value
Records (problem × attempt)
6,181
Distinct problems
848
Attempts per problem
7.29 (mean), 8 (max)
Accepted (answer_correct = true)
3,705… See the full description on the dataset page: https://huggingface.co/datasets/zakoman/glm-5.3-flash-mathnet-bon.napoleon_bonaparte
Napoleon Bonaparte
The Napoleon Bonaparte dataset is a collection of information and data related to Napoleon Bonaparte's life and reign. It includes details on his military campaigns, battles, conquests, and political career as Emperor of France. The dataset also contains information on the social and economic reforms he implemented in France, such as the establishment of the Napoleonic Code. The data is gathered from various sources, including historical records, biographies, and… See the full description on the dataset page: https://huggingface.co/datasets/MH0386/napoleon_bonaparte.bondshift-organic-chemistry
BondShift: Organic Chemistry Mechanism Dataset
10,000 ground-truth-separated records for mechanism diagnosis, misconception repair, and chemistry tutoring.
A narrow, auditable V1 dataset built from independently constructed scenario blueprints and deterministic answer keys.
TL;DR
BondShift addresses the right answer, wrong mechanism problem. It trains models to examine electron flow, formal
charge, intermediates, pathway choice, and stereochemical… See the full description on the dataset page: https://huggingface.co/datasets/prathmeshadsod/bondshift-organic-chemistry.ox-alpha-mathnet-bon
ox-alpha-mathnet-bon
Verified chain-of-thought reasoning traces for competition mathematics, generated with the ox-alpha reasoning model via best-of-N rejection sampling against the ShadenA/MathNet dataset (ICLR 2026).
ox-alpha stealth model spec
On August 20, 2026, an anonymous model designated stealth/ox-alpha appeared on OpenRouter (and OpenCode) with no disclosed developer, a 1M-token context window, multimodal input (text, image, video), and a roughly… See the full description on the dataset page: https://huggingface.co/datasets/zakoman/ox-alpha-mathnet-bon.bondfoundry-quant-sampleBondFoundry Quant Sample — 10 Records
Premium synthetic instruction-tuning data for enterprise AI teams building domain-specific quantitative finance models.
What's inside
10 of the highest-depth records from BondFoundry's quantitative finance catalogue. Each record is generated by a senior quant PM-level persona operating under real institutional constraints — margin call pressure, risk committee pushback, regulatory deadlines, portfolio drawdown scenarios.
Average word count: 568 words per… See the full description on the dataset page: https://huggingface.co/datasets/BondFoundry/bondfoundry-quant-sample.bonsai2-drafter-eval
Bonsai 2 drafter evaluation corpora
Prompts and greedy responses from PrismML's
Ternary-Bonsai-2-27B, recorded
as token ids together with the per-round accepted lengths of the speculative decoding loop that
produced them. The data exists to evaluate and train DFlash 2 drafters against this one target.
Target
prism-ml/Ternary-Bonsai-2-27B-mlx-2bit (MLX pack), greedy, temperature 0
Loop
mlx-dspark 0.18.0 DFlash 2 loop, stock drafter z-lab/Qwen3.8-27B-DFlash2, draft… See the full description on the dataset page: https://huggingface.co/datasets/Schiltmans/bonsai2-drafter-eval.bondfoundry-legal-sampleBondFoundry Legal Sample — 10 Records
Premium synthetic instruction-tuning data for enterprise AI teams building domain-specific legal reasoning models.
What's inside
10 of the highest-depth records from BondFoundry's legal catalogue. Each record is generated by a Magic Circle partner-level persona operating under real institutional constraints — client deadlines, regulatory pressure, multi-jurisdiction complexity, risk committee scrutiny.
Average word count: 1,015 words per record
Sample… See the full description on the dataset page: https://huggingface.co/datasets/BondFoundry/bondfoundry-legal-sample.ralphthon-release-readiness
Ralphthon Release Readiness Environment
This package is a public-safe, deterministic replay sandbox for synthetic release-readiness tasks. It is an OpenEnv interaction adapter over frozen synthetic fixtures, not a historical corpus, live deployment system, company digital twin, or causal simulator.
Run the deterministic demonstration:
uv run ralphthon-env demo --fixture-id test-ready-001
The command prints one canonical JSON line. Other public commands are compile, validate… See the full description on the dataset page: https://huggingface.co/datasets/bong-9/ralphthon-release-readiness.bondfoundry-cyber-sample--BondFoundry Cybersecurity Sample — 10 Records
Premium synthetic instruction-tuning data for enterprise AI teams building domain-specific cybersecurity models.
What's inside:
10 of the highest-depth records from BondFoundry's cybersecurity catalogue. Each record is generated by a CISO-level persona operating under real institutional constraints — budget fights, risk committee pushback, regulatory deadlines, active incident pressure.
Average word count: 914 words per record
Sample record… See the full description on the dataset page: https://huggingface.co/datasets/BondFoundry/bondfoundry-cyber-sample.
