datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Alexandria_geometry_optimization_paths_PBE_2D
Cite this dataset Schmidt, J., Hoffmann, N., Wang, H., Borlido, P., Carriço, P. J. M. A., Cerqueira, T. F. T., Botti, S., and Marques, M. A. L. Alexandria geometry optimization paths PBE 2D. ColabFit, 2025. https://doi.org/10.60732/8781419f
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:
https://materials.colabfit.org/id/DS_6pieq95jrqpn_0
Visit the ColabFit… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/Alexandria_geometry_optimization_paths_PBE_2D.speculators-ci-datasets
speculator-tutorial
Raw vs. on-policy regenerated conversation data for training speculative-decoding
drafters (EAGLE-3 / DFlash / DSpark style), with the original source data kept alongside
so you can see exactly what regeneration changes and why it matters.
Prompts come from UltraChat-200k. The verifier / teacher model is Qwen/Qwen3-8B.
Why regenerate at all?
A speculative-decoding drafter is trained to predict what the verifier would say next.
If you train it… See the full description on the dataset page: https://huggingface.co/datasets/inference-optimization/speculators-ci-datasets.Alexandria_geometry_optimization_paths_PBE_1D
Cite this dataset Schmidt, J., Hoffmann, N., Wang, H., Borlido, P., Carriço, P. J. M. A., Cerqueira, T. F. T., Botti, S., and Marques, M. A. L. Alexandria geometry optimization paths PBE 1D. ColabFit, 2025. https://doi.org/10.60732/12246d46
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:
https://materials.colabfit.org/id/DS_xnio123pebli_0
Visit the ColabFit… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/Alexandria_geometry_optimization_paths_PBE_1D.dflash-code-multilingual-teacher-responses-qwen235b
Code + Multilingual Teacher Responses (Qwen3-235B-A22B-Instruct-2507)
This repo now contains 302,800 total samples across the main blended
data.jsonl / .parquet file plus a second Nemotron-only file
(nemotron_code_teacher_responses.jsonl / .parquet). All responses were
generated by Qwen3-235B-A22B-Instruct-2507 in non-thinking mode
(enable_thinking=false) to match downstream speculator training and eval.
Built in two batches: an initial 59,506-row batch (50K code + 9.5K… See the full description on the dataset page: https://huggingface.co/datasets/inference-optimization/dflash-code-multilingual-teacher-responses-qwen235b.nl-optimization-instantiation-metrics
Natural-Language Optimization Instantiation Metrics (v1.2)
This dataset is a text-free metrics release for a retrieval-assisted natural-language optimization instantiation pipeline. It contains per-example and aggregate evaluation outcomes for a frozen pipeline evaluated on the NLP4LP benchmark, the OptMath external validation domain (added in v1.1), and, as of v1.2, a 13-method schema-retrieval/grounding diagnostic suite (per-query breakdowns, bottleneck taxonomy, threshold… See the full description on the dataset page: https://huggingface.co/datasets/SoroushVahidi/nl-optimization-instantiation-metrics.SWE-bench_MultilingualCARMO-UltraFeedbackquantum-optimization
Neura Parse — Quantum Optimization, Annealing & Finance: QAOA, Adiabatic Methods & the Advantage Question
A research-plus-practitioner vertical on quantum approaches to combinatorial and continuous optimization and their most-piloted enterprise use cases. Covers QAOA theory and variants, adiabatic/annealing methods and D-Wave, QUBO/Ising encodings, amplitude-estimation Monte Carlo for finance, and the rigorous question of whether and where quantum beats classical (including… See the full description on the dataset page: https://huggingface.co/datasets/Neura-parse/quantum-optimization.manufacturing-cost-optimization-2026Q3
Manufacturing Quarterly Cost Optimization Dataset (2026Q3)
Unified quarterly cost-analysis dataset for the manufacturing group, merged from
the China / Japan / India factory datasets hosted on Hugging Face.
Contents
11,100 records (>= 10,000) covering 9 plants across 3 regions.
Source datasets:
toolathon123/manufacturing-cn-energy-2026Q3 — China energy & raw material (4,200 rows)
toolathon123/manufacturing-jp-maintenance-2026Q3 — Japan maintenance & downtime (3… See the full description on the dataset page: https://huggingface.co/datasets/toolathon123/manufacturing-cost-optimization-2026Q3.optimization-benchmark-datasetPaper available on arXiv
CARMO-UltraFeedback-BinarizedAMPO-OPT-Selectionmath-optimization-reasoning-v3ai-text-detection-datasetnigerian_transport_and_logistics_route_optimization
Nigeria Transport & Logistics – Route Optimization | Africa (Electric Sheep Africa metadata inventory)
Size category: 100K<n<1M - Formats: parquet - Sector: infrastructure_transport - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/nigerian_transport_and_logistics_route_optimization.nigerian_energy_and_utilities_ai_grid_optimization
Nigerian Energy & Utilities – AI Grid Optimization | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: parquet - Sector: energy - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/nigerian_energy_and_utilities_ai_grid_optimization.classification-ie-optimizationembedding-ie-optimizationmath-optimization-reasoning-v1math-optimization-reasoning-v2SWE-bench_Lite
Dataset Summary
SWE-bench Lite is subset of SWE-bench, a dataset that tests systems’ ability to solve GitHub issues automatically. The dataset collects 300 test Issue-Pull Request pairs from 11 popular Python. Evaluation is performed by unit test verification using post-PR behavior as the reference solution.
The dataset was released as part of SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
Want to run inference now?
This dataset only contains the… See the full description on the dataset page: https://huggingface.co/datasets/inference-optimization/SWE-bench_Lite.london-cvrptw-dyamaic-optimization-rl
london-dynamic-routing
Task dataset for the LondonDynamicRouting OpenReward environment:
dynamic, multi-horizon, weather- and traffic-aware vehicle routing on the
real London road network.
100 tasks of monotonically increasing difficulty (1..100), partitioned into
tutorial (5), train (70), test (25) splits.
Format
A single tasks.parquet file. Each row is a fully-specified, deterministic
episode. The columns match the TaskSpec schema documented in the
environment repo. Key… See the full description on the dataset page: https://huggingface.co/datasets/kosiasuzu/london-cvrptw-dyamaic-optimization-rl.huberman_lab_Dr._Kyle_Gillett_Tools_for_Hormone_Optimization_in_MalesSWE-bench_VerifiedDataset Summary
SWE-bench Verified is a subset of 500 samples from the SWE-bench test set, which have been human-validated for quality. SWE-bench is a dataset that tests systems’ ability to solve GitHub issues automatically. See this post for more details on the human-validation process.
The dataset collects 500 test Issue-Pull Request pairs from popular Python repositories. Evaluation is performed by unit test verification using post-PR behavior as the reference solution.
The original… See the full description on the dataset page: https://huggingface.co/datasets/inference-optimization/SWE-bench_Verified.vision-embedding-ie-optimizationhuberman_lab_Dr__Duncan_French_How_to_Exercise_for_Strength_Gains__Hormone_Optimizationhuberman_lab_Dr__Kyle_Gillett_Tools_for_Hormone_Optimization_in_MalesAMPO-Coreset-selectionOrca-Direct-Preference-Optimizationhyperparameter-optimization-dataset
