datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ClimbMixClimbMix is a high-quality pre-training corpus released by NVIDIA. Here is the description:
ClimbMix is a compact yet powerful 400-billion-token dataset designed for efficient pre-training that delivers superior performance under an equal token budget. It was introduced in this paper.
We proposed a new algorithm to filter and mix the dataset. First, we grouped the data into 1,000 groups based on topic information. Then we applied two classifiers: one to detect advertisements and another to… See the full description on the dataset page: https://huggingface.co/datasets/OptimalScale/ClimbMix.optiq-lab-traces
OptiQ Lab Traces
Research and tool-calling sessions produced by OptiQ Lab, the local web UI that ships with mlx-optiq. Each session is a complete run: a deep-research report built from live web sources, or a multi-turn agent loop driving the Lab's own sandboxed tools.
The dataset is 866 sessions in HuggingFace Session-Traces format (the agent-traces viewer). Each .jsonl file is one session: a header line carrying the run's metadata, then one message per turn.
The two… See the full description on the dataset page: https://huggingface.co/datasets/mlx-community/optiq-lab-traces.speculators-ci-datasets
speculator-tutorial
Raw vs. on-policy regenerated conversation data for training speculative-decoding
drafters (EAGLE-3 / DFlash / DSpark style), with the original source data kept alongside
so you can see exactly what regeneration changes and why it matters.
Prompts come from UltraChat-200k. The verifier / teacher model is Qwen/Qwen3-8B.
Why regenerate at all?
A speculative-decoding drafter is trained to predict what the verifier would say next.
If you train it… See the full description on the dataset page: https://huggingface.co/datasets/inference-optimization/speculators-ci-datasets.optiq-code-traces
OptiQ Code Traces
Gold-verified agentic software-engineering trajectories, produced by OptiQ Code, the terminal coding agent for local models on a Mac. Each trajectory is a full tool-calling run against a real repository bug, and every resolved label is set by executing the gold tests (FAIL_TO_PASS + PASS_TO_PASS) after applying the model's patch, never by the agent's own self-report.
The dataset is 1,789 agent sessions in HuggingFace Session-Traces format (the agent-traces… See the full description on the dataset page: https://huggingface.co/datasets/mlx-community/optiq-code-traces.deep-space-optical-chip-thermal-dataset
🚀 Deep Space Optical Chip Thermal Dataset 🪐
🌡️ 40,000 scenario-based prompt and response pairs on thermal mitigation for photonic chips in scientific instruments aboard deep-space probes, covering refractive index drift, waveguide misalignment, and thermal stress across materials, instruments, and environments.
⚠️ Disclaimer: All entries are synthetically generated. Material coefficients are drawn from published typical values, but no row is based on mission logs or flight… See the full description on the dataset page: https://huggingface.co/datasets/Taylor658/deep-space-optical-chip-thermal-dataset.quantum-optimization
Neura Parse — Quantum Optimization, Annealing & Finance: QAOA, Adiabatic Methods & the Advantage Question
A research-plus-practitioner vertical on quantum approaches to combinatorial and continuous optimization and their most-piloted enterprise use cases. Covers QAOA theory and variants, adiabatic/annealing methods and D-Wave, QUBO/Ising encodings, amplitude-estimation Monte Carlo for finance, and the rigorous question of whether and where quantum beats classical (including… See the full description on the dataset page: https://huggingface.co/datasets/Neura-parse/quantum-optimization.Code_Opt_Triton
Overview
This dataset, TEEN-D/Code_Opt_Triton, is an extended version of the publicly available GPUMODE/Inductor_Created_Data_Permissive dataset. It contains pairs of original (PyTorch or Triton) programs and their equivalent Triton code (generated by torch inductor), intended for training models in PyTorch-to-Triton code translation and optimization.
The primary modification in this extended version is that each optimized Triton code snippet is paired with both its original source… See the full description on the dataset page: https://huggingface.co/datasets/Teen-Different/Code_Opt_Triton.swiss-web-premium-ch
*.ch Swiss Web Premium (A+)
Grade A+ -- 112.4M tokens -- 110,491 records -- 29 languages -- Full provenance -- PII-redacted -- RAG-ready -- SFT-formatted
A production-grade Swiss web corpus from the .ch TLD namespace. 110,491 documents independently quality-scored (avg 93.3/100, minimum 90), PII-redacted, and SHA256-verified. Built for LLM training, RAG pipelines, SFT fine-tuning, and multilingual NLP.
OptiTransferData Portfolio
Premium sovereign web corpora for… See the full description on the dataset page: https://huggingface.co/datasets/OptiTransferData/swiss-web-premium-ch.tropt-optbench-triggers
TROPT — OptBench Triggers (Exp1: optimizer benchmark)
A store of optimized adversarial trigger suffixes produced by the
TROPT white-box optimizers, together with
the transfer evaluation of each trigger across a held-out set of harmful
instructions.
⚠️ Intended use — defensive security research only. These are adversarial
artifacts for evaluating and hardening LLM robustness (red-teaming,
jailbreak-robustness benchmarking). The harmful instructions come from the
public ClearHarm… See the full description on the dataset page: https://huggingface.co/datasets/MatanBT/tropt-optbench-triggers.swiss-web-premium-ch-full
*.ch Swiss Web Premium (A+) -- Full Dataset
Grade A+ -- 112.4M tokens -- 110,491 records -- 29 languages -- 22 production files -- 554.1 MB
The complete production release of OptiTransfer's Swiss web corpus. 110,491 documents from the .ch TLD namespace, independently quality-scored (avg 93.3/100, minimum 90), PII-redacted, and SHA256-verified. Delivered in Parquet, JSONL, language splits, and pre-built RAG chunks.
This is the full commercial dataset. For evaluation, see the free… See the full description on the dataset page: https://huggingface.co/datasets/OptiTransferData/swiss-web-premium-ch-full.
