datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
quantum-simulation-chemistry-materials
Neura Parse — Quantum Simulation of Chemistry & Materials: Encodings, VQE/QPE & Dynamics
An application-deep, code-backed vertical on simulating quantum matter: electronic-structure problems, fermion-to-qubit encodings, Hamiltonian factorizations, ground/excited-state and real-time-dynamics algorithms, and analog simulation, with end-to-end resource estimates and honest classical-competitor accounting. Built with Qiskit Nature, OpenFermion, PennyLane-QChem, and PySCF — far… See the full description on the dataset page: https://huggingface.co/datasets/Neura-parse/quantum-simulation-chemistry-materials.quants
QuAnTS: Question Answering on Time Series
QuAnTS is a challenging dataset designed to bridge the gap in question-answering research on time series data.
The dataset features a wide variety of questions and answers concerning human movements, presented as tracked skeleton trajectories.
QuAnTS also includes human reference performance to benchmark the practical usability of models trained on this dataset.
At present, there is no official leaderboard for this dataset.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/dasyd/quants.colab-strong-benchmark-reasoning
Colab Strong Benchmark: Reasoning & Trap (1.15GB Version)
A large-scale reasoning benchmark dataset with 2,875,000 carefully crafted questions containing logical traps and challenges.
☕ Dataset Statistics
📊 Dataset Visualizations
Topic
Plot
Category Distribution
Difficulty Distribution
Question Length Distribution
Trap Inclusion Distribution
Total Samples: 2,875,000
Total Size: 1.164 GB
Number of Shards: 58
Categories: logical_reasoning… See the full description on the dataset page: https://huggingface.co/datasets/QuantaSparkLabs/colab-strong-benchmark-reasoning.quantqa
QuantQA: Quantitative Finance Interview Questions
QuantQA is a curated dataset of 519 interview questions sourced from leading quantitative trading firms including Jane Street, Citadel, Two Sigma, Optiver, and SIG, in collaboration with CoachQuant.
Topic Distribution
Topic
Coverage
Probability
67%
Combinatorics
22%
Expected Value
21%
Conditional Probability
14%
Game Theory
11%
Note: Questions may cover multiple topics
Training Results… See the full description on the dataset page: https://huggingface.co/datasets/ReinforceNow/quantqa.quantum-computing
Neura Parse — Quantum Computing
A multi-format quantum computing dataset spanning theory and hardware — from qubits, gates, and algorithms to QPUs, error correction, quantum software (Qiskit/Cirq/PennyLane), and quantum machine learning. Records come as instruction/response pairs, open and multiple-choice Q&A, runnable code tasks, encyclopedic concepts, and pretraining-style text, so the dataset supports SFT, evaluation, and continued pretraining under one schema.
Part of… See the full description on the dataset page: https://huggingface.co/datasets/Neura-parse/quantum-computing.gmat-quant-corpus
GMAT Quant Corpus for Solver + Retrieval
This dataset is intended for retrieval over GMAT-style quantitative teaching content.
Files
gmat_hf_chunks.jsonl — retrieval chunks used by the app
gmat_question_seed.jsonl — question seed data
gmat_topic_index.json — topic metadata/index
Default dataset viewer
The default dataset viewer is configured to load only:
gmat_hf_chunks.jsonl
This avoids schema conflicts with the other support files in the repository.
srb-data
SRB Data
Benchmark datasets for Scientific Recall Bench — a config-driven framework for evaluating LLM agent memory systems and retrieval architectures on scientific corpora.
Two scientific-domain benchmark slices, designed with a shared schema so findings can be validated across domains and corpus sizes:
public_ai_memory
public_transformers
Domain
LLM agent memory research
Transformer architecture research
Papers
103 structured notes + 39 full-text mirrors¹
252… See the full description on the dataset page: https://huggingface.co/datasets/quantellence/srb-data.WildChat-1M
Dataset Card for WildChat
Dataset Description
Paper: https://arxiv.org/abs/2405.01470
Interactive Search Tool: https://wildvisualizer.com (paper)
License: ODC-BY
Language(s) (NLP): multi-lingual
Point of Contact: Yuntian Deng
Dataset Summary
WildChat is a collection of 1 million conversations between human users and ChatGPT, alongside demographic data, including state, country, hashed IP addresses, and request headers. We collected WildChat by… See the full description on the dataset page: https://huggingface.co/datasets/quantcalc/WildChat-1M.quant-finance-hft-trading-2026
⚡ Quantitative Finance & High-Frequency Trading (HFT) SFT/DPO Suite (2026)
Institutional-grade instruction fine-tuning and preference alignment dataset for training domain-expert Large Language Models in Quantitative Finance, Algorithmic Execution, and Ultra-Low-Latency HFT Systems.
Engineered to the Mandatory Tier-1 Quality Standard: 80–150 lines of dense, production-grade C++20 and Rust per code snippet. Zero stubs, zero toy snippets, zero heap allocations on the critical… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/quant-finance-hft-trading-2026.quantibias
QuantiBias
Benchmarking quantization-induced bias in large language models.
QuantiBias measures a specific, under-audited failure mode: post-training quantization can leave a
model's short-form safety behavior almost untouched while the bias it volunteers in open-ended
generation rises. A standard audit that reads refusal rates and multiple-choice bias scores reports
the compressed model unchanged; QuantiBias shows what that audit misses.
Content warning. QuantiBias evaluates… See the full description on the dataset page: https://huggingface.co/datasets/emilioferrara/quantibias.quantum-information-and-complexity-theory
Neura Parse — Quantum Information & Complexity Theory: Channels, Entropies, Classes & the Structure of Advantage
A proof-based theoretical-foundations vertical uniting quantum information theory (channels, entropies, entanglement measures, distinguishability, capacities, Shannon theory) with quantum complexity theory and the structure of quantum advantage (classes, Hamiltonian complexity, sampling-based advantage and its verification, pseudorandomness, dequantization).… See the full description on the dataset page: https://huggingface.co/datasets/Neura-parse/quantum-information-and-complexity-theory.fault-tolerant-quantum-computing
Neura Parse — Fault-Tolerant Quantum Computing: QEC Codes, Decoders, Magic States & Resource Estimation
A deep, Stim-informed vertical on fault tolerance — QEC code families, decoders, fault-tolerant gate constructions, and the full physical-to-logical resource-estimation pipeline. Expands the general dataset's handful of error-correction topics into research-grade coverage including the 2024-2026 milestones: surface-code below threshold, qLDPC/bivariate-bicycle memories… See the full description on the dataset page: https://huggingface.co/datasets/Neura-parse/fault-tolerant-quantum-computing.quantum-compilation-and-programming
Neura Parse — Quantum Compilation & Programming
A code-heavy vertical on the quantum software/compilation stack: turning abstract quantum circuits and unitaries into device-executable programs. Covers unitary decomposition and circuit synthesis (Euler/ZYZ, KAK/Cartan, Solovay-Kitaev, Ross-Selinger gridsynth, numerical synthesis with BQSKit), gate-set/basis transpilation to native gate sets, qubit layout/mapping and routing under connectivity constraints (SABRE, VF2, SWAP… See the full description on the dataset page: https://huggingface.co/datasets/Neura-parse/quantum-compilation-and-programming.quantum-worldline-research
Quantum Worldline Research Data
Structured research data from the Quantum Worldline project - an AI-assisted research program investigating holographic forces in MERA tensor networks, worldline path integrals on AdS spacetime, and quantum simulation of lattice gauge theories.
Dataset Description
This dataset contains the complete structured output of the Quantum Worldline multi-agent research system, which automates the research cycle: discover - hypothesize - gate - test… See the full description on the dataset page: https://huggingface.co/datasets/Jonboy648/quantum-worldline-research.quantum-machine-learning-models
Neura Parse — Quantum Machine Learning Models: Encodings, Kernels, QNNs & Generative/Deep Architectures
A hands-on, code-first vertical on quantum models that learn from data. Spans data encodings/feature maps, variational classifiers, quantum kernels/QSVMs, and quantum neural networks through modern generative and deep architectures (quantum GANs, circuit Born machines, quantum Boltzmann machines, QCNNs, quantum autoencoders, quantum RL, and quantum… See the full description on the dataset page: https://huggingface.co/datasets/Neura-parse/quantum-machine-learning-models.agentic-lightweight-envs-runtime-20260528
Lightweight Agentic RL Runtime Environments
This dataset contains prebuilt runtime environments for lightweight agentic reinforcement learning. It is intended to be used with the accompanying environment server/runtime code. The LLM synthesis pipeline used to create these environments is not required for serving this dataset.
Files
runtime_catalog.json.gz: prebuilt runtime task catalog consumed by the env server.
task_drafts.json.gz: source task drafts and… See the full description on the dataset page: https://huggingface.co/datasets/quantumfr/agentic-lightweight-envs-runtime-20260528.quantum-error-mitigation-and-benchmarking
Neura Parse — Quantum Error Mitigation, Characterization & Benchmarking
A pre-fault-tolerance, code-backed vertical on getting trustworthy answers from noisy hardware and rigorously measuring device quality: error-mitigation techniques, characterization/tomography protocols, and benchmarking suites. Runnable Mitiq, pyGSTi, and Qiskit Experiments pipelines with honest sampling-overhead and bias/variance accounting — the practitioner and research toolkit the general dataset… See the full description on the dataset page: https://huggingface.co/datasets/Neura-parse/quantum-error-mitigation-and-benchmarking.ai-auto-train-datasets-cuda-5d-quantum-mindmap-simulations-generator-zkevms-immutablexadvanced-quantum-algorithms
Neura Parse — Advanced Quantum Algorithms: Derivations, QSVT/Block-Encoding & Hamiltonian Simulation
A derivation- and resource-analyzed algorithms vertical spanning the canonical fault-tolerant canon (with full proofs, complexity, and worked traces) and the modern QSVT/block-encoding toolkit through Hamiltonian simulation, amplitude estimation, and quantum linear systems. Turns the general dataset's one-topic-per-algorithm summaries into line-by-line derivations, lower… See the full description on the dataset page: https://huggingface.co/datasets/Neura-parse/advanced-quantum-algorithms.quantum-networking-and-distributed
Neura Parse — Quantum Networking, Repeaters & Distributed Quantum Computing
A systems-frontier vertical on connecting quantum devices: entanglement distribution and distillation, quantum repeaters, quantum-internet protocol stacks, quantum memories/transduction, and modular/distributed quantum computing (nonlocal gates, circuit knitting across nodes, blind/verifiable delegated computation). Covers protocol and simulation methods used with tools such as NetSquid and SeQUeNCe… See the full description on the dataset page: https://huggingface.co/datasets/Neura-parse/quantum-networking-and-distributed.ai-for-quantum
Neura Parse — AI for Quantum: ML & LLMs for Decoding, Control, Characterization & Software
The reverse quantum-AI direction — classical machine learning, RL, and LLMs/agents applied to make quantum computers work. Covers neural/transformer QEC decoders (AlphaQubit-style), RL/ML pulse and calibration control, neural-network quantum states, ML tomography and Hamiltonian/noise learning, learned circuit optimization, and LLM/agentic quantum software engineering (code generation… See the full description on the dataset page: https://huggingface.co/datasets/Neura-parse/ai-for-quantum.quantumchem-reasoningReasoning dataset for post-training Raven and QuantaMind chemisty LLMs with GRPO and RLVR.
QuantumAIquantum-machine-learning-theory
Neura Parse — Quantum Machine Learning Theory: Trainability, Generalization & Learning From Quantum Data
A research-depth, proof-oriented vertical on the learning theory of quantum models and quantum data. Covers why parameterized quantum circuits train or don't (barren plateaus), what they can represent, when they generalize or provably beat classical models, and — for quantum data — how to predict properties of unknown states/channels with few measurements (classical… See the full description on the dataset page: https://huggingface.co/datasets/Neura-parse/quantum-machine-learning-theory.QuantumMechanics
Quantum Mechanics Reasoning Dataset
🧠 High-quality physics reasoning chains for training thinking LLMs
Dataset Overview
This dataset provides systematic reasoning chains for quantum mechanics concepts, designed specifically for training thinking LLMs like GPT-OSS-20B. Each entry contains step-by-step logical progressions with mathematical expressions and physical interpretations.
🎯 Current Status: Comprehensive Release (v1.5.0) 🆕 CHAPTER 7 ADDED
1,350 high-quality… See the full description on the dataset page: https://huggingface.co/datasets/themanaspandey/QuantumMechanics.public_ai_memory_slice
Public AI Memory Slice
A scientific-domain benchmark for evaluating LLM agent memory systems on the AI / agent-memory research literature.
103 structured paper notes (~448K tokens) covering LLM agent memory, memory benchmarks, and adjacent cognitive-architecture / theory-formation work
81 full-text paper mirrors (~1.47M tokens), OCR extracted from open-access arXiv PDFs
66 main queries + 10 holdout queries with rubric-style ground truth, every must-have fact traceable to a verbatim… See the full description on the dataset page: https://huggingface.co/datasets/quantellence/public_ai_memory_slice.quantum-optimization
Neura Parse — Quantum Optimization, Annealing & Finance: QAOA, Adiabatic Methods & the Advantage Question
A research-plus-practitioner vertical on quantum approaches to combinatorial and continuous optimization and their most-piloted enterprise use cases. Covers QAOA theory and variants, adiabatic/annealing methods and D-Wave, QUBO/Ising encodings, amplitude-estimation Monte Carlo for finance, and the rigorous question of whether and where quantum beats classical (including… See the full description on the dataset page: https://huggingface.co/datasets/Neura-parse/quantum-optimization.gpqa
Dataset Card for GPQA
GPQA is a multiple-choice, Q&A dataset of very hard questions written and validated by experts in biology, physics, and chemistry. When attempting questions out of their own domain (e.g., a physicist answers a chemistry question), these experts get only 34% accuracy, despite spending >30m with full access to Google.
We request that you do not reveal examples from this dataset in plain text or images online, to reduce the risk of leakage into foundation… See the full description on the dataset page: https://huggingface.co/datasets/quantiles/gpqa.post-quantum-crypto-en
Post-Quantum Cryptography - Dataset EN
Comprehensive English dataset on post-quantum cryptography (PQC) migration, covering NIST algorithms (FIPS 203-206), migration phases, protocol impacts, quantum threats, and 70 detailed Q&A pairs.
Dataset Contents
Category
Entries
Description
PQC Algorithms
15
ML-KEM (Kyber), ML-DSA (Dilithium), SLH-DSA (SPHINCS+), FN-DSA (FALCON), HQC, RSA, ECDSA, Ed25519, X25519, AES-256
Migration Phases
12
Cryptographic inventory… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/post-quantum-crypto-en.AV-QuantBench-Dataset
AV-QuantBench
AV-QuantBench is a procedural audio-visual benchmark for evaluating multimodal foundation models on abstract temporal reasoning, cross-modal conflict detection, and synchronized data interpretation across finance, medical, and industrial domains.
This Hugging Face dataset repository is structured as a benchmark-style release. It contains:
split metadata in JSONL format,
question-answer annotations,
audio-visual sample assets,
manifest files by domain,
and… See the full description on the dataset page: https://huggingface.co/datasets/gfcfirefly/AV-QuantBench-Dataset.
