datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CMAPSS_Jet_Engine_Simulated_Datatau2-simulated
tau2 Simulated Training Set
Made with the whileai SDK · Collections: Simulation, Start here: foundational post-training datasets
The training set that took a base model from 5% to 30% on tau2-bench
telecom, made from nothing but the agent's tool list and policy.
If you build a customer-facing agent, you already have the two files this
dataset was made from: the tools it can call and the policy it follows.
The whileai SDK turned those into 1,057 graded conversations across the… See the full description on the dataset page: https://huggingface.co/datasets/while-ai/tau2-simulated.UserBehavioralDivergence-simulated-conversationsfineweb-ir-simulated-search-queries
fineweb-ir-simulated-search-queries
An English web-retrieval dataset built by generating simulated search queries for FineWeb-style positive target documents.
This dataset contains English query-document pairs derived from HuggingFaceFW/fineweb-edu.
Each row is designed so that the associated document is a positive retrieval target for the generated query.
The queries were generated with a Qwen3.5-35B-A3B model quantized to NVFP4 and served through OpenAI-compatible vLLM serve… See the full description on the dataset page: https://huggingface.co/datasets/hotchpotch/fineweb-ir-simulated-search-queries.arxiv-ir-simulated-search-queries
arxiv-ir-simulated-search-queries
An arXiv retrieval dataset with more than 2.8 million simulated specialist search queries and paper-level positive targets.
This dataset contains 2,875,637 query-document pairs derived from arXiv title-and-abstract records.
Each row is designed so that the associated arXiv paper record is a positive retrieval target for the generated query.
The queries were generated with a Qwen3.5-35B-A3B model quantized to NVFP4, and were intentionally… See the full description on the dataset page: https://huggingface.co/datasets/hotchpotch/arxiv-ir-simulated-search-queries.AEMO_simulated_trade
AEMO Battery Trading Dataset
Note (Aug 2026): The SDP-teacher trajectory dataset (data/aemo_dt_sdp/ in the repo) is
now the preferred training data for the shipped model. The original FCAS dataset below was the
training source for the Jul 2026 v2 pretrained model and the GRPO study. Both are historical —
the Stage C standalone DT (models/aemo/dt/aemo_dt_sdp_jtsoc_fullcorpus.pt) was trained on
SDP-teacher trajectories with J_t(soc) RTG prompts.
Files
File… See the full description on the dataset page: https://huggingface.co/datasets/mrvictoru/AEMO_simulated_trade.wikipedia-english-ir-simulated-search-queries
wikipedia-english-ir-simulated-search-queries
An English Wikipedia retrieval dataset with more than 29 million simulated search queries and paragraph-level positive targets.
This dataset contains 29,366,101 English query-document pairs derived from Wikipedia.
Each row is designed so that the associated Wikipedia paragraph is a positive retrieval target for the generated query.
The queries were generated with a Qwen3.5-35B-A3B model quantized to NVFP4, and were intentionally… See the full description on the dataset page: https://huggingface.co/datasets/hotchpotch/wikipedia-english-ir-simulated-search-queries.simulated_spectrapubmed-abstract-ir-simulated-search-queries
pubmed-abstract-ir-simulated-search-queries
A PubMed retrieval dataset with simulated specialist search queries and abstract-level positive targets.
This dataset contains 2,355,329 query-document pairs derived from PubMed title-and-abstract records.
Each row is designed so that the associated PubMed record is a positive retrieval target for the generated query.
The queries were generated with a Qwen3.5-35B-A3B model quantized to NVFP4, and were intentionally constructed to… See the full description on the dataset page: https://huggingface.co/datasets/hotchpotch/pubmed-abstract-ir-simulated-search-queries.ccnews-ir-simulated-search-queries
ccnews-ir-simulated-search-queries
An English news-retrieval dataset with 1.84 million simulated search queries paired with positive CC-News-style document targets.
This dataset contains 1,839,547 English query-document pairs derived from the English subset of multilingual CC-News.
Each row is designed so that the associated news document is a positive retrieval target for the generated query.
The queries were generated with a Qwen3.5-35B-A3B model quantized to NVFP4 and… See the full description on the dataset page: https://huggingface.co/datasets/hotchpotch/ccnews-ir-simulated-search-queries.AEMO_simulated_trade_sdp
AEMO SDP-Teacher Trajectories
Offline trajectories for Decision Transformer training, generated by replaying the honest SDP/MPC executor on historical Australian NEM (AEMO) market data. These are the teacher trajectories from the energydecision research codebase.
Each row is a single 5-minute market interval with a self-consistent (normalized observation, 9-dim action, reward) triple: the action is what the honest optimal planner dispatched, the reward is what it earned, and the… See the full description on the dataset page: https://huggingface.co/datasets/mrvictoru/AEMO_simulated_trade_sdp.wtd_simulated_dataspeech-simulated-medical-exams
Speech Simulated Medical Exams
Simulated patient-physician medical exam conversations with rich speech metadata annotations. Built for training single-step ASR models that transcribe and annotate multiple concepts simultaneously, including speaker changes, emotions, intents, and roles.
Dataset Details
Property
Value
Examples
25,706
Language
English
Audio
16 kHz WAV
Source
Simulated medical interviews (respiratory focus)
Features… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/speech-simulated-medical-exams.simulated_rirs_dataset
Simulated Rirs Dataset
Dataset Description
This dataset contains 400 samples organized across multiple splits and 4 subsets.
The dataset includes audio data.
Dataset Structure
Subsets
This dataset includes the following subsets:
original: 100 samples
train: 100 samples
largeroom: 100 samples
train: 100 samples
mediumroom: 100 samples
train: 100 samples
smallroom: 100 samples
train: 100 samples
Usage
Load specific subset and… See the full description on the dataset page: https://huggingface.co/datasets/sujalappa/simulated_rirs_dataset.binary-function-code-simulatederc8004-simulated-agents
ERC-8004 Simulated Agents — labeled synthetic dataset (6,000 agents)
⚠️ This dataset is fully synthetic. No public labeled dataset of malicious ERC-8004 agents exists (the standard reached mainnet in 2026 and exposes no trust label), so this dataset simulates the feature distributions the three ERC-8004 registries would expose, for training/evaluating trustworthiness models. For real on-chain data see the companion Base mainnet census.
Composition
6,000 agents, 1… See the full description on the dataset page: https://huggingface.co/datasets/rsoft-latam/erc8004-simulated-agents.AEMO_simulated_trade_impact
AEMO Simulated Trade — Impact-Aware Episodes
Impact-aware battery trading episodes from Australia's National Electricity
Market (AEMO/NEM), generated under an endogenous market-impact model for
retraining a Decision Transformer (DT) to avoid self-impact.
This is the Phase 4 companion dataset to
mrvictoru/AEMO_simulated_trade:
where the original is price-taking, every episode here is rolled out with a
piecewise-linear merit-order market-impact model enabled — the battery's own… See the full description on the dataset page: https://huggingface.co/datasets/mrvictoru/AEMO_simulated_trade_impact.anais112-simulated-backgroundrare_disease_clinical_profiles_simulatedCOMPUTER_SIMULATED_PLANT_DESIGN_for_WASTE_MINIMIZATION_POLLUTION_PREVENTIONhttps://drive.google.com/file/d/1O8-cxFCJWZ6n5NSQ5ZNnPo0ftm8jIPBL/view?usp=drivesdk
simulated_bank_data_2012_2026
Simulated Bank Marketing Dataset (2012-2026)
Description
Simulated extension of UCI Bank Marketing dataset for predicting term deposit subscriptions. Original data from 2008-2010; simulated for 2012-2026 with similar distributions.
Data Source
Based on UCI ML Repository: https://archive.ics.uci.edu/dataset/222/bank+marketing
41,188 instances, 21 features. Quinlan, J. (1987). Credit Approval [Dataset]. UCI Machine Learning Repository.… See the full description on the dataset page: https://huggingface.co/datasets/supersam7/simulated_bank_data_2012_2026.HVDC-SIMULATED-FAULTS-FINAL-COMBINED
