datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
HVDC-SIMULATED-FAULTSigsm-med-120Mproblems
Overview
This repository contains datasets to partially reproduce the paper Physics of Language Models: Part 2.1.These are artefacts of our independent reproduction effort.
Models trained on this data are available on Hugginf Face here.
Content
Main training dataset: ./igsm_train_120M
~120 million iGSM problems
difficulties (operations needed to solve each problem): 1-15
variable dependency probe training dataset: ./probes/dep/vprobe_dep_train_20k
dep probe… See the full description on the dataset page: https://huggingface.co/datasets/SimulatedScience/igsm-med-120Mproblems.general_light_curve_benchmark_dataset_collection_roman_simulated_variable_star_datasetCMAPSS_Jet_Engine_Simulated_DataSCATSVAD-Simulated-Data
SCATSVAD Dataset
This repository contains the SCATSVAD dataset, including training, evaluation, and test sets. The dataset is split into multiple files, most of them with a size of 40GB.
Dataset Structure
Training Set (train)
The training set consists of five compressed files:
train_dataset.tar.gz00
train_dataset.tar.gz01
train_dataset.tar.gz02
train_dataset.tar.gz03
train_dataset.tar.gz04 - 27GB
Evaluation Set (eval)
The evaluation set contains… See the full description on the dataset page: https://huggingface.co/datasets/SCATSVAD/SCATSVAD-Simulated-Data.tau2-simulated
tau2 Simulated Training Set
Made with the whileai SDK · Collections: Simulation, Start here: foundational post-training datasets
The training set that took a base model from 5% to 30% on tau2-bench
telecom, made from nothing but the agent's tool list and policy.
If you build a customer-facing agent, you already have the two files this
dataset was made from: the tools it can call and the policy it follows.
The whileai SDK turned those into 1,057 graded conversations across the… See the full description on the dataset page: https://huggingface.co/datasets/while-ai/tau2-simulated.UserBehavioralDivergence-simulated-conversationsfineweb-ir-simulated-search-queries
fineweb-ir-simulated-search-queries
An English web-retrieval dataset built by generating simulated search queries for FineWeb-style positive target documents.
This dataset contains English query-document pairs derived from HuggingFaceFW/fineweb-edu.
Each row is designed so that the associated document is a positive retrieval target for the generated query.
The queries were generated with a Qwen3.5-35B-A3B model quantized to NVFP4 and served through OpenAI-compatible vLLM serve… See the full description on the dataset page: https://huggingface.co/datasets/hotchpotch/fineweb-ir-simulated-search-queries.arxiv-ir-simulated-search-queries
arxiv-ir-simulated-search-queries
An arXiv retrieval dataset with more than 2.8 million simulated specialist search queries and paper-level positive targets.
This dataset contains 2,875,637 query-document pairs derived from arXiv title-and-abstract records.
Each row is designed so that the associated arXiv paper record is a positive retrieval target for the generated query.
The queries were generated with a Qwen3.5-35B-A3B model quantized to NVFP4, and were intentionally… See the full description on the dataset page: https://huggingface.co/datasets/hotchpotch/arxiv-ir-simulated-search-queries.simulated_runs_50AEMO_simulated_trade
AEMO Battery Trading Dataset
Note (Aug 2026): The SDP-teacher trajectory dataset (data/aemo_dt_sdp/ in the repo) is
now the preferred training data for the shipped model. The original FCAS dataset below was the
training source for the Jul 2026 v2 pretrained model and the GRPO study. Both are historical —
the Stage C standalone DT (models/aemo/dt/aemo_dt_sdp_jtsoc_fullcorpus.pt) was trained on
SDP-teacher trajectories with J_t(soc) RTG prompts.
Files
File… See the full description on the dataset page: https://huggingface.co/datasets/mrvictoru/AEMO_simulated_trade.wikipedia-english-ir-simulated-search-queries
wikipedia-english-ir-simulated-search-queries
An English Wikipedia retrieval dataset with more than 29 million simulated search queries and paragraph-level positive targets.
This dataset contains 29,366,101 English query-document pairs derived from Wikipedia.
Each row is designed so that the associated Wikipedia paragraph is a positive retrieval target for the generated query.
The queries were generated with a Qwen3.5-35B-A3B model quantized to NVFP4, and were intentionally… See the full description on the dataset page: https://huggingface.co/datasets/hotchpotch/wikipedia-english-ir-simulated-search-queries.simulated_spectrapubmed-abstract-ir-simulated-search-queries
pubmed-abstract-ir-simulated-search-queries
A PubMed retrieval dataset with simulated specialist search queries and abstract-level positive targets.
This dataset contains 2,355,329 query-document pairs derived from PubMed title-and-abstract records.
Each row is designed so that the associated PubMed record is a positive retrieval target for the generated query.
The queries were generated with a Qwen3.5-35B-A3B model quantized to NVFP4, and were intentionally constructed to… See the full description on the dataset page: https://huggingface.co/datasets/hotchpotch/pubmed-abstract-ir-simulated-search-queries.developer-productivity-simulated-behavioral-data
Synthetic AI Developer Productivity Dataset — Behavioral + Cognitive Simulation
A synthetic data generation resource for modeling behavioral and cognitive dynamics in developers.
📘 About This Dataset
This dataset simulates productivity data from AI-assisted software developers. It blends behavioral signals, physiological inputs, and productivity metrics to explore the nuanced relationships between deep work, distractions, caffeine, AI usage, and cognitive strain.… See the full description on the dataset page: https://huggingface.co/datasets/strova-ai/developer-productivity-simulated-behavioral-data.ccnews-ir-simulated-search-queries
ccnews-ir-simulated-search-queries
An English news-retrieval dataset with 1.84 million simulated search queries paired with positive CC-News-style document targets.
This dataset contains 1,839,547 English query-document pairs derived from the English subset of multilingual CC-News.
Each row is designed so that the associated news document is a positive retrieval target for the generated query.
The queries were generated with a Qwen3.5-35B-A3B model quantized to NVFP4 and… See the full description on the dataset page: https://huggingface.co/datasets/hotchpotch/ccnews-ir-simulated-search-queries.AEMO_simulated_trade_sdp
AEMO SDP-Teacher Trajectories
Offline trajectories for Decision Transformer training, generated by replaying the honest SDP/MPC executor on historical Australian NEM (AEMO) market data. These are the teacher trajectories from the energydecision research codebase.
Each row is a single 5-minute market interval with a self-consistent (normalized observation, 9-dim action, reward) triple: the action is what the honest optimal planner dispatched, the reward is what it earned, and the… See the full description on the dataset page: https://huggingface.co/datasets/mrvictoru/AEMO_simulated_trade_sdp.HVDC-SIMULATED-FAULTS-FINALUNIINR-Fastec-Simulated-Datasetafter_visit_summary_simulated_edits
Dataset Card: AVS edits Dataset
Dataset Summary
The AVS edits dataset is designed to support human feedback research in for clinical summarization. It contains synthetic edit feedback generated by large language models (LLMs) to improve the factual consistency and quality of summaries. The dataset includes training, evaluation, and test splits with specific fields for modeling and evaluation tasks.
Dataset Structure
Train Split
Keys:
article: The… See the full description on the dataset page: https://huggingface.co/datasets/PrabhakarSai/after_visit_summary_simulated_edits.SimulatedScatteringtracked_CIRS_simulated
Tracked CIRS simulated
CIRS phantom RF acquisition with a timestamped probe-tracking stream.
This repository contains both the original raw source files and the converted OpenH-RF/ZEA dataset.
Files
Raw source data:
raw/cirs_imaging.hdf5: original imaging HDF5 source.
raw/cirs_tracking.ts: original timestamped probe tracking stream.
OpenH-RF/ZEA converted data:
zea/cirs_imaging_zea.hdf5: RF data, scan metadata, and probe pose metadata.
zea/config.yaml:… See the full description on the dataset page: https://huggingface.co/datasets/Felixdu11/tracked_CIRS_simulated.wtd_simulated_dataak47-acoustic-rul-simulated
AK-47 Acoustic Run-to-Failure (RUL) Simulation Dataset
A synthetic Run-to-Failure dataset for Remaining Useful Life (RUL) estimation of an
AK-47's recoil spring from gunshot audio. Because real run-to-failure recordings of a
wearing firearm are practically impossible to collect, this dataset is generated by a
physics-based Digital Twin that takes a small set of real, healthy gunshot recordings and
mathematically simulates the acoustic signature of mechanical wear over thousands… See the full description on the dataset page: https://huggingface.co/datasets/karankhatavkar/ak47-acoustic-rul-simulated.robco_simulatedspeech-simulated-medical-exams
Speech Simulated Medical Exams
Simulated patient-physician medical exam conversations with rich speech metadata annotations. Built for training single-step ASR models that transcribe and annotate multiple concepts simultaneously, including speaker changes, emotions, intents, and roles.
Dataset Details
Property
Value
Examples
25,706
Language
English
Audio
16 kHz WAV
Source
Simulated medical interviews (respiratory focus)
Features… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/speech-simulated-medical-exams.HarborMap-SimulatedRoutes
HarborMap Simulated Routes
This dataset contains synthetic vessel-routing dialogues assembled for the port-planning sandbox.
Reuse register
Material
License
Status
Distribution score
Notice actions
Navigation prompt library
MPL-2.0
Approved
9
1
Beacon interpolation component
BSD-2-Clause
Approved
9
1
Dockside validation samples
Apache-2.0
Approved
9
2
Retired charting draft
CC-BY-NC-4.0
Withdrawn
12
0
This register is the authoritative record… See the full description on the dataset page: https://huggingface.co/datasets/SOTAagi2030/HarborMap-SimulatedRoutes.simulated_rirs_dataset
Simulated Rirs Dataset
Dataset Description
This dataset contains 400 samples organized across multiple splits and 4 subsets.
The dataset includes audio data.
Dataset Structure
Subsets
This dataset includes the following subsets:
original: 100 samples
train: 100 samples
largeroom: 100 samples
train: 100 samples
mediumroom: 100 samples
train: 100 samples
smallroom: 100 samples
train: 100 samples
Usage
Load specific subset and… See the full description on the dataset page: https://huggingface.co/datasets/sujalappa/simulated_rirs_dataset.simulated-binaural-speech-directivity
Simulated Binaural Speech Source Directivity Dataset
This dataset contains 96,000 simulated binaural speech recordings generated for the study of speech source directivity classification. The dataset is designed to support the binary classification of whether a speech source is oriented toward or away from a listener.
The recordings were generated under controlled acoustic and spatial conditions using an acoustic simulation workflow based on RAVEN… See the full description on the dataset page: https://huggingface.co/datasets/sguajardo799/simulated-binaural-speech-directivity.real_simulated_compareThis a new test set for comparing real and simulated APIs in StableToolBench-MirrorAPI. This dataset is in the ToolBench/StableToolBench test set format and you can directly use it in ToolBench or StableToolBench.
