simulated
whisper-large-v2-tr-ft-03-04-26-full-ft-50ksamples-simulated-databge-m3-w8a16-simulatedwhisper-large-v2-tr-ft-12-04-26-encoder-only-100ksamples-simulated-datawhisper-large-v2-tr-ft-18-04-26-full-ft-100ksamples-simulated-datasmolvla-dwip-simulatedwhisper-large-v2-tr-ft-31-03-26-encoder-only-50ksamples-simulated-datawhisper-large-v2-tr-ft-24-03-26-encoder-only-25ksamples-simulatedwhisper-medium-tr-ft-17-04-26-encoder-only-100ksamples-simulated-data
HVDC-SIMULATED-FAULTSigsm-med-120Mproblems
Overview
This repository contains datasets to partially reproduce the paper Physics of Language Models: Part 2.1.These are artefacts of our independent reproduction effort.
Models trained on this data are available on Hugginf Face here.
Content
Main training dataset: ./igsm_train_120M
~120 million iGSM problems
difficulties (operations needed to solve each problem): 1-15
variable dependency probe training dataset: ./probes/dep/vprobe_dep_train_20k
dep probe… See the full description on the dataset page: https://huggingface.co/datasets/SimulatedScience/igsm-med-120Mproblems.general_light_curve_benchmark_dataset_collection_roman_simulated_variable_star_datasetCMAPSS_Jet_Engine_Simulated_DataSCATSVAD-Simulated-Data
SCATSVAD Dataset
This repository contains the SCATSVAD dataset, including training, evaluation, and test sets. The dataset is split into multiple files, most of them with a size of 40GB.
Dataset Structure
Training Set (train)
The training set consists of five compressed files:
train_dataset.tar.gz00
train_dataset.tar.gz01
train_dataset.tar.gz02
train_dataset.tar.gz03
train_dataset.tar.gz04 - 27GB
Evaluation Set (eval)
The evaluation set contains… See the full description on the dataset page: https://huggingface.co/datasets/SCATSVAD/SCATSVAD-Simulated-Data.tau2-simulated
tau2 Simulated Training Set
Made with the whileai SDK · Collections: Simulation, Start here: foundational post-training datasets
The training set that took a base model from 5% to 30% on tau2-bench
telecom, made from nothing but the agent's tool list and policy.
If you build a customer-facing agent, you already have the two files this
dataset was made from: the tools it can call and the policy it follows.
The whileai SDK turned those into 1,057 graded conversations across the… See the full description on the dataset page: https://huggingface.co/datasets/while-ai/tau2-simulated.
