CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01tahabou /HVDC-SIMULATED-FAULTS0 likes619 downloads1y agoHugging Face02SimulatedScience /igsm-med-120Mproblems Overview This repository contains datasets to partially reproduce the paper Physics of Language Models: Part 2.1.These are artefacts of our independent reproduction effort. Models trained on this data are available on Hugginf Face here. Content Main training dataset: ./igsm_train_120M ~120 million iGSM problems difficulties (operations needed to solve each problem): 1-15 variable dependency probe training dataset: ./probes/dep/vprobe_dep_train_20k dep probe… See the full description on the dataset page: https://huggingface.co/datasets/SimulatedScience/igsm-med-120Mproblems.10M<n<100M0 likes579 downloads1mo agoHugging Face03golmschenk /general_light_curve_benchmark_dataset_collection_roman_simulated_variable_star_datasetimage1K<n<10K0 likes435 downloads1y agoHugging Face04DeveloperMindset123 /CMAPSS_Jet_Engine_Simulated_Datadocument100K<n<1M0 likes323 downloads11mo agoHugging Face05SCATSVAD /SCATSVAD-Simulated-Data SCATSVAD Dataset This repository contains the SCATSVAD dataset, including training, evaluation, and test sets. The dataset is split into multiple files, most of them with a size of 40GB. Dataset Structure Training Set (train) The training set consists of five compressed files: train_dataset.tar.gz00 train_dataset.tar.gz01 train_dataset.tar.gz02 train_dataset.tar.gz03 train_dataset.tar.gz04 - 27GB Evaluation Set (eval) The evaluation set contains… See the full description on the dataset page: https://huggingface.co/datasets/SCATSVAD/SCATSVAD-Simulated-Data.0 likes317 downloads2y agoHugging Face06while-ai /tau2-simulated tau2 Simulated Training Set Made with the whileai SDK · Collections: Simulation, Start here: foundational post-training datasets The training set that took a base model from 5% to 30% on tau2-bench telecom, made from nothing but the agent's tool list and policy. If you build a customer-facing agent, you already have the two files this dataset was made from: the tools it can call and the policy it follows. The whileai SDK turned those into 1,057 graded conversations across the… See the full description on the dataset page: https://huggingface.co/datasets/while-ai/tau2-simulated.texttext-generation1K<n<10K0 likes145 downloads4d agoHugging Face07shuhaibmehri /UserBehavioralDivergence-simulated-conversationstext100K<n<1M2 likes127 downloads4mo agoHugging Face08hotchpotch /fineweb-ir-simulated-search-queries fineweb-ir-simulated-search-queries An English web-retrieval dataset built by generating simulated search queries for FineWeb-style positive target documents. This dataset contains English query-document pairs derived from HuggingFaceFW/fineweb-edu. Each row is designed so that the associated document is a positive retrieval target for the generated query. The queries were generated with a Qwen3.5-35B-A3B model quantized to NVFP4 and served through OpenAI-compatible vLLM serve… See the full description on the dataset page: https://huggingface.co/datasets/hotchpotch/fineweb-ir-simulated-search-queries.text1M<n<10M0 likes126 downloads5mo agoHugging Face09hotchpotch /arxiv-ir-simulated-search-queries arxiv-ir-simulated-search-queries An arXiv retrieval dataset with more than 2.8 million simulated specialist search queries and paper-level positive targets. This dataset contains 2,875,637 query-document pairs derived from arXiv title-and-abstract records. Each row is designed so that the associated arXiv paper record is a positive retrieval target for the generated query. The queries were generated with a Qwen3.5-35B-A3B model quantized to NVFP4, and were intentionally… See the full description on the dataset page: https://huggingface.co/datasets/hotchpotch/arxiv-ir-simulated-search-queries.text1M<n<10M0 likes118 downloads6mo agoHugging Face10rorymcgowan /simulated_runs_50image10K<n<100K0 likes112 downloads6mo agoHugging Face11mrvictoru /AEMO_simulated_trade AEMO Battery Trading Dataset Note (Aug 2026): The SDP-teacher trajectory dataset (data/aemo_dt_sdp/ in the repo) is now the preferred training data for the shipped model. The original FCAS dataset below was the training source for the Jul 2026 v2 pretrained model and the GRPO study. Both are historical — the Stage C standalone DT (models/aemo/dt/aemo_dt_sdp_jtsoc_fullcorpus.pt) was trained on SDP-teacher trajectories with J_t(soc) RTG prompts. Files File… See the full description on the dataset page: https://huggingface.co/datasets/mrvictoru/AEMO_simulated_trade.tabular10M<n<100M0 likes105 downloads1mo agoHugging Face12hotchpotch /wikipedia-english-ir-simulated-search-queries wikipedia-english-ir-simulated-search-queries An English Wikipedia retrieval dataset with more than 29 million simulated search queries and paragraph-level positive targets. This dataset contains 29,366,101 English query-document pairs derived from Wikipedia. Each row is designed so that the associated Wikipedia paragraph is a positive retrieval target for the generated query. The queries were generated with a Qwen3.5-35B-A3B model quantized to NVFP4, and were intentionally… See the full description on the dataset page: https://huggingface.co/datasets/hotchpotch/wikipedia-english-ir-simulated-search-queries.text10M<n<100M0 likes102 downloads4mo agoHugging Face13jablonkagroup /simulated_spectratext1M<n<10M0 likes90 downloads1mo agoHugging Face14hotchpotch /pubmed-abstract-ir-simulated-search-queries pubmed-abstract-ir-simulated-search-queries A PubMed retrieval dataset with simulated specialist search queries and abstract-level positive targets. This dataset contains 2,355,329 query-document pairs derived from PubMed title-and-abstract records. Each row is designed so that the associated PubMed record is a positive retrieval target for the generated query. The queries were generated with a Qwen3.5-35B-A3B model quantized to NVFP4, and were intentionally constructed to… See the full description on the dataset page: https://huggingface.co/datasets/hotchpotch/pubmed-abstract-ir-simulated-search-queries.text1M<n<10M0 likes81 downloads6mo agoHugging Face15strova-ai /developer-productivity-simulated-behavioral-data Synthetic AI Developer Productivity Dataset — Behavioral + Cognitive Simulation A synthetic data generation resource for modeling behavioral and cognitive dynamics in developers. 📘 About This Dataset This dataset simulates productivity data from AI-assisted software developers. It blends behavioral signals, physiological inputs, and productivity metrics to explore the nuanced relationships between deep work, distractions, caffeine, AI usage, and cognitive strain.… See the full description on the dataset page: https://huggingface.co/datasets/strova-ai/developer-productivity-simulated-behavioral-data.tabular1K<n<10K503 likes75 downloads1y agoHugging Face16hotchpotch /ccnews-ir-simulated-search-queries ccnews-ir-simulated-search-queries An English news-retrieval dataset with 1.84 million simulated search queries paired with positive CC-News-style document targets. This dataset contains 1,839,547 English query-document pairs derived from the English subset of multilingual CC-News. Each row is designed so that the associated news document is a positive retrieval target for the generated query. The queries were generated with a Qwen3.5-35B-A3B model quantized to NVFP4 and… See the full description on the dataset page: https://huggingface.co/datasets/hotchpotch/ccnews-ir-simulated-search-queries.text1M<n<10M0 likes60 downloads5mo agoHugging Face17mrvictoru /AEMO_simulated_trade_sdp AEMO SDP-Teacher Trajectories Offline trajectories for Decision Transformer training, generated by replaying the honest SDP/MPC executor on historical Australian NEM (AEMO) market data. These are the teacher trajectories from the energydecision research codebase. Each row is a single 5-minute market interval with a self-consistent (normalized observation, 9-dim action, reward) triple: the action is what the honest optimal planner dispatched, the reward is what it earned, and the… See the full description on the dataset page: https://huggingface.co/datasets/mrvictoru/AEMO_simulated_trade_sdp.tabular1M<n<10M0 likes59 downloads12d agoHugging Face18tahabou /HVDC-SIMULATED-FAULTS-FINAL1 likes58 downloads1y agoHugging Face19yunfanlu /UNIINR-Fastec-Simulated-Dataset0 likes45 downloads11mo agoHugging Face20PrabhakarSai /after_visit_summary_simulated_edits Dataset Card: AVS edits Dataset Dataset Summary The AVS edits dataset is designed to support human feedback research in for clinical summarization. It contains synthetic edit feedback generated by large language models (LLMs) to improve the factual consistency and quality of summaries. The dataset includes training, evaluation, and test splits with specific fields for modeling and evaluation tasks. Dataset Structure Train Split Keys: article: The… See the full description on the dataset page: https://huggingface.co/datasets/PrabhakarSai/after_visit_summary_simulated_edits.summarization1K<n<10K0 likes40 downloads2y agoHugging Face21joonhyukseow /SimulatedScatteringimage1K<n<10K0 likes28 downloads8mo agoHugging Face22Felixdu11 /tracked_CIRS_simulated Tracked CIRS simulated CIRS phantom RF acquisition with a timestamped probe-tracking stream. This repository contains both the original raw source files and the converted OpenH-RF/ZEA dataset. Files Raw source data: raw/cirs_imaging.hdf5: original imaging HDF5 source. raw/cirs_tracking.ts: original timestamped probe tracking stream. OpenH-RF/ZEA converted data: zea/cirs_imaging_zea.hdf5: RF data, scan metadata, and probe pose metadata. zea/config.yaml:… See the full description on the dataset page: https://huggingface.co/datasets/Felixdu11/tracked_CIRS_simulated.0 likes27 downloads12d agoHugging Face23Abhijnan /wtd_simulated_datatabular10K<n<100K0 likes26 downloads1y agoHugging Face24karankhatavkar /ak47-acoustic-rul-simulated AK-47 Acoustic Run-to-Failure (RUL) Simulation Dataset A synthetic Run-to-Failure dataset for Remaining Useful Life (RUL) estimation of an AK-47's recoil spring from gunshot audio. Because real run-to-failure recordings of a wearing firearm are practically impossible to collect, this dataset is generated by a physics-based Digital Twin that takes a small set of real, healthy gunshot recordings and mathematically simulates the acoustic signature of mechanical wear over thousands… See the full description on the dataset page: https://huggingface.co/datasets/karankhatavkar/ak47-acoustic-rul-simulated.audioaudio-classification1 likes26 downloads4mo agoHugging Face25niklashcs /robco_simulatedtabularn<1K0 likes23 downloads10mo agoHugging Face26WhissleAI /speech-simulated-medical-examsgated Speech Simulated Medical Exams Simulated patient-physician medical exam conversations with rich speech metadata annotations. Built for training single-step ASR models that transcribe and annotate multiple concepts simultaneously, including speaker changes, emotions, intents, and roles. Dataset Details Property Value Examples 25,706 Language English Audio 16 kHz WAV Source Simulated medical interviews (respiratory focus) Features… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/speech-simulated-medical-exams.audioautomatic-speech-recognition10K<n<100K4 likes21 downloads4mo agoHugging Face27SOTAagi2030 /HarborMap-SimulatedRoutes HarborMap Simulated Routes This dataset contains synthetic vessel-routing dialogues assembled for the port-planning sandbox. Reuse register Material License Status Distribution score Notice actions Navigation prompt library MPL-2.0 Approved 9 1 Beacon interpolation component BSD-2-Clause Approved 9 1 Dockside validation samples Apache-2.0 Approved 9 2 Retired charting draft CC-BY-NC-4.0 Withdrawn 12 0 This register is the authoritative record… See the full description on the dataset page: https://huggingface.co/datasets/SOTAagi2030/HarborMap-SimulatedRoutes.0 likes21 downloads2d agoHugging Face28sujalappa /simulated_rirs_dataset Simulated Rirs Dataset Dataset Description This dataset contains 400 samples organized across multiple splits and 4 subsets. The dataset includes audio data. Dataset Structure Subsets This dataset includes the following subsets: original: 100 samples train: 100 samples largeroom: 100 samples train: 100 samples mediumroom: 100 samples train: 100 samples smallroom: 100 samples train: 100 samples Usage Load specific subset and… See the full description on the dataset page: https://huggingface.co/datasets/sujalappa/simulated_rirs_dataset.audioautomatic-speech-recognitionn<1K0 likes19 downloads7mo agoHugging Face29sguajardo799 /simulated-binaural-speech-directivity Simulated Binaural Speech Source Directivity Dataset This dataset contains 96,000 simulated binaural speech recordings generated for the study of speech source directivity classification. The dataset is designed to support the binary classification of whether a speech source is oriented toward or away from a listener. The recordings were generated under controlled acoustic and spatial conditions using an acoustic simulation workflow based on RAVEN… See the full description on the dataset page: https://huggingface.co/datasets/sguajardo799/simulated-binaural-speech-directivity.audioaudio-classification10K<n<100K0 likes18 downloads5mo agoHugging Face30stabletoolbench /real_simulated_compareThis a new test set for comparing real and simulated APIs in StableToolBench-MirrorAPI. This dataset is in the ToolBench/StableToolBench test set format and you can directly use it in ToolBench or StableToolBench. 0 likes17 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.