CoolFace
22 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01DeveloperMindset123 /CMAPSS_Jet_Engine_Simulated_Datadocument100K<n<1M0 likes323 downloads11mo agoHugging Face02while-ai /tau2-simulated tau2 Simulated Training Set Made with the whileai SDK · Collections: Simulation, Start here: foundational post-training datasets The training set that took a base model from 5% to 30% on tau2-bench telecom, made from nothing but the agent's tool list and policy. If you build a customer-facing agent, you already have the two files this dataset was made from: the tools it can call and the policy it follows. The whileai SDK turned those into 1,057 graded conversations across the… See the full description on the dataset page: https://huggingface.co/datasets/while-ai/tau2-simulated.texttext-generation1K<n<10K0 likes145 downloads4d agoHugging Face03shuhaibmehri /UserBehavioralDivergence-simulated-conversationstext100K<n<1M2 likes127 downloads4mo agoHugging Face04hotchpotch /fineweb-ir-simulated-search-queries fineweb-ir-simulated-search-queries An English web-retrieval dataset built by generating simulated search queries for FineWeb-style positive target documents. This dataset contains English query-document pairs derived from HuggingFaceFW/fineweb-edu. Each row is designed so that the associated document is a positive retrieval target for the generated query. The queries were generated with a Qwen3.5-35B-A3B model quantized to NVFP4 and served through OpenAI-compatible vLLM serve… See the full description on the dataset page: https://huggingface.co/datasets/hotchpotch/fineweb-ir-simulated-search-queries.text1M<n<10M0 likes126 downloads5mo agoHugging Face05hotchpotch /arxiv-ir-simulated-search-queries arxiv-ir-simulated-search-queries An arXiv retrieval dataset with more than 2.8 million simulated specialist search queries and paper-level positive targets. This dataset contains 2,875,637 query-document pairs derived from arXiv title-and-abstract records. Each row is designed so that the associated arXiv paper record is a positive retrieval target for the generated query. The queries were generated with a Qwen3.5-35B-A3B model quantized to NVFP4, and were intentionally… See the full description on the dataset page: https://huggingface.co/datasets/hotchpotch/arxiv-ir-simulated-search-queries.text1M<n<10M0 likes118 downloads6mo agoHugging Face06mrvictoru /AEMO_simulated_trade AEMO Battery Trading Dataset Note (Aug 2026): The SDP-teacher trajectory dataset (data/aemo_dt_sdp/ in the repo) is now the preferred training data for the shipped model. The original FCAS dataset below was the training source for the Jul 2026 v2 pretrained model and the GRPO study. Both are historical — the Stage C standalone DT (models/aemo/dt/aemo_dt_sdp_jtsoc_fullcorpus.pt) was trained on SDP-teacher trajectories with J_t(soc) RTG prompts. Files File… See the full description on the dataset page: https://huggingface.co/datasets/mrvictoru/AEMO_simulated_trade.tabular10M<n<100M0 likes105 downloads1mo agoHugging Face07hotchpotch /wikipedia-english-ir-simulated-search-queries wikipedia-english-ir-simulated-search-queries An English Wikipedia retrieval dataset with more than 29 million simulated search queries and paragraph-level positive targets. This dataset contains 29,366,101 English query-document pairs derived from Wikipedia. Each row is designed so that the associated Wikipedia paragraph is a positive retrieval target for the generated query. The queries were generated with a Qwen3.5-35B-A3B model quantized to NVFP4, and were intentionally… See the full description on the dataset page: https://huggingface.co/datasets/hotchpotch/wikipedia-english-ir-simulated-search-queries.text10M<n<100M0 likes102 downloads4mo agoHugging Face08jablonkagroup /simulated_spectratext1M<n<10M0 likes90 downloads1mo agoHugging Face09hotchpotch /pubmed-abstract-ir-simulated-search-queries pubmed-abstract-ir-simulated-search-queries A PubMed retrieval dataset with simulated specialist search queries and abstract-level positive targets. This dataset contains 2,355,329 query-document pairs derived from PubMed title-and-abstract records. Each row is designed so that the associated PubMed record is a positive retrieval target for the generated query. The queries were generated with a Qwen3.5-35B-A3B model quantized to NVFP4, and were intentionally constructed to… See the full description on the dataset page: https://huggingface.co/datasets/hotchpotch/pubmed-abstract-ir-simulated-search-queries.text1M<n<10M0 likes81 downloads6mo agoHugging Face10hotchpotch /ccnews-ir-simulated-search-queries ccnews-ir-simulated-search-queries An English news-retrieval dataset with 1.84 million simulated search queries paired with positive CC-News-style document targets. This dataset contains 1,839,547 English query-document pairs derived from the English subset of multilingual CC-News. Each row is designed so that the associated news document is a positive retrieval target for the generated query. The queries were generated with a Qwen3.5-35B-A3B model quantized to NVFP4 and… See the full description on the dataset page: https://huggingface.co/datasets/hotchpotch/ccnews-ir-simulated-search-queries.text1M<n<10M0 likes60 downloads5mo agoHugging Face11mrvictoru /AEMO_simulated_trade_sdp AEMO SDP-Teacher Trajectories Offline trajectories for Decision Transformer training, generated by replaying the honest SDP/MPC executor on historical Australian NEM (AEMO) market data. These are the teacher trajectories from the energydecision research codebase. Each row is a single 5-minute market interval with a self-consistent (normalized observation, 9-dim action, reward) triple: the action is what the honest optimal planner dispatched, the reward is what it earned, and the… See the full description on the dataset page: https://huggingface.co/datasets/mrvictoru/AEMO_simulated_trade_sdp.tabular1M<n<10M0 likes59 downloads12d agoHugging Face12Abhijnan /wtd_simulated_datatabular10K<n<100K0 likes26 downloads1y agoHugging Face13WhissleAI /speech-simulated-medical-examsgated Speech Simulated Medical Exams Simulated patient-physician medical exam conversations with rich speech metadata annotations. Built for training single-step ASR models that transcribe and annotate multiple concepts simultaneously, including speaker changes, emotions, intents, and roles. Dataset Details Property Value Examples 25,706 Language English Audio 16 kHz WAV Source Simulated medical interviews (respiratory focus) Features… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/speech-simulated-medical-exams.audioautomatic-speech-recognition10K<n<100K4 likes21 downloads4mo agoHugging Face14sujalappa /simulated_rirs_dataset Simulated Rirs Dataset Dataset Description This dataset contains 400 samples organized across multiple splits and 4 subsets. The dataset includes audio data. Dataset Structure Subsets This dataset includes the following subsets: original: 100 samples train: 100 samples largeroom: 100 samples train: 100 samples mediumroom: 100 samples train: 100 samples smallroom: 100 samples train: 100 samples Usage Load specific subset and… See the full description on the dataset page: https://huggingface.co/datasets/sujalappa/simulated_rirs_dataset.audioautomatic-speech-recognitionn<1K0 likes19 downloads7mo agoHugging Face15pcy12345BSU /binary-function-code-simulatedtextn<1K1 likes12 downloads1y agoHugging Face16rsoft-latam /erc8004-simulated-agents ERC-8004 Simulated Agents — labeled synthetic dataset (6,000 agents) ⚠️ This dataset is fully synthetic. No public labeled dataset of malicious ERC-8004 agents exists (the standard reached mainnet in 2026 and exposes no trust label), so this dataset simulates the feature distributions the three ERC-8004 registries would expose, for training/evaluating trustworthiness models. For real on-chain data see the companion Base mainnet census. Composition 6,000 agents, 1… See the full description on the dataset page: https://huggingface.co/datasets/rsoft-latam/erc8004-simulated-agents.tabular1K<n<10K0 likes9 downloads2mo agoHugging Face17mrvictoru /AEMO_simulated_trade_impact AEMO Simulated Trade — Impact-Aware Episodes Impact-aware battery trading episodes from Australia's National Electricity Market (AEMO/NEM), generated under an endogenous market-impact model for retraining a Decision Transformer (DT) to avoid self-impact. This is the Phase 4 companion dataset to mrvictoru/AEMO_simulated_trade: where the original is price-taking, every episode here is rolled out with a piecewise-linear merit-order market-impact model enabled — the battery's own… See the full description on the dataset page: https://huggingface.co/datasets/mrvictoru/AEMO_simulated_trade_impact.tabular10M<n<100M0 likes8 downloads2mo agoHugging Face18scidata-hub /anais112-simulated-backgroundtabular1K<n<10K0 likes7 downloads3mo agoHugging Face19Shoriful025 /rare_disease_clinical_profiles_simulatedtabularn<1K0 likes5 downloads9mo agoHugging Face20introvoyz041 /COMPUTER_SIMULATED_PLANT_DESIGN_for_WASTE_MINIMIZATION_POLLUTION_PREVENTIONhttps://drive.google.com/file/d/1O8-cxFCJWZ6n5NSQ5ZNnPo0ftm8jIPBL/view?usp=drivesdk textn<1K0 likes4 downloads1y agoHugging Face21supersam7 /simulated_bank_data_2012_2026 Simulated Bank Marketing Dataset (2012-2026) Description Simulated extension of UCI Bank Marketing dataset for predicting term deposit subscriptions. Original data from 2008-2010; simulated for 2012-2026 with similar distributions. Data Source Based on UCI ML Repository: https://archive.ics.uci.edu/dataset/222/bank+marketing 41,188 instances, 21 features. Quinlan, J. (1987). Credit Approval [Dataset]. UCI Machine Learning Repository.… See the full description on the dataset page: https://huggingface.co/datasets/supersam7/simulated_bank_data_2012_2026.tabular100K<n<1M0 likes3 downloads8mo agoHugging Face22tahabou /HVDC-SIMULATED-FAULTS-FINAL-COMBINEDtabular10K<n<100K0 likes2 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.