CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01annahbanannah /synthetic-math-toolcall-deception Synthetic Math Tool-Call Deception 200 paired multi-turn math-assistant trajectories (400 rows) for evaluating deception detectors on mid-trajectory tool-call misreporting. Each trajectory: a system prompt instructs the model to compute via an execute_python tool under a stated tool-call limit, and requires every call to carry a running call_index argument (1 for the first call, 2 for the second, …). The platform enforcing the limit is said to only see the reported call_index… See the full description on the dataset page: https://huggingface.co/datasets/annahbanannah/synthetic-math-toolcall-deception.tabulartext-classificationn<1K0 likes6.3k downloads2mo agoHugging Face02riotu-lab /Synthetic-UAV-Flight-Trajectories UAV Trajectory Dataset Summary This dataset comprises over 5000 random UAV (Unmanned Aerial Vehicle) trajectories collected over 20 hours of flight time. It is intended for training AI models such as trajectory prediction applications. The dataset is generated through an automated pipeline for the creation and preprocessing of UAV synthetic trajectories, making it ready for direct AI model training. Data Description The dataset features parameterized… See the full description on the dataset page: https://huggingface.co/datasets/riotu-lab/Synthetic-UAV-Flight-Trajectories.tabular100K<n<1M17 likes5.1k downloads2y agoHugging Face03sophia1ch /zendo-synthetic-data Zendo Synthetic Visual Reasoning Dataset Synthetic Zendo-style scenes with associated rules and per-scene tensor representations. Each scene either follows ("positive", label=1) or violates ("negative", label=0) a rule that is given in natural language and as a Prolog query. Splits split scenes train 56475 test 3344 rules total 3439 Layout images/<split>/<batch>/<rule_id>/<scene_id>.png — rendered scene… See the full description on the dataset page: https://huggingface.co/datasets/sophia1ch/zendo-synthetic-data.imageimage-classification10K<n<100K1 likes4.7k downloads4mo agoHugging Face04yxchng /laion_synthetic_filtered_large_part3image10M<n<100M0 likes3.2k downloads3y agoHugging Face05yxchng /laion_synthetic_filtered_large_part1image10M<n<100M2 likes3.2k downloads3y agoHugging Face06yxchng /laion_synthetic_filtered_large_part2image10M<n<100M0 likes3k downloads3y agoHugging Face07gretelai /synthetic_pii_finance_multilingual Image generated by DALL-E. See prompt for more details 💼 📊 Synthetic Financial Domain Documents with PII Labels gretelai/synthetic_pii_finance_multilingual is a dataset of full length synthetic financial documents containing Personally Identifiable Information (PII), generated using Gretel Navigator and released under Apache 2.0. This dataset is designed to assist with the following use cases: 🏷️ Training NER (Named Entity Recognition) models to detect and label PII in… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/synthetic_pii_finance_multilingual.tabulartext-classification10K<n<100K81 likes2.2k downloads2y agoHugging Face08physicl /synthetic-bathroom-dataset-for-robotic-perception Synthetic Bathroom Dataset for Robotic Perception Generated by datapack-import.ts This dataset mirrors public data-pack render outputs from Physicl. Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under /data; image_path stores the relative repository path and data_commit_sha pins the Hugging Face dataset commit used by those URLs. Files are uploaded as downloaded unless optional PNG recompression is enabled by the… See the full description on the dataset page: https://huggingface.co/datasets/physicl/synthetic-bathroom-dataset-for-robotic-perception.imagen<1K0 likes2k downloads3mo agoHugging Face09physicl /synthetic-living-room-dataset-for-robotic-perception Synthetic Living Room Dataset for Robotic Perception Generated by datapack-import.ts This dataset mirrors public data-pack render outputs from Physicl. Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under /data; image_path stores the relative repository path and data_commit_sha pins the Hugging Face dataset commit used by those URLs. Files are uploaded as downloaded unless optional PNG recompression is enabled by… See the full description on the dataset page: https://huggingface.co/datasets/physicl/synthetic-living-room-dataset-for-robotic-perception.imagen<1K0 likes1.7k downloads3mo agoHugging Face10vidore /syntheticDocQA_artificial_intelligence_test_beirBEIR version of vidore/syntheticDocQA_artificial_intelligence_test. imagedocument-question-answering1K<n<10K0 likes1.7k downloads1y agoHugging Face11vidore /syntheticDocQA_energy_test_beirBEIR version of vidore/syntheticDocQA_energy_test. imagedocument-question-answering1K<n<10K0 likes1.7k downloads1y agoHugging Face12vidore /syntheticDocQA_healthcare_industry_test_beirBEIR version of vidore/syntheticDocQA_healthcare_industry_test. imagedocument-question-answering1K<n<10K0 likes1.7k downloads1y agoHugging Face13vidore /syntheticDocQA_government_reports_test_beirBEIR version of vidore/syntheticDocQA_government_reports_test. imagedocument-question-answering1K<n<10K1 likes1.7k downloads1y agoHugging Face14yxchng /laion_synthetic_filtered_large_part4image10M<n<100M0 likes1.5k downloads3y agoHugging Face15mteb /syntheticDocQA_artificial_intelligence_test_beirBEIR version of vidore/syntheticDocQA_artificial_intelligence_test. imagedocument-question-answering1K<n<10K0 likes992 downloads7mo agoHugging Face16mteb /syntheticDocQA_healthcare_industry_test_beirBEIR version of vidore/syntheticDocQA_healthcare_industry_test. imagedocument-question-answering1K<n<10K0 likes981 downloads7mo agoHugging Face17GOD111111111 /synthetic-timeseries-data cruscy data — evaluation sample Three full days of real crypto market microstructure (Binance spot), prepared for public evaluation: absolute prices, dates, and the instrument are withheld — the shape of the day (tick-by-tick relative price, normalized volumes, trade side, book imbalance) is fully preserved. The full feed — 27+ streams (raw L2 depth, 1-second trade tape, order-book metrics, derived features, regime labels) with SQL console, backtest runner and MCP access for AI… See the full description on the dataset page: https://huggingface.co/datasets/GOD111111111/synthetic-timeseries-data.tabular1B<n<10B0 likes978 downloads2h agoHugging Face18mteb /syntheticDocQA_government_reports_test_beirBEIR version of vidore/syntheticDocQA_government_reports_test. imagedocument-question-answering1K<n<10K0 likes948 downloads7mo agoHugging Face19mteb /syntheticDocQA_energy_test_beirBEIR version of vidore/syntheticDocQA_energy_test. imagedocument-question-answering1K<n<10K0 likes946 downloads7mo agoHugging Face20jinaai /airbnb-synthetic-retrieval_beirThis is a copy of https://huggingface.co/datasets/jinaai/airbnb-synthetic-retrieval reformatted into the BEIR format. For any further information like license, please refer to the original dataset. Disclaimer This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data (at)… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/airbnb-synthetic-retrieval_beir.image1K<n<10K0 likes515 downloads1y agoHugging Face21aaaaliou /pi-synthetic Coding agent session traces for aaaaliou/pi-synthetic This dataset contains redacted coding agent session traces collected while working on git@github.com:aliou/pi-synthetic.git. The traces were exported with pi-share-hf from a local pi workspace and filtered to keep only sessions that passed deterministic redaction and LLM review. Data description Each *.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where each line is a structured… See the full description on the dataset page: https://huggingface.co/datasets/aaaaliou/pi-synthetic.tabulartext-generationn<1K0 likes491 downloads5mo agoHugging Face22lum-ai /metal-python-synthetic-explanations-gpt4-graphcodeberttabular1M<n<10M0 likes478 downloads3y agoHugging Face23jinaai /tweet-stock-synthetic-retrieval_beirThis is a copy of https://huggingface.co/datasets/jinaai/tweet-stock-synthetic-retrieval reformatted into the BEIR format. For any further information like license, please refer to the original dataset. Disclaimer This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/tweet-stock-synthetic-retrieval_beir.image1K<n<10K0 likes453 downloads1y agoHugging Face24hotchpotch /wikipedia-multilingual-synthetic-ir-query wikipedia-multilingual-synthetic-ir-query This dataset contains multilingual Wikipedia-derived synthetic query-document pairs for information retrieval training. It was created with the query-crafter-multilingual model, which generates search-like queries from Wikipedia text. The current release contains two different retrieval settings: short_doc: pairs of (query, short document) long_doc: pairs of (query, long document) These two subsets are not generated in the same way… See the full description on the dataset page: https://huggingface.co/datasets/hotchpotch/wikipedia-multilingual-synthetic-ir-query.tabulartext-retrieval10M<n<100M0 likes449 downloads3mo agoHugging Face25NurErtug /mmlu-synthetictabular10K<n<100K0 likes446 downloads2mo agoHugging Face26PrimeIntellect /SYNTHETIC-2 SYNTHETIC-2 SYNTHETIC-2 is an open reasoning dataset spanning a variety of math, coding and general reasoning tasks along with reasoning traces generated in a collaborative manner. The dataset contains both high quality reasoning traces from Deepseek-R1-0528 ideally suited for SFT, as well as multiple reasoning traces from smaller models which can be used for difficulty estimation. To read more about our data collection approach, check out our blog post. We release the following… See the full description on the dataset page: https://huggingface.co/datasets/PrimeIntellect/SYNTHETIC-2.tabular10K<n<100K16 likes437 downloads1y agoHugging Face27yxchng /ccs_synthetic_filtered_largeimage10M<n<100M0 likes430 downloads3y agoHugging Face28apockill /myarm-8-synthetic-cube-to-cup-largeThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": null, "total_episodes": 874, "total_frames": 421190, "total_tasks": 1, "total_videos": 1748, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:874" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/apockill/myarm-8-synthetic-cube-to-cup-large.tabularrobotics100K<n<1M0 likes413 downloads2y agoHugging Face29Alignment-Lab-AI /synthetic-bn-subseqtabular1M<n<10M0 likes407 downloads1y agoHugging Face30MoritzLaurer /synthetic_zeroshot_mixtral_v0.1tabular1M<n<10M9 likes366 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.