CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SAIRfoundation /equational-theories-selected-problems Equational Theories Selected Problems Update (September 11, 2026) This dataset was updated on September 11, 2026. Main changes: released the official Stage 2 evaluation problems: stage2_evaluation_main (200 problems; ground truth withheld — answer is null until Stage 2 concludes) and stage2_evaluation_research (100 order-5 research problems with no ground truth) added metadata/stage2_evaluation_main.json and metadata/stage2_evaluation_research.json… See the full description on the dataset page: https://huggingface.co/datasets/SAIRfoundation/equational-theories-selected-problems.tabular1K<n<10K11 likes12k downloads15d agoHugging Face02saidutta69 /fable-5-premium 🧠 Fable-5 Premium Dataset 🚀 V2 is out! This dataset has a successor: fable-5-premium-v2 — new users should start there. A rigorously cleaned, high-quality supervised fine-tuning (SFT) dataset built from Claude Fable-5 agent traces. Priorities: Quality > Ease of Access > Quantity 📊 Dataset Overview Property Value Total Records 12,730 Train Split 5,728 (45.0%) Validation Split 318 (2.5%) Test Split 319 (2.5%) Created 2026-07-30… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/fable-5-premium.texttext-generation10K<n<100K134 likes11k downloads13d agoHugging Face03SaifPunjwani /slo-rlvr-results0 likes10k downloads3d agoHugging Face04sailor2 /sea-commoncrawltext100M<n<1B1 likes7.5k downloads2y agoHugging Face05sail /regmix-data RegMix Data Dataset Description The RegMix Data is a curated dataset derived from the Pile-Uncopyrighted, specifically designed for the RegMix paper (https://huggingface.co/papers/2407.01492). This dataset aims to facilitate the automatic identification of high-performing data mixtures for language model pre-training by formulating it as a regression task. Key Features: Size: Approximately 1TB disk space, 250B tokens Distribution: Follows the natural token… See the full description on the dataset page: https://huggingface.co/datasets/sail/regmix-data.text10M<n<100M4 likes6.2k downloads2y agoHugging Face06BangumiBase /saikyounoousamanidomenojinseiwananiwosuru Bangumi Image Base of Saikyou No Ousama, Nidome No Jinsei Wa Nani Wo Suru? This is the image base of bangumi Saikyou no Ousama, Nidome no Jinsei wa Nani wo Suru?, we detected 71 characters, 4913 images in total. The full dataset is here. Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/saikyounoousamanidomenojinseiwananiwosuru.image1K<n<10K0 likes4.7k downloads1y agoHugging Face07sailplane /SWE-bench_Lite_filtered0 likes4.7k downloads2y agoHugging Face08simular-ai /sai-osworld-v2-benchmark-runs Sai on OSWorld-V2 — benchmark runs of record Two complete 108-task runs of the Sai computer-use agent on OSWorld-V2, with full per-task evidence: scores, trajectories, agent runtime logs, evaluator logs, API protocol logs, and run manifests. Run Date Tasks scored Mean score Perfect (1.0) Zeros run1/ 2026-08-12 108/108 0.7276 28 7 run2/ 2026-08-20 108/108 0.7329 33 5 Model anthropic/claude-opus-5, thinking max, --max_steps 500, screenshot-only observation… See the full description on the dataset page: https://huggingface.co/datasets/simular-ai/sai-osworld-v2-benchmark-runs.textother1 likes4k downloads1mo agoHugging Face09BangumiBase /sailormoon1990s Bangumi Image Base of Sailor Moon (1990s) This is the image base of bangumi Sailor Moon (1990s), we detected 132 characters, 14684 images in total. The full dataset is here. Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability). Here is… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/sailormoon1990s.image10K<n<100K0 likes3.6k downloads3y agoHugging Face10sailplane /SWE-bench_Lite_filtered_10 likes3.3k downloads2y agoHugging Face11sailor2 /sea-synthetictext10M<n<100M0 likes3.2k downloads2y agoHugging Face12sailor2 /sailor2-pretrain-data-stage1The pre-training dataset (stage1) for the Sailor2 models, including 1B, 8B and 20B. text100M<n<1B0 likes2.9k downloads2y agoHugging Face13SAIS-Life-Science /Aneumo Aneumo Datasets AneumoDataset is a comprehensive multi-modal dataset containing 3D anatomical structures and simulated hemodynamic parameters for cerebral aneurysms, supporting both computational modeling and AI-based analysis. textn<1K6 likes2.6k downloads6mo agoHugging Face14saidmoh /ahmedvv0 likes2.3k downloads5mo agoHugging Face15SAINetset /SAINetset_v8.0 SAINetset - Wildfire Smoke Detection Dataset Dataset of real-world images captured by SAI (Sistema de Alerta de Incendios / Fire Alert System) surveillance nodes for wildfire smoke detection in Cordoba, Argentina. Current version: v8.0 (January 2026) About SAI The SAI (Fire Alert System) is an open-source early wildfire detection platform developed by AlterMundi, a civil association in Argentina. The system uses distributed camera nodes with YOLO-based AI (powered by… See the full description on the dataset page: https://huggingface.co/datasets/SAINetset/SAINetset_v8.0.imageobject-detection1K<n<10K1 likes2.2k downloads9mo agoHugging Face16sailor2 /sea-commoncrawl-high-qualitytext10M<n<100M0 likes1.9k downloads2y agoHugging Face17saibi /gs-lrm3d4 likes1.8k downloads2y agoHugging Face18sailor2 /sea-internettext10M<n<100M1 likes1.6k downloads2y agoHugging Face19sailor2 /sea-pdf-texttext10M<n<100M1 likes1.6k downloads2y agoHugging Face20IlyaGusev /saiga_scoredSFT dataset for the Saiga family of models collected from various sources. tabular10K<n<100K23 likes1.6k downloads2y agoHugging Face21saidurga001301 /mathmetics-dataset Transformer Math Dataset (100,000,000 Samples Sharded) High-precision synthetic mathematical expression dataset generated for training sequence-to-sequence math Transformers in JAX/Flax. Dataset Structure Total Samples: 100,000,000 Shard Format: JSONL sharded files (100,000 samples per shard) Supported Operations: +, -, *, /, ^, sin, cos, tan, log, ln, exp, sqrt, abs Expression Depth Range: Depth 1 to 3 Integer Operand Ratio: 50% Data Fields Each… See the full description on the dataset page: https://huggingface.co/datasets/saidurga001301/mathmetics-dataset.text-generation0 likes1.6k downloads1mo agoHugging Face22SaifPunjwani /ruleforge-checkpoints0 likes1.6k downloads3d agoHugging Face23Saiukk /POSEJEPA_Training What is? A set of prerender 2D images from GSO datasets, including png, mask, camera calibration and poses. Why this DATASET Could be used to others?? When training Deep Learning models for Novel View Synthesis (NVS) or 3D-to-2D Representation Learning, loading 3D meshes and rendering views on-the-fly inside PyTorch DataLoaders creates massive bottlenecks. This script solves critical problems: Time consumption Ram OOM 0 likes1.5k downloads5d agoHugging Face24saillab /taco-datasetsThis repo consists of the datasets used for the TaCo paper. There are four datasets: Multilingual Alpaca-52K GPT-4 dataset Multilingual Dolly-15K GPT-4 dataset TaCo dataset Multilingual Vicuna Benchmark dataset We translated the first three datasets using Google Cloud Translation. The TaCo dataset is created by using the TaCo approach as described in our paper, combining the Alpaca-52K and Dolly-15K datasets. If you would like to create the TaCo dataset for a specific language, you can… See the full description on the dataset page: https://huggingface.co/datasets/saillab/taco-datasets.text1M<n<10M17 likes1.5k downloads3y agoHugging Face25saifkhichi96 /spinetrackPlease visit the project homepage for more details about the dataset, associated research paper, and citation information. If you use our dataset in your work, we request proper attribution and a citation to our paper "Towards Unconstrained 2D Pose Estimation of the Human Spine". imagekeypoint-detection10K<n<100K0 likes1.5k downloads6mo agoHugging Face26Weibang2004 /Complet4R_SAILVOS3D1 likes1.4k downloads4mo agoHugging Face27BangumiBase /saikyounoshienshokuwajutsushidearuorewasekaisaikyouclanwoshitagaeru Bangumi Image Base of Saikyou No Shienshoku "wajutsushi" De Aru Ore Wa Sekai Saikyou Clan Wo Shitagaeru This is the image base of bangumi Saikyou no Shienshoku "Wajutsushi" de Aru Ore wa Sekai Saikyou Clan wo Shitagaeru, we detected 70 characters, 4558 images in total. The full dataset is here. Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/saikyounoshienshokuwajutsushidearuorewasekaisaikyouclanwoshitagaeru.image1K<n<10K0 likes1.3k downloads2y agoHugging Face28stephaniegwright2770 /sailboat0 likes1.3k downloads3d agoHugging Face29SaintsStudios /Tumbuka_Text-Speech_Audioaudion<1K0 likes1.2k downloads26d agoHugging Face30SaintsStudios /Tumbuka_Text-Speechaudio1K<n<10K0 likes1.1k downloads26d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.