CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SAIRfoundation /equational-theories-selected-problems Equational Theories Selected Problems Update (September 11, 2026) This dataset was updated on September 11, 2026. Main changes: released the official Stage 2 evaluation problems: stage2_evaluation_main (200 problems; ground truth withheld — answer is null until Stage 2 concludes) and stage2_evaluation_research (100 order-5 research problems with no ground truth) added metadata/stage2_evaluation_main.json and metadata/stage2_evaluation_research.json… See the full description on the dataset page: https://huggingface.co/datasets/SAIRfoundation/equational-theories-selected-problems.tabular1K<n<10K11 likes12k downloads15d agoHugging Face02saidutta69 /fable-5-premium 🧠 Fable-5 Premium Dataset 🚀 V2 is out! This dataset has a successor: fable-5-premium-v2 — new users should start there. A rigorously cleaned, high-quality supervised fine-tuning (SFT) dataset built from Claude Fable-5 agent traces. Priorities: Quality > Ease of Access > Quantity 📊 Dataset Overview Property Value Total Records 12,730 Train Split 5,728 (45.0%) Validation Split 318 (2.5%) Test Split 319 (2.5%) Created 2026-07-30… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/fable-5-premium.texttext-generation10K<n<100K134 likes10k downloads13d agoHugging Face03sailor2 /sea-commoncrawltext100M<n<1B1 likes7.1k downloads2y agoHugging Face04sail /regmix-data RegMix Data Dataset Description The RegMix Data is a curated dataset derived from the Pile-Uncopyrighted, specifically designed for the RegMix paper (https://huggingface.co/papers/2407.01492). This dataset aims to facilitate the automatic identification of high-performing data mixtures for language model pre-training by formulating it as a regression task. Key Features: Size: Approximately 1TB disk space, 250B tokens Distribution: Follows the natural token… See the full description on the dataset page: https://huggingface.co/datasets/sail/regmix-data.text10M<n<100M4 likes6.4k downloads2y agoHugging Face05BangumiBase /saikyounoousamanidomenojinseiwananiwosuru Bangumi Image Base of Saikyou No Ousama, Nidome No Jinsei Wa Nani Wo Suru? This is the image base of bangumi Saikyou no Ousama, Nidome no Jinsei wa Nani wo Suru?, we detected 71 characters, 4913 images in total. The full dataset is here. Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/saikyounoousamanidomenojinseiwananiwosuru.image1K<n<10K0 likes4.7k downloads1y agoHugging Face06simular-ai /sai-osworld-v2-benchmark-runs Sai on OSWorld-V2 — benchmark runs of record Two complete 108-task runs of the Sai computer-use agent on OSWorld-V2, with full per-task evidence: scores, trajectories, agent runtime logs, evaluator logs, API protocol logs, and run manifests. Run Date Tasks scored Mean score Perfect (1.0) Zeros run1/ 2026-08-12 108/108 0.7276 28 7 run2/ 2026-08-20 108/108 0.7329 33 5 Model anthropic/claude-opus-5, thinking max, --max_steps 500, screenshot-only observation… See the full description on the dataset page: https://huggingface.co/datasets/simular-ai/sai-osworld-v2-benchmark-runs.textother1 likes4.2k downloads1mo agoHugging Face07sailor2 /sea-synthetictext10M<n<100M0 likes3.3k downloads2y agoHugging Face08sailor2 /sailor2-pretrain-data-stage1The pre-training dataset (stage1) for the Sailor2 models, including 1B, 8B and 20B. text100M<n<1B0 likes2.9k downloads2y agoHugging Face09SAIS-Life-Science /Aneumo Aneumo Datasets AneumoDataset is a comprehensive multi-modal dataset containing 3D anatomical structures and simulated hemodynamic parameters for cerebral aneurysms, supporting both computational modeling and AI-based analysis. textn<1K6 likes2.3k downloads6mo agoHugging Face10SAINetset /SAINetset_v8.0 SAINetset - Wildfire Smoke Detection Dataset Dataset of real-world images captured by SAI (Sistema de Alerta de Incendios / Fire Alert System) surveillance nodes for wildfire smoke detection in Cordoba, Argentina. Current version: v8.0 (January 2026) About SAI The SAI (Fire Alert System) is an open-source early wildfire detection platform developed by AlterMundi, a civil association in Argentina. The system uses distributed camera nodes with YOLO-based AI (powered by… See the full description on the dataset page: https://huggingface.co/datasets/SAINetset/SAINetset_v8.0.imageobject-detection1K<n<10K1 likes2.2k downloads9mo agoHugging Face11sailor2 /sea-commoncrawl-high-qualitytext10M<n<100M0 likes1.9k downloads2y agoHugging Face12sailor2 /sea-internettext10M<n<100M1 likes1.8k downloads2y agoHugging Face13IlyaGusev /saiga_scoredSFT dataset for the Saiga family of models collected from various sources. tabular10K<n<100K23 likes1.6k downloads2y agoHugging Face14saillab /taco-datasetsThis repo consists of the datasets used for the TaCo paper. There are four datasets: Multilingual Alpaca-52K GPT-4 dataset Multilingual Dolly-15K GPT-4 dataset TaCo dataset Multilingual Vicuna Benchmark dataset We translated the first three datasets using Google Cloud Translation. The TaCo dataset is created by using the TaCo approach as described in our paper, combining the Alpaca-52K and Dolly-15K datasets. If you would like to create the TaCo dataset for a specific language, you can… See the full description on the dataset page: https://huggingface.co/datasets/saillab/taco-datasets.text1M<n<10M17 likes1.5k downloads3y agoHugging Face15sailor2 /sea-pdf-texttext10M<n<100M1 likes1.5k downloads2y agoHugging Face16BangumiBase /saikyounoshienshokuwajutsushidearuorewasekaisaikyouclanwoshitagaeru Bangumi Image Base of Saikyou No Shienshoku "wajutsushi" De Aru Ore Wa Sekai Saikyou Clan Wo Shitagaeru This is the image base of bangumi Saikyou no Shienshoku "Wajutsushi" de Aru Ore wa Sekai Saikyou Clan wo Shitagaeru, we detected 70 characters, 4558 images in total. The full dataset is here. Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/saikyounoshienshokuwajutsushidearuorewasekaisaikyouclanwoshitagaeru.image1K<n<10K0 likes1.3k downloads2y agoHugging Face17SaintsStudios /Tumbuka_Text-Speechaudio1K<n<10K0 likes1.1k downloads27d agoHugging Face18tli-legumes /sainfoin-seed-datasetimage100K<n<1M0 likes1k downloads3mo agoHugging Face19marianbasti /jurisprudencia-Argentina-SAIJ Jurisprudencia de la Repùblica Argentina - Sistema Argentino de Información Jurídica Este dataset es actualizado diariamente con la información de SAIJ utilizando la librería de SandboxAI Formato El formato del dataset es el siguiente: { "numero-sumario": "Número de identificación del sumario", "materia": "Área del derecho a la que pertenece el caso", "timestamp": "Fecha y hora de creación del registro", "timestamp-m": "Fecha y hora de la última modificación del… See the full description on the dataset page: https://huggingface.co/datasets/marianbasti/jurisprudencia-Argentina-SAIJ.text100K<n<1M1 likes980 downloads5mo agoHugging Face20saidutta69 /fable-5-premium-v2 🧠 Fable-5 Premium V2 A rigorously cleaned, high-quality supervised fine-tuning (SFT) dataset of 100,000 agent traces, built for training tool-using models. Successor to fable-5-premium. Priorities: Quality > Ease of Access > Quantity 📊 Dataset Overview Property Value Total Traces 100,000 Train Split 85,000 (85.0%) Validation Split 7,500 (7.5%) Test Split 7,500 (7.5%) Average Quality 0.966 (0.8–1.0 band) Distilled From Claude… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/fable-5-premium-v2.texttext-generation100K<n<1M16 likes888 downloads13d agoHugging Face21saileshbro /nepali-cs-asr Nepali–English Code-Switched ASR A ~59-hour corpus of spontaneous Nepali–English code-switched speech clipped from publicly available STEM and CS lecture videos on YouTube. The dataset targets ASR model training and evaluation for code-switched (CS) Nepali–English speech — a variety commonly used in Nepali higher education and online tutoring, where teachers fluidly mix Nepali grammar with English technical vocabulary. v2 (2026-07) — the current revision. Splits are… See the full description on the dataset page: https://huggingface.co/datasets/saileshbro/nepali-cs-asr.audioautomatic-speech-recognition10K<n<100K1 likes882 downloads2mo agoHugging Face22saidutta69 /PhishTrap PhishTrap Catch phishing URLs before they catch you — 16 features, 19,954 URLs, balanced 50/50. Cross-verified from 496K phishing domains + Tranco top 10K. Automatically refreshed every 6 hours. Priorities: Quality > Ease of Access > Quantity Build pipeline (open source): github.com/instax-dutta/PhishTrap — see how every row is fetched, merged, deduplicated, validated and published. Dataset Overview PhishTrap is a curated phishing URL detection dataset… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/PhishTrap.tabular10K<n<100K0 likes867 downloads2h agoHugging Face23SaifPunjwani /mroot-q17-scoreband-runtime-v1textn<1K0 likes773 downloads21d agoHugging Face24saidurga001301 /mathmetics-dataset-custom Transformer Math Dataset (54,000,000 Samples Sharded) High-precision synthetic mathematical expression dataset generated for training sequence-to-sequence math Transformers in JAX/Flax. Dataset Structure Total Samples: 54,000,000 Shard Format: JSONL sharded files (100,000 samples per shard) Supported Operations: +, -, *, /, ^, sin, cos, tan, log, ln, exp, sqrt, abs Expression Depth Range: Depth 1 to 2 Integer Operand Ratio: 0% Data Fields Each… See the full description on the dataset page: https://huggingface.co/datasets/saidurga001301/mathmetics-dataset-custom.texttext-generation100M<n<1B0 likes713 downloads1mo agoHugging Face25BangumiBase /sailormoon2010s Bangumi Image Base of Sailor Moon (2010s) This is the image base of bangumi Sailor Moon (2010s), we detected 46 characters, 3463 images in total. The full dataset is here. Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability). Here is… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/sailormoon2010s.image1K<n<10K0 likes682 downloads3y agoHugging Face26SandboxAQ /SAIRgated Announcing SAIR Structurally-Augmented IC50 Repository In collaboration with Nvidia The Largest Publicly Available Binding Affinity Dataset with Cofolded 3D Structures SAIR (Structurally Augmented IC50 Repository), is the largest public dataset of protein--ligand 3D structures paired with binding potency measurements. SAIR contains over one million protein--ligand complexes (1,048,857 unique pairs) and a total of 5.2 million 3D structures, curated from the ChEMBL and… See the full description on the dataset page: https://huggingface.co/datasets/SandboxAQ/SAIR.tabular1M<n<10M56 likes631 downloads6mo agoHugging Face27SaifPunjwani /mroot-q17-scoreband-runtime-bridge-v1textn<1K0 likes628 downloads20d agoHugging Face28saifgazali /reglu2_fr_eval_10BT_jetons_fineweb2_culturax_wikipedia_pdnewspapertext10M<n<100M0 likes572 downloads3mo agoHugging Face29saidurga001301 /mathmetics-dataset-intmax Transformer Math Dataset (200,000,000 Samples Sharded) High-precision synthetic mathematical expression dataset generated for training sequence-to-sequence math Transformers in JAX/Flax. Dataset Structure Total Samples: 200,000,000 Shard Format: JSONL sharded files (100,000 samples per shard) Supported Operations: +, -, *, /, ^, sin, cos, tan, log, ln, exp, sqrt, abs Expression Depth Range: Depth 4 to 6 Integer Operand Ratio: 80% Data Fields Each… See the full description on the dataset page: https://huggingface.co/datasets/saidurga001301/mathmetics-dataset-intmax.texttext-generation100M<n<1B0 likes545 downloads1mo agoHugging Face30BangumiBase /saikyouonmyoujinoisekaitenseiki Bangumi Image Base of Saikyou Onmyouji No Isekai Tenseiki This is the image base of bangumi Saikyou Onmyouji no Isekai Tenseiki, we detected 63 characters, 5099 images in total. The full dataset is here. Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/saikyouonmyoujinoisekaitenseiki.image1K<n<10K0 likes536 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.