datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
equational-theories-selected-problems
Equational Theories Selected Problems
Update (September 11, 2026)
This dataset was updated on September 11, 2026.
Main changes:
released the official Stage 2 evaluation problems: stage2_evaluation_main (200 problems; ground truth withheld — answer is null until Stage 2 concludes) and stage2_evaluation_research (100 order-5 research problems with no ground truth)
added metadata/stage2_evaluation_main.json and metadata/stage2_evaluation_research.json… See the full description on the dataset page: https://huggingface.co/datasets/SAIRfoundation/equational-theories-selected-problems.fable-5-premium
🧠 Fable-5 Premium Dataset
🚀 V2 is out! This dataset has a successor: fable-5-premium-v2 — new users should start there.
A rigorously cleaned, high-quality supervised fine-tuning (SFT) dataset built from Claude Fable-5 agent traces.
Priorities: Quality > Ease of Access > Quantity
📊 Dataset Overview
Property
Value
Total Records
12,730
Train Split
5,728 (45.0%)
Validation Split
318 (2.5%)
Test Split
319 (2.5%)
Created
2026-07-30… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/fable-5-premium.sea-commoncrawlregmix-data
RegMix Data
Dataset Description
The RegMix Data is a curated dataset derived from the Pile-Uncopyrighted, specifically designed for the RegMix paper (https://huggingface.co/papers/2407.01492). This dataset aims to facilitate the automatic identification of high-performing data mixtures for language model pre-training by formulating it as a regression task.
Key Features:
Size: Approximately 1TB disk space, 250B tokens
Distribution: Follows the natural token… See the full description on the dataset page: https://huggingface.co/datasets/sail/regmix-data.saikyounoousamanidomenojinseiwananiwosuru
Bangumi Image Base of Saikyou No Ousama, Nidome No Jinsei Wa Nani Wo Suru?
This is the image base of bangumi Saikyou no Ousama, Nidome no Jinsei wa Nani wo Suru?, we detected 71 characters, 4913 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/saikyounoousamanidomenojinseiwananiwosuru.sai-osworld-v2-benchmark-runs
Sai on OSWorld-V2 — benchmark runs of record
Two complete 108-task runs of the Sai computer-use agent on OSWorld-V2, with full
per-task evidence: scores, trajectories, agent runtime logs, evaluator logs, API
protocol logs, and run manifests.
Run
Date
Tasks scored
Mean score
Perfect (1.0)
Zeros
run1/
2026-08-12
108/108
0.7276
28
7
run2/
2026-08-20
108/108
0.7329
33
5
Model anthropic/claude-opus-5, thinking max, --max_steps 500, screenshot-only
observation… See the full description on the dataset page: https://huggingface.co/datasets/simular-ai/sai-osworld-v2-benchmark-runs.sea-syntheticsailor2-pretrain-data-stage1The pre-training dataset (stage1) for the Sailor2 models, including 1B, 8B and 20B.
Aneumo
Aneumo Datasets
AneumoDataset is a comprehensive multi-modal dataset containing 3D anatomical structures and simulated hemodynamic parameters for cerebral aneurysms, supporting both computational modeling and AI-based analysis.
SAINetset_v8.0
SAINetset - Wildfire Smoke Detection Dataset
Dataset of real-world images captured by SAI (Sistema de Alerta de Incendios / Fire Alert System) surveillance nodes for wildfire smoke detection in Cordoba, Argentina.
Current version: v8.0 (January 2026)
About SAI
The SAI (Fire Alert System) is an open-source early wildfire detection platform developed by AlterMundi, a civil association in Argentina. The system uses distributed camera nodes with YOLO-based AI (powered by… See the full description on the dataset page: https://huggingface.co/datasets/SAINetset/SAINetset_v8.0.sea-commoncrawl-high-qualitysea-internetsaiga_scoredSFT dataset for the Saiga family of models collected from various sources.
taco-datasetsThis repo consists of the datasets used for the TaCo paper. There are four datasets:
Multilingual Alpaca-52K GPT-4 dataset
Multilingual Dolly-15K GPT-4 dataset
TaCo dataset
Multilingual Vicuna Benchmark dataset
We translated the first three datasets using Google Cloud Translation.
The TaCo dataset is created by using the TaCo approach as described in our paper, combining the Alpaca-52K and Dolly-15K datasets.
If you would like to create the TaCo dataset for a specific language, you can… See the full description on the dataset page: https://huggingface.co/datasets/saillab/taco-datasets.sea-pdf-textsaikyounoshienshokuwajutsushidearuorewasekaisaikyouclanwoshitagaeru
Bangumi Image Base of Saikyou No Shienshoku "wajutsushi" De Aru Ore Wa Sekai Saikyou Clan Wo Shitagaeru
This is the image base of bangumi Saikyou no Shienshoku "Wajutsushi" de Aru Ore wa Sekai Saikyou Clan wo Shitagaeru, we detected 70 characters, 4558 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/saikyounoshienshokuwajutsushidearuorewasekaisaikyouclanwoshitagaeru.Tumbuka_Text-Speechsainfoin-seed-datasetjurisprudencia-Argentina-SAIJ
Jurisprudencia de la Repùblica Argentina - Sistema Argentino de Información Jurídica
Este dataset es actualizado diariamente con la información de SAIJ utilizando la librería de SandboxAI
Formato
El formato del dataset es el siguiente:
{
"numero-sumario": "Número de identificación del sumario",
"materia": "Área del derecho a la que pertenece el caso",
"timestamp": "Fecha y hora de creación del registro",
"timestamp-m": "Fecha y hora de la última modificación del… See the full description on the dataset page: https://huggingface.co/datasets/marianbasti/jurisprudencia-Argentina-SAIJ.fable-5-premium-v2
🧠 Fable-5 Premium V2
A rigorously cleaned, high-quality supervised fine-tuning (SFT) dataset of 100,000 agent traces, built for training tool-using models. Successor to fable-5-premium.
Priorities: Quality > Ease of Access > Quantity
📊 Dataset Overview
Property
Value
Total Traces
100,000
Train Split
85,000 (85.0%)
Validation Split
7,500 (7.5%)
Test Split
7,500 (7.5%)
Average Quality
0.966 (0.8–1.0 band)
Distilled From
Claude… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/fable-5-premium-v2.nepali-cs-asr
Nepali–English Code-Switched ASR
A ~59-hour corpus of spontaneous Nepali–English code-switched speech clipped from publicly available STEM and CS lecture videos on YouTube. The dataset targets ASR model training and evaluation for code-switched (CS) Nepali–English speech — a variety commonly used in Nepali higher education and online tutoring, where teachers fluidly mix Nepali grammar with English technical vocabulary.
v2 (2026-07) — the current revision. Splits are… See the full description on the dataset page: https://huggingface.co/datasets/saileshbro/nepali-cs-asr.PhishTrap
PhishTrap
Catch phishing URLs before they catch you — 16 features, 19,954 URLs, balanced 50/50. Cross-verified from 496K phishing domains + Tranco top 10K. Automatically refreshed every 6 hours.
Priorities: Quality > Ease of Access > Quantity
Build pipeline (open source): github.com/instax-dutta/PhishTrap — see how every row is fetched, merged, deduplicated, validated and published.
Dataset Overview
PhishTrap is a curated phishing URL detection dataset… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/PhishTrap.mroot-q17-scoreband-runtime-v1mathmetics-dataset-custom
Transformer Math Dataset (54,000,000 Samples Sharded)
High-precision synthetic mathematical expression dataset generated for training sequence-to-sequence math Transformers in JAX/Flax.
Dataset Structure
Total Samples: 54,000,000
Shard Format: JSONL sharded files (100,000 samples per shard)
Supported Operations: +, -, *, /, ^, sin, cos, tan, log, ln, exp, sqrt, abs
Expression Depth Range: Depth 1 to 2
Integer Operand Ratio: 0%
Data Fields
Each… See the full description on the dataset page: https://huggingface.co/datasets/saidurga001301/mathmetics-dataset-custom.sailormoon2010s
Bangumi Image Base of Sailor Moon (2010s)
This is the image base of bangumi Sailor Moon (2010s), we detected 46 characters, 3463 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).
Here is… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/sailormoon2010s.SAIR
Announcing SAIR
Structurally-Augmented IC50 Repository
In collaboration with Nvidia
The Largest Publicly Available Binding Affinity Dataset with Cofolded 3D Structures
SAIR (Structurally Augmented IC50 Repository), is the largest public
dataset of protein--ligand 3D structures paired with binding potency
measurements. SAIR contains over one million protein--ligand complexes
(1,048,857 unique pairs) and a total of 5.2 million 3D structures,
curated from the ChEMBL and… See the full description on the dataset page: https://huggingface.co/datasets/SandboxAQ/SAIR.mroot-q17-scoreband-runtime-bridge-v1reglu2_fr_eval_10BT_jetons_fineweb2_culturax_wikipedia_pdnewspapermathmetics-dataset-intmax
Transformer Math Dataset (200,000,000 Samples Sharded)
High-precision synthetic mathematical expression dataset generated for training sequence-to-sequence math Transformers in JAX/Flax.
Dataset Structure
Total Samples: 200,000,000
Shard Format: JSONL sharded files (100,000 samples per shard)
Supported Operations: +, -, *, /, ^, sin, cos, tan, log, ln, exp, sqrt, abs
Expression Depth Range: Depth 4 to 6
Integer Operand Ratio: 80%
Data Fields
Each… See the full description on the dataset page: https://huggingface.co/datasets/saidurga001301/mathmetics-dataset-intmax.saikyouonmyoujinoisekaitenseiki
Bangumi Image Base of Saikyou Onmyouji No Isekai Tenseiki
This is the image base of bangumi Saikyou Onmyouji no Isekai Tenseiki, we detected 63 characters, 5099 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/saikyouonmyoujinoisekaitenseiki.
