datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
equational-theories-selected-problems
Equational Theories Selected Problems
Update (September 11, 2026)
This dataset was updated on September 11, 2026.
Main changes:
released the official Stage 2 evaluation problems: stage2_evaluation_main (200 problems; ground truth withheld — answer is null until Stage 2 concludes) and stage2_evaluation_research (100 order-5 research problems with no ground truth)
added metadata/stage2_evaluation_main.json and metadata/stage2_evaluation_research.json… See the full description on the dataset page: https://huggingface.co/datasets/SAIRfoundation/equational-theories-selected-problems.fable-5-premium
🧠 Fable-5 Premium Dataset
🚀 V2 is out! This dataset has a successor: fable-5-premium-v2 — new users should start there.
A rigorously cleaned, high-quality supervised fine-tuning (SFT) dataset built from Claude Fable-5 agent traces.
Priorities: Quality > Ease of Access > Quantity
📊 Dataset Overview
Property
Value
Total Records
12,730
Train Split
5,728 (45.0%)
Validation Split
318 (2.5%)
Test Split
319 (2.5%)
Created
2026-07-30… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/fable-5-premium.slo-rlvr-resultssea-commoncrawlregmix-data
RegMix Data
Dataset Description
The RegMix Data is a curated dataset derived from the Pile-Uncopyrighted, specifically designed for the RegMix paper (https://huggingface.co/papers/2407.01492). This dataset aims to facilitate the automatic identification of high-performing data mixtures for language model pre-training by formulating it as a regression task.
Key Features:
Size: Approximately 1TB disk space, 250B tokens
Distribution: Follows the natural token… See the full description on the dataset page: https://huggingface.co/datasets/sail/regmix-data.saikyounoousamanidomenojinseiwananiwosuru
Bangumi Image Base of Saikyou No Ousama, Nidome No Jinsei Wa Nani Wo Suru?
This is the image base of bangumi Saikyou no Ousama, Nidome no Jinsei wa Nani wo Suru?, we detected 71 characters, 4913 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/saikyounoousamanidomenojinseiwananiwosuru.SWE-bench_Lite_filteredsai-osworld-v2-benchmark-runs
Sai on OSWorld-V2 — benchmark runs of record
Two complete 108-task runs of the Sai computer-use agent on OSWorld-V2, with full
per-task evidence: scores, trajectories, agent runtime logs, evaluator logs, API
protocol logs, and run manifests.
Run
Date
Tasks scored
Mean score
Perfect (1.0)
Zeros
run1/
2026-08-12
108/108
0.7276
28
7
run2/
2026-08-20
108/108
0.7329
33
5
Model anthropic/claude-opus-5, thinking max, --max_steps 500, screenshot-only
observation… See the full description on the dataset page: https://huggingface.co/datasets/simular-ai/sai-osworld-v2-benchmark-runs.sailormoon1990s
Bangumi Image Base of Sailor Moon (1990s)
This is the image base of bangumi Sailor Moon (1990s), we detected 132 characters, 14684 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).
Here is… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/sailormoon1990s.SWE-bench_Lite_filtered_1sea-syntheticsailor2-pretrain-data-stage1The pre-training dataset (stage1) for the Sailor2 models, including 1B, 8B and 20B.
Aneumo
Aneumo Datasets
AneumoDataset is a comprehensive multi-modal dataset containing 3D anatomical structures and simulated hemodynamic parameters for cerebral aneurysms, supporting both computational modeling and AI-based analysis.
ahmedvvSAINetset_v8.0
SAINetset - Wildfire Smoke Detection Dataset
Dataset of real-world images captured by SAI (Sistema de Alerta de Incendios / Fire Alert System) surveillance nodes for wildfire smoke detection in Cordoba, Argentina.
Current version: v8.0 (January 2026)
About SAI
The SAI (Fire Alert System) is an open-source early wildfire detection platform developed by AlterMundi, a civil association in Argentina. The system uses distributed camera nodes with YOLO-based AI (powered by… See the full description on the dataset page: https://huggingface.co/datasets/SAINetset/SAINetset_v8.0.sea-commoncrawl-high-qualitygs-lrmsea-internetsea-pdf-textsaiga_scoredSFT dataset for the Saiga family of models collected from various sources.
mathmetics-dataset
Transformer Math Dataset (100,000,000 Samples Sharded)
High-precision synthetic mathematical expression dataset generated for training sequence-to-sequence math Transformers in JAX/Flax.
Dataset Structure
Total Samples: 100,000,000
Shard Format: JSONL sharded files (100,000 samples per shard)
Supported Operations: +, -, *, /, ^, sin, cos, tan, log, ln, exp, sqrt, abs
Expression Depth Range: Depth 1 to 3
Integer Operand Ratio: 50%
Data Fields
Each… See the full description on the dataset page: https://huggingface.co/datasets/saidurga001301/mathmetics-dataset.ruleforge-checkpointsPOSEJEPA_Training
What is?
A set of prerender 2D images from GSO datasets, including png, mask, camera calibration and poses.
Why this DATASET Could be used to others??
When training Deep Learning models for Novel View Synthesis (NVS) or 3D-to-2D Representation Learning, loading 3D meshes and rendering views on-the-fly inside PyTorch DataLoaders creates massive bottlenecks. This script solves critical problems:
Time consumption
Ram OOM
taco-datasetsThis repo consists of the datasets used for the TaCo paper. There are four datasets:
Multilingual Alpaca-52K GPT-4 dataset
Multilingual Dolly-15K GPT-4 dataset
TaCo dataset
Multilingual Vicuna Benchmark dataset
We translated the first three datasets using Google Cloud Translation.
The TaCo dataset is created by using the TaCo approach as described in our paper, combining the Alpaca-52K and Dolly-15K datasets.
If you would like to create the TaCo dataset for a specific language, you can… See the full description on the dataset page: https://huggingface.co/datasets/saillab/taco-datasets.spinetrackPlease visit the project homepage for more details about the dataset, associated research paper, and citation information.
If you use our dataset in your work, we request proper attribution and a citation to our paper "Towards Unconstrained 2D Pose Estimation of the Human Spine".
Complet4R_SAILVOS3Dsaikyounoshienshokuwajutsushidearuorewasekaisaikyouclanwoshitagaeru
Bangumi Image Base of Saikyou No Shienshoku "wajutsushi" De Aru Ore Wa Sekai Saikyou Clan Wo Shitagaeru
This is the image base of bangumi Saikyou no Shienshoku "Wajutsushi" de Aru Ore wa Sekai Saikyou Clan wo Shitagaeru, we detected 70 characters, 4558 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/saikyounoshienshokuwajutsushidearuorewasekaisaikyouclanwoshitagaeru.sailboatTumbuka_Text-Speech_AudioTumbuka_Text-Speech
