datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
epoch_ai_swebench_verified
Epoch AI SWE-bench Verified Traces
Complete public trace archives and an analysis-ready Parquet conversion of Epoch AI's SWE-bench Verified evaluations.
Contents
34 published evaluation runs covering 16,456 traces (484 SWE-bench instances per run).
data/: loadable Parquet data, one exact trace per row.
original/: the byte-identical .eval archives published by Epoch AI.
run_metadata/: non-sample files from each .eval archive (header.json, summaries, reductions… See the full description on the dataset page: https://huggingface.co/datasets/naderalfares/epoch_ai_swebench_verified.GeneralThought-430K-filteredData from https://huggingface.co/datasets/GeneralReasoning/GeneralThought-430K, removed prompts with non commercial data
NatureBench-traces
NatureBench-traces
NatureBench-traces contains the full solving process
of coding agents on the 90 tasks of NatureBench.
The task packages themselves (task brief, data, evaluator, SOTA
scores) live in the sibling repository FrontisAI/NatureBench.
Harbor-compatible task packages are available in
FrontisAI/NatureBench-Harbor.
The traces released here were collected with NatureBench's native task format, not from the Harbor tasks.
This repository releases only the process traces:… See the full description on the dataset page: https://huggingface.co/datasets/FrontisAI/NatureBench-traces.weather-geo-era5
Weather Geo ERA5 Dataset (Optimized)
📊 Dataset Overview
This dataset contains 1.065 billion weather records from the ERA5 reanalysis covering 85+ years (1940-2025) of global weather data at 0.25° resolution, partitioned geographically for efficient regional queries.
Key Features
🌍 Global Coverage: Complete worldwide historical weather data
⏰ Time Range: 1940-2025 (85+ years) - UPDATED
📍 Resolution: 0.25° x 0.25° (~28km grid)
🗂️ Geographic Partitioning: 48… See the full description on the dataset page: https://huggingface.co/datasets/NaaVrug/weather-geo-era5.multi_session_chat
Dataset Card for "multi_session_chat"
More Information needed
relaion2b-natural-embeddings
LAION-Natural Embeddings: CLIP ViT-H/14 Features for ~500M Natural Photographs (CCN 2025, Roth & Hebart)
LAION-Natural Embeddings provides pre-computed CLIP ViT-H/14 embeddings for ~500 million natural photographs from ReLAION-2B, filtered using the LAION-Natural naturalness classifier (score > 0.7).
Also known as: LAION-Natural Embeddings · ReLAION-Natural Embeddings · LAION-2B-Natural Embeddings
Part of the LAION-Natural dataset family, introduced in: How to sample the… See the full description on the dataset page: https://huggingface.co/datasets/andropar/relaion2b-natural-embeddings.mtnwx-trainingstock-price-history
SBI Stock Price History
Partitioned Parquet archive generated from Mnie/SBI price history.
Layout:
data/source=sbi/market={MARKET}/timeframe={TIMEFRAME}/year={YYYY}/part-000.parquet
status/source=sbi/market={MARKET}/timeframe={TIMEFRAME}/part-000.parquet
schemas/price-history.v1.schema.json
Canonical storage is this Hugging Face Dataset repo. Local history/ directories are temporary
fetch/cache artifacts and should be removed after upload.
jamendo_arraynaaseh1-bottle-holder-calib-090326This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_yaw.pos",
"wrist_roll.pos",
"gripper.pos"
]… See the full description on the dataset page: https://huggingface.co/datasets/VibeCuisine/naaseh1-bottle-holder-calib-090326.nadora-global-industries
NADORA Global Industries
A synthetic multinational, built to be developed against rather than
demonstrated with.
One fictional company — $5.20bn revenue, $716m EBITDA, 24,000 employees, 18
countries, 35 legal entities, five business units — traded daily from January
2022 to December 2026 and rendered at six fidelities, from a 3 MB unit-test
fixture to a 5 GB full-scale corpus.
38,964,663 rows · 11 GB · 2,319 verification assertions, all passing.
100% synthetic. No real company… See the full description on the dataset page: https://huggingface.co/datasets/hemanthreddy901/nadora-global-industries.FineWeb-Nano
FineWeb-Nano
Dataset Description
FineWeb-Nano is a highly curated, premium subset extracted from nampdn-ai/mini-fineweb.
How "The Best" Was Determined
This dataset was created programmatically by streaming the original dataset and sorting chunks based on a rigorous quality scoring algorithm. The heuristic heavily favors:
High language_score (if provided by the upstream extraction).
Optimal document length (penalizing abnormally short snippets and excessively… See the full description on the dataset page: https://huggingface.co/datasets/ray0rf1re/FineWeb-Nano.baby_names
Dataset Card for "baby_names"
More Information needed
gender-by-name
Dataset Card for "Gender-by-Name"
This dataset attributes first names to genders, giving counts and probabilities. It combines open-source government data from the US, UK, Canada, and Australia. The dataset is taken from UCI Machine Learning Repository
Dataset Information
This dataset combines raw counts for first/given names of male and female babies in those time periods, and then calculates a probability for a name given the aggregate count. Source datasets are from… See the full description on the dataset page: https://huggingface.co/datasets/erickrribeiro/gender-by-name.so101_hand_blue_napkin
SO-ARM101 — "Hand me the blue napkin"
Teleoperated demonstrations of a human–robot handover on a real
SO-ARM101 — a 6-DOF, ~$300 open-source arm with
STS3215 servos. The robot picks up a pack of blue tissues from the table and places it into a
human hand.
Recorded with LeRobot (codebase_version: v2.1).
Robot
so101_follower, 6 DOF
Task
"Hand me the blue napkin" (single task)
Episodes
101 (complete set)
Frames
40,493
Episode length
400–401 frames ≈ 13.4 s each… See the full description on the dataset page: https://huggingface.co/datasets/Twu31/so101_hand_blue_napkin.ancient-scripts-datasets
Ancient Scripts Decipherment Datasets
Collated datasets for the paper:
Deciphering Undersegmented Ancient Scripts Using Phonetic Prior
Jiaming Luo, Frederik Hartmann, Enrico Santus, Regina Barzilay, Yuan Cao
Transactions of the Association for Computational Linguistics, 2021
arXiv:2010.11054
This repository gathers the training datasets used in the paper — both those hosted in the authors' GitHub repos and the external cited sources.
Repository Structure
data/
├──… See the full description on the dataset page: https://huggingface.co/datasets/Nacryos/ancient-scripts-datasets.japanese-conversion
Awesome Japanese IME Training Data
Awesome Japanese Corpus の本文を直接 KyTea で解析し、文脈付きかな漢字変換の
ランキング学習例を作成したデータセットです。中間の読み付きデータセットは
作りません。任意の検証モードでは、抽出範囲についてMeCabの読みとも一致した
例だけを採用できます。
context: 変換対象より前の本文
input: 変換対象のひらがな読み
correct: 元コーパスにある正解表記
incorrect: predict.py で全体または一部分を再変換した誤候補の配列
n_words: 抽出した連続形態素数
source_text と target_start / target_end により、元文章中の抽出位置を
復元できます。元データの利用条件は from と from_license を参照して
ください。
tiny-textbooks
Textbook-like Dataset: A High-Quality Resource for Small Language Models
The idea is simply inspired by the Textbooks Are All You Need II: phi-1.5 technical report paper. The source texts in this dataset have been gathered and carefully select the best of the falcon-refinedweb and minipile datasets to ensure the diversity, quality while tiny in size. The dataset was synthesized using 4x3090 Ti cards over a period of 500 hours, thanks to Nous-Hermes-Llama2-13b finetuned model.
Why… See the full description on the dataset page: https://huggingface.co/datasets/nampdn-ai/tiny-textbooks.nasa-exoplanets
NASA Exoplanet Archive
Credit: NASA/JPL-Caltech
Part of a dataset collection on Hugging Face.
Dataset description
Confirmed exoplanets with orbital, stellar, and discovery parameters from the NASA Exoplanet Archive.
The NASA Exoplanet Archive is the authoritative database of confirmed exoplanets, maintained by Caltech/IPAC under contract with NASA. Each entry represents a confirmed planet with its best-available physical and orbital parameters, host star… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/nasa-exoplanets.Pluto-Nano-1.0-Pretrain-v2
ASTRAI Pluto Nano 1.0 — Pretrain Mix (v2)
Curated multilingual pretraining corpus (~50 GB parquet, ~12 B tokens after tokenization) used for ASTRAI Pluto Nano 1.0, a 1 B-total / 50 M-active MoE model with 64 k vocabulary and 5 target languages (EN, PT, ES, ZH, HI).
v2 additions vs v1: OpenThoughts3 (CoT reasoning), openstax textbooks + peS2o (science), and reweighting for better balance. NOTE: factsense (openbmb) was used at training time but is not redistributed here due to its… See the full description on the dataset page: https://huggingface.co/datasets/ASTRAI-labs/Pluto-Nano-1.0-Pretrain-v2.africa-owid-natural-gas-proved-reserves
Natural Gas Proved Reserves | Africa (Our World in Data) | Africa (Electric Sheep Africa metadata inventory)
Size category: n<1K - Formats: parquet - Sector: energy - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-owid-natural-gas-proved-reserves.Sombench-pretraining-data
SomBench Pre-training Corpus: Multimodal Lunar Tiles
Dataset Summary
This includes a small sample from SomBench: a corpus of co-registered, multimodal lunar image tiles built for
large-scale self-supervised (foundation-model) pre-training. It contains a subset of modalities from the
low-resolution (WAC-anchored) and high-resolution (NAC-anchored) tracks specifically used in pretraining.
Tiles are anchored to individual LROC Experiment Data Record (EDR) image… See the full description on the dataset page: https://huggingface.co/datasets/nasa-ibm-ai4science/Sombench-pretraining-data.Sombench-WAC-Crater-Detection
SomBench Benchmark: Robbins Crater Detection, WAC
Science theme: Impact processes
Task: Object detection
Dataset Summary
An impact-crater object-detection benchmark built from the
Robbins (2019) global lunar crater
catalog, a manually compiled, near-complete census of
lunar impact craters (≥ ~1–2 km). Catalog crater centers and diameters are
converted to bounding boxes and packaged over LROC WAC visible tiles drawn
from the pre-training corpus test split, in COCO… See the full description on the dataset page: https://huggingface.co/datasets/nasa-ibm-ai4science/Sombench-WAC-Crater-Detection.AraMix-Native
AraMix-Native
A native-Arabic-filtered version of
AdaMLLab/AraMix (minhash_deduped),
derived from SultanR/AraMix-Translation-Scores:
machine-translated and garbled-MT documents removed, 162,887,010 rows kept of
178,883,241 (91.06%). All columns preserved.
Filter rules
A document is kept iff all of:
mmbert_translated_score < 0.1, or a classical-text rescue: diacritic
(tashkeel) ratio ≥ 0.02 over Arabic letters and ≥ 3 distinct diacritic
classes (fully/partially… See the full description on the dataset page: https://huggingface.co/datasets/SultanR/AraMix-Native.note-articles
note articles
note.com の公開ページから抽出した本文と IPADIC による MeCab 解析結果です。
各行は text と、構造化されたトークン列 mecab を持ちます。有料記事は公開されている範囲のみです。
grip_oThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
11
],
"names": [
"vel_x",
"vel_y",
"vel_z",
"room_vel_x",
"room_vel_y",
"wrist_speed",
"finger_speed"… See the full description on the dataset page: https://huggingface.co/datasets/naavox/grip_o.Natural-Reasoning-STEM-25Kreally_long_task_name_for_testing_20260729_111042This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
6
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/makermods/really_long_task_name_for_testing_20260729_111042.Online-RetailClawBench
ClawBench — A Benchmark for AI Web Agents
Can AI Agents Complete Everyday Online Tasks?
|💻 Github | 🏆 Leaderboard | 📖 Paper | 🌐 Website |
ClawBench is an open benchmark for AI web agents — the systems that drive a real browser to complete a user's task end-to-end. It scores agents on real, everyday online tasks (booking flights, ordering groceries, submitting job applications) across live websites. The corpus ships in two slices: V1 — 153 tasks across 144 websites (the original… See the full description on the dataset page: https://huggingface.co/datasets/NAIL-Group/ClawBench.
