datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Aiice
Dataset
Aiice benchmark dataset for Arctic sea ice concentration (SIC) forecasting,
based on OSI-SAF satellite products (CC BY 4.0).
Coverage
Period: October 1978 – April 2026
Resolution: 25 km spatial, daily temporal
Grid: 432×432 (Lambert Azimuthal Equal Area, EPSG:6931)
Source products
Product
Source
Period
OSI-450-a
SMMR, SSM/I, SSMIS
1978–2020
OSI-430-a
SSMIS
2021–Jul 2025
OSI-438
AMSR2
Jul 2025–present… See the full description on the dataset page: https://huggingface.co/datasets/ITMO-NSS/Aiice.gaming-500-hours
Gaming Dataset (gaming-1) — 494.7 Hours
Native PC/console gameplay screen-recordings, organized by game. Each workflow
is one play session, trimmed to pure gameplay — login screens, launchers,
desktop, collection-app references, and any watching/streaming are removed.
In-game menus, lobbies, loading, and cutscenes are retained as part of the session.
Workflows: 776
Total gameplay: 494.7 hours
Distinct games: 168
Clip duration (min): median 24.0, p90 90.9, max 457.7
Platforms:… See the full description on the dataset page: https://huggingface.co/datasets/markov-ai/gaming-500-hours.hplt2_edu_scores
HPLT2-Edu-scores
Dataset summary
HPLT2-JQL-Education is a model-annotated language subset of HPLT2, spanning 35 languages.
Our model-annotations allow for a filtering that achieves higher-quality training outcomes without excessively aggressive data reduction.
The original FW2 heuristic filtering method serves as our baseline, providing reference points for both the volume of retained tokens and downstream model performance.
For example, in the Spanish language case… See the full description on the dataset page: https://huggingface.co/datasets/JQL-AI/hplt2_edu_scores.fw2_edu_scores
Fineweb2-Edu-scores
Dataset summary
FineWeb2-JQL-Education is a model-annotated language subset of FineWeb2, spanning 36 languages.
Our model-annotations allow for a filtering that achieves higher-quality training outcomes without excessively aggressive data reduction.
The original FW2 heuristic filtering method serves as our baseline, providing reference points for both the volume of retained tokens and downstream model performance.
For example, in the Spanish language… See the full description on the dataset page: https://huggingface.co/datasets/JQL-AI/fw2_edu_scores.SuperMemory-VQA
SuperMemoryVQA
SuperMemory-VQA is an egocentric visual question answering benchmark for
evaluating long-horizon memory in augmented reality assistant settings. The
dataset is designed around practical questions a person might ask a wearable
memory assistant, such as where an object was left, what someone said earlier,
whether a planned step was completed, or what happened next in a longer event.
The benchmark contains 4,853 human-verified question-answer pairs grounded in
52.9… See the full description on the dataset page: https://huggingface.co/datasets/OSU-AIoT-MLSys-Lab/SuperMemory-VQA.GMAI-VL-5.5M
GMAI-VL-5.5M Dataset
GMAI-VL-5.5M is a comprehensive, large-scale medical General Medical AI Vision-Language (GMAI-VL) dataset built specifically for training multimodal foundation models in the medical domain. It contains an extraordinary scale of high-quality instructions encompassing over 5.5 million multimodal question-answering pairs, carefully constructed based on hundreds of medical classification, segmentation, and detection datasets.
This repository… See the full description on the dataset page: https://huggingface.co/datasets/General-Medical-AI/GMAI-VL-5.5M.africa-synth-aid-flows-medical-multimodal-fracture-all
Africa Synth Aid Flows Medical Multimodal Fracture All | Africa (Electric Sheep Africa metadata inventory)
Size category: 1K<n<10K - Formats: json - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Health… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-aid-flows-medical-multimodal-fracture-all.Auto-ClawEval
Auto-ClawEval
Auto-generated agent evaluation benchmark with 1,040 tasks across 104 unique scenarios created by ClawEnvKit.
Statistics
Tasks
1,040
Categories
24
Mock services
20
Task types
API-based (77%) + file-dependent (23%)
Quick Start
# Download
huggingface-cli download AIcell/Auto-ClawEval --repo-type dataset --local-dir Auto-ClawEval
# Evaluate with ClawEnvKit (Docker harness)
bash run_harnesses.sh --harness claudecode… See the full description on the dataset page: https://huggingface.co/datasets/AIcell/Auto-ClawEval.ai-ecosystem-daily
TensorFeed AI Ecosystem Daily
Daily snapshots of the AI ecosystem: news, model pricing, benchmarks, service status, GPU rental prices, MCP registry growth, LLM endpoint latency probes, agent traffic, and the AFTA adopter directory. Captured once per day from the public tensorfeed.ai API and committed to this repo as JSONL.
Each daily snapshot lives in a YYYY-MM-DD/ subfolder with one JSONL file per feed plus a manifest.json summarizing what was captured.
What's in… See the full description on the dataset page: https://huggingface.co/datasets/tensorfeed/ai-ecosystem-daily.ATBench
ATBench: Agent Trajectory Safety Benchmark Family
💻 GitHub |
📄 ATBench Paper |
📄 AgentDoG Paper (ATBench500) |
🤗 Hugging Face Collection
ATBench is a family of trajectory-level safety benchmarks for long-horizon, tool-using AI agents. The latest release is introduced in ATBench: A Diverse and Realistic Agent Trajectory Benchmark for Safety Evaluation and Diagnosis. This repository now follows a versioned naming scheme:
ATBench: the latest 1… See the full description on the dataset page: https://huggingface.co/datasets/AI45Research/ATBench.thai_exam
Dataset Card for Thai_Exam
ThaiExam is a Thai knowledge benchmarking dataset, consisting of multiple-choice questions from examinations in Thailand. The dataset was originally developed for evaluating Typhoon (Thai LLM). This dataset contains 5 splits corresponding to 5 examinations as follows:
ONET: The Ordinary National Educational Test (ONET) is an examination for students in Thailand. This dataset is based on the grade-12 ONET exam, comprising 4 subjects and each question has 5… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/thai_exam.JQL-LLM-Edu-Annotations
📚 JQL Educational Quality Annotations from LLMs
This dataset provides 17,186,606 documents with high-quality LLM annotations for evaluating the educational value of web documents, and serves as a benchmark for training and evaluating multilingual LLM annotators as described in the JQL paper.
📝 Dataset Summary
Multilingual document-level quality annotations scored on a 0–5 educational value scale by three state-of-the-art LLMs:
Gemma-3-27B-it, Mistral-3.1-24B-it… See the full description on the dataset page: https://huggingface.co/datasets/JQL-AI/JQL-LLM-Edu-Annotations.SEED-Data-Edit-Part1-Openimages
SEED-Data-Edit
SEED-Data-Edit is a hybrid dataset for instruction-guided image editing with a total of 3.7 image editing pairs, which comprises three distinct types of data:
Part-1: Large-scale high-quality editing data produced by automated pipelines (3.5M editing pairs).
Part-2: Real-world scenario data collected from the internet (52K editing pairs).
Part-3: High-precision multi-turn editing data annotated by humans (95K editing pairs, 21K multi-turn rounds with a maximum of 5… See the full description on the dataset page: https://huggingface.co/datasets/AILab-CVC/SEED-Data-Edit-Part1-Openimages.swe-prbench
SWE-PRBench
Benchmarking AI Code Review Quality Against Human Pull Request Feedback
Blog: Read the blog
GitHub Repository: View the code
arXiv Paper: View the paper
Overview
SWE-PRBench is a benchmark of 350 pull requests with human-annotated
ground truth for evaluating whether LLMs can identify the same issues
that real human reviewers flag in production code.
Existing benchmarks like SWE-Bench measure whether models can produce
correct code. SWE-PRBench… See the full description on the dataset page: https://huggingface.co/datasets/foundry-ai/swe-prbench.gspc-ai-economy-index
GSPC — ai adoption components facts (Eurostat)
SWIFT census (live): https://councilof.ai/api/swift
XRPL reader (live): https://councilof.ai/api/xrpl
Live axis name: ai-adoption-components — MEASURED as two Eurostat series (deterministic-facts, n=2). Not an index. No composite formula. No MEASURED-INDEX-v0.1 sticker (C-2026-0826-05: do not restore).
This Hub repo id keeps the legacy slug gspc-ai-economy-index for inbound links only. Cite the live axis name. Do not stamp an… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-ai-economy-index.thai-aligner-bench
Thai Aligner Bench
🚧 Development in progress.
How accurately can a forced aligner place Thai token and word boundaries in
speech? This is a self-contained benchmark: one Python file
(aligner_bench.py) plus 1,572 clips of Thai speech with frame-exact timing
ground truth. No Thai NLP stack or other code is needed — just
numpy soundfile torch torchaudio transformers.
The ground truth is what makes the dataset useful: the audio was rendered by a
TTS model whose duration predictor… See the full description on the dataset page: https://huggingface.co/datasets/wayu-ai/thai-aligner-bench.DecodingTrust
DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models
Overview
This repo contains the source code of DecodingTrust. This research endeavor is designed to help researchers better understand the capabilities, limitations, and potential risks associated with deploying these state-of-the-art Large Language Models (LLMs). See our paper for details.
DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models
Boxin Wang, Weixin Chen, Hengzhi… See the full description on the dataset page: https://huggingface.co/datasets/AI-Secure/DecodingTrust.voice-code-bench
VoiceCodeBench
VoiceCodeBench is a test-only benchmark for evaluating whether automatic
speech recognition (ASR) systems preserve exact structured values in English
workplace speech.
Paper: VoiceCodeBench: Evaluating Exact Structured-Token Recovery in Automatic Speech Recognition
The benchmark targets cases where a transcript is software input: callback
numbers, email addresses, command-line flags, file paths, URLs, account
identifiers, dates, measurements, and similar values… See the full description on the dataset page: https://huggingface.co/datasets/besimple-ai/voice-code-bench.traffic-sign-bench
Traffic Sign Bench
Official per-sign SUMO maps for TrafficRuleBench: real Moscow OSM
layouts, 25 signs, 2500 maps. Protocol size is
80 train + 20 test maps per sign.
Road geometry is derived from OpenStreetMap
© OpenStreetMap contributors and is released under ODbL 1.0.
Download
All scenes land under data/scenes/<sign>/<scene_id>/, which is what eval
expects:
huggingface-cli download emb-ai/traffic-sign-bench \
--repo-type dataset \… See the full description on the dataset page: https://huggingface.co/datasets/emb-ai/traffic-sign-bench.political-bias-in-ai
Political Bias in AI — Where the Major AI Models Stand
An open, monthly measurement of where the major AI models land on value‑loaded political and
ethical questions. Each model is asked the same battery of questions many times, with web
search turned off, so the result reflects the trained weights rather than whatever the model
retrieves that day. Every answer is classified by a neutral coder onto a left–right economic
axis and a libertarian–authoritarian social axis, and… See the full description on the dataset page: https://huggingface.co/datasets/trakkr-ai/political-bias-in-ai.ro-aya_collectionThis dataset is a translation of CohereLabs/aya_collection, an instruction dataset, using LLMic, a bilingual Romanian-English LLM.
The Aya Collection is a massive multilingual collection consisting of 513 million instances of prompts and completions covering a wide
range of tasks. This collection incorporates instruction-style templates from fluent speakers and applies them to a curated
list of datasets, as well as translations of instruction-style datasets into 101 languages.
Only the… See the full description on the dataset page: https://huggingface.co/datasets/faur-ai/ro-aya_collection.jbdprompt-swap-mixed12-5xlr-e1-mxfp4-mergedprompt-swap-mixed12-5xlr-e2-mxfp4-mergedcurated_edu_scoresstruct-ir
SSRB: Direct Natural Language Querying to Massive Heterogeneous Semi-Structured Data
github
We employ LLM-based automatic evaluation and build a large-scale semi-structured retrieval benchmark (SSRB) using LLM generation and filtering, containing 14M structured objects from 99 different schemas across 6 domains, along with 8,485 test queries that combine both exact and fuzzy matching conditions.
This repository contains the data for SSRB.
Data Download
Data can be… See the full description on the dataset page: https://huggingface.co/datasets/vec-ai/struct-ir.AgentJudgeBench
AgentJudgeBench: Evaluating LLM Judge Reliability on Agentic Tool-Calling
A benchmark for systematically evaluating how reliably LLM judges assess
agentic tool-calling workflows across structured, dependency-driven tasks.
Why this benchmark?
AgentJudgeBench measures how reliably LLM judges assess agentic tool-calling outputs. It provides 3,808 benchmark records spanning six DAG topologies and three difficulty… See the full description on the dataset page: https://huggingface.co/datasets/ServiceNow-AI/AgentJudgeBench.smol-worldcup
🏟️ Smol AI WorldCup — SHIFT Benchmark
The world's first 5-axis evaluation framework for small language models.
Not just "how smart?" — but "how honest? how fast? how small? how efficient?"
🏟️ Leaderboard
huggingface.co/spaces/ginigen-ai/smol-worldcup
📊 Dataset
huggingface.co/datasets/ginigen-ai/smol-worldcup
🏅 ALL Bench
huggingface.co/spaces/FINAL-Bench/all-bench-leaderboard
🏆 Official Ranking: WCS (WorldCup Score)
WCS = √( SHIFT × PIR_norm )… See the full description on the dataset page: https://huggingface.co/datasets/ginigen-ai/smol-worldcup.Auto-ClawEval-mini
Auto-ClawEval-mini
Compact agent evaluation benchmark with 104 tasks created by ClawEnvKit.
Statistics
Tasks
104
Categories
24
Mock services
20
Task types
API-based (77%) + file-dependent (23%)
Quick Start
# Download
huggingface-cli download AIcell/Auto-ClawEval-mini --repo-type dataset --local-dir Auto-ClawEval-mini
# Evaluate with ClawEnvKit (Docker harness)
bash run_harnesses.sh --harness claudecode --dataset Auto-ClawEval-mini… See the full description on the dataset page: https://huggingface.co/datasets/AIcell/Auto-ClawEval-mini.airbnb_embeddings
Overview
This dataset consists of AirBnB listings with property descriptions, reviews, and other metadata.
It also contains text embeddings of the property descriptions as well as image embeddings of the listing image. The text embeddings were created using OpenAI's text-embedding-3-small model and the image embeddings using OpenAI's clip-vit-base-patch32 model available on Hugging Face.
The text embeddings have 1536 dimensions, while the image embeddings have 512 dimensions.… See the full description on the dataset page: https://huggingface.co/datasets/MongoDB/airbnb_embeddings.
