datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
chm-corr-prj-giangmmlu-prox-eval-predictions
MMLU-ProX Multilingual Model Predictions
Raw per-sample model predictions on MMLU-ProX
across 29 languages and 25 open-weight LLMs, produced with
lm-evaluation-harness.
This dataset releases the full prediction logs (not just aggregate scores) so that
item-level responses can be re-analysed — e.g. for Item Response Theory (IRT) modelling
of multilingual benchmarks, error analysis, or per-item difficulty estimation.
Repository structure
mmlu_prox_<lang>/
└──… See the full description on the dataset page: https://huggingface.co/datasets/gililior/mmlu-prox-eval-predictions.GUIOdyssey
Dataset Card for GUIOdyssey
Repository: https://github.com/OpenGVLab/GUI-Odyssey
Paper: https://arxiv.org/pdf/2406.08451
News⭐️
Latest version of GUIOdyssey released!🎉
This updated version features a larger dataset with 8,334 episodes, as well as richer semantic annotations. Compared to the previous version, we have added more fine-grained low-level instructions, image descriptions, action intentions, and context review for each step. Additionally, we provide bounding… See the full description on the dataset page: https://huggingface.co/datasets/hflqf88888/GUIOdyssey.gspc-hub-cards
GSPC hub cards — mill cards, not board axes
SWIFT census (live): https://councilof.ai/api/swift
XRPL reader (live): https://councilof.ai/api/xrpl
One row per signed measurement card: one model, one axis, one date, Ed25519 over the body. A row is MEASURED only when a signed card verifies. Absent (model, axis) pairs are absent — not zero.
Measurement, not certification. Cards are evidence of bytes on a frozen bank at a time — never approval, rating, or safety guarantee.… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-hub-cards.ipo-text
SEC IPO Filings Dataset
A large-scale, comprehensive dataset of 100,000+ filings (S-1 and F-1 filings) filed with the SEC EDGAR system, spanning 1994–2026 and over 20,000 unique registrants.
Every filing has been downloaded and then parsed using the IPO-Mine Python Package. We have extracted three common sections found in these documents (Prospectus Summary, Risk Factors, Legal Matters), and then used an LLM classifier to group them into three categories. For this dataset, we have… See the full description on the dataset page: https://huggingface.co/datasets/gtfintechlab/ipo-text.GUIGuard-Bench
GUIGuard-Bench (Public Ladder)
GUIGuard-Bench is a cross-platform GUI agent benchmark for studying privacy risks and privacy-preserving execution in multimodal GUI agents.
This public-ladder release contains 121 GUI interaction trajectories (68 Android + 53 PC) for benchmark evaluation, with 26,407 region-level privacy annotations across 2,002 screenshots.
For the anonymous review version of the evaluation toolkit, see GUIGaurd-Bench-CA4F.
Dataset Summary
GUI agents… See the full description on the dataset page: https://huggingface.co/datasets/ShaofantuoshuzhengzhiSha/GUIGuard-Bench.gaming-500-hours
Gaming Dataset (gaming-1) — 494.7 Hours
Native PC/console gameplay screen-recordings, organized by game. Each workflow
is one play session, trimmed to pure gameplay — login screens, launchers,
desktop, collection-app references, and any watching/streaming are removed.
In-game menus, lobbies, loading, and cutscenes are retained as part of the session.
Workflows: 776
Total gameplay: 494.7 hours
Distinct games: 168
Clip duration (min): median 24.0, p90 90.9, max 457.7
Platforms:… See the full description on the dataset page: https://huggingface.co/datasets/markov-ai/gaming-500-hours.LEMUR
EU Law Dataset – Category 15.10: Environment
This dataset contains official legal documents from the European Union, collected from the EUR-Lex website, specifically under category 15.10: "Environment". The documents span from the year 1961 to 2025 and are provided in multiple European "languages. The original documents are in PDF format and have been converted into various text-based formats using OLMCR.
The dataset splits represent the different "languages available for each… See the full description on the dataset page: https://huggingface.co/datasets/G4KMU/LEMUR.gspc-boards
GSPC signed board archive
SWIFT census (live): https://councilof.ai/api/swift
XRPL reader (live): https://councilof.ai/api/xrpl
Snapshot/archive of board payloads. Printer language only.
Not a certificate. Hub is a printer of live GET, not a second engine. Fetch fail → UNCHECKABLE.
The live board is the authority
GET https://councilof.ai/api/gspc — quote totals.public_count. This Hub card is a printer of that GET, never a second
engine. If the fetch fails… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-boards.GSM-Symbolic
GSM-Symbolic
This project accompanies the research paper, GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models.
Getting Started
In our first release, we provide data for GSM-Symbolic, GSM-Symbolic-P1, and GSM-Symbolic-P2 variants. For each variant, we have released both templates (available on Github), and a sample of generated data that can be used for evaluation.
To load the data, you can use the following code. Note that in… See the full description on the dataset page: https://huggingface.co/datasets/apple/GSM-Symbolic.faience-games
Faïence: human-vs-net Azul games
Every game played on Faïence, a
free browser implementation of the rules of Azul (Michael Kiesling) against
a neural net trained by self-play, unless the player switched sharing off.
This dataset is the training pile the playing page tells its players about,
and it is public precisely so that a player can read everything the project
collects. Records are anonymous by construction: moves, deals, which net
played, and the score. No names, no… See the full description on the dataset page: https://huggingface.co/datasets/RemiFabre/faience-games.rl-run-archive-2026
RL run archive 2026
Archived raw run artifacts (rollout trajectories, rendered frames, policy and optimizer
checkpoints, configs, logs) from simulation reinforcement-learning experiments, published for
long-term preservation and reproducibility.
Layout mirrors the verified backup trees they were copied from:
tilde/20260915-102000/ and taurus/20260915-085631/: batched tar archives. Every archive
carries a per-file SHA-256 manifest inside it; the batch inventories (9998.json.gz… See the full description on the dataset page: https://huggingface.co/datasets/gavinlaw/rl-run-archive-2026.t2-ragbench
Dataset Card for T2-RAGBench
Project Page | Paper | Code
IMPORTANT NOTICE:
We deleted VQAonBD from the dataset due to low quality of the question reformulations. If you still want to use it you will find the data in the previous commit history.
Dataset Description
Dataset Summary
T2-RAGBench is a benchmark dataset designed to evaluate Retrieval-Augmented Generation (RAG) on financial documents containing both text and tables. It consists of 23,088… See the full description on the dataset page: https://huggingface.co/datasets/G4KMU/t2-ragbench.MAPBench-V2For more details, please check our project page.
Paper: https://arxiv.org/abs/2601.05432
Repository: https://github.com/AMAP-ML/Thinking-with-Map
Granary
Granary: Speech Recognition and Translation Dataset in 25 European Languages
Granary is a large-scale, open-source multilingual speech dataset covering 25 European languages for Automatic Speech Recognition (ASR) and Automatic Speech Translation (AST) tasks.
Overview
Granary addresses the scarcity of high-quality speech data for low-resource languages by consolidating multiple datasets under a unified framework:
🗣️ ~1M hours of… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Granary.GUI-Odyssey
Dataset Card for GUI Odyssey
News⭐️
A new and improved version of the GUIOdyssey dataset has been released! 🎉🎉
👉 Please use the latest version and refer to the updated README for the most up-to-date information.
We highly recommend using the new version for all training and evaluation!
Repository: https://github.com/OpenGVLab/GUI-Odyssey
Latest Version of Dataset: hflqf88888/GUIOdyssey
Paper: https://arxiv.org/pdf/2406.08451
Introduction
GUI Odyssey is… See the full description on the dataset page: https://huggingface.co/datasets/OpenGVLab/GUI-Odyssey.matterport3d_region_mcmc_3dgsLicense Notice:This dataset is derived from Matterport3D.It follows the Matterport End User License Agreement for Academic Use of Model Data.See Matterport3D License for details.
GMAI-VL-5.5M
GMAI-VL-5.5M Dataset
GMAI-VL-5.5M is a comprehensive, large-scale medical General Medical AI Vision-Language (GMAI-VL) dataset built specifically for training multimodal foundation models in the medical domain. It contains an extraordinary scale of high-quality instructions encompassing over 5.5 million multimodal question-answering pairs, carefully constructed based on hundreds of medical classification, segmentation, and detection datasets.
This repository… See the full description on the dataset page: https://huggingface.co/datasets/General-Medical-AI/GMAI-VL-5.5M.any4hdmi-g1-100stylexauusd-gold-price-historical-data-2004-2025
XAUUSD Gold Price Historical Data 2004-2025
This dataset contains historical price data for XAUUSD (Gold vs US Dollar) from 2004 to 2025.
Source: Kaggle dataset "novandraanugrah/xauusd-gold-price-historical-data-2004-2024"
Content:
The dataset includes CSV files with different time granularities (e.g., 1 minute, 5 minutes, 1 hour, 1 day). Each file typically contains the following columns:
Date
Open
High
Low
Close
Volume
Usage:
This dataset can be used for analyzing historical… See the full description on the dataset page: https://huggingface.co/datasets/ZombitX64/xauusd-gold-price-historical-data-2004-2025.style-dpo
gijl style dataset (multi-type)
Generated by meta-models/Muse-Glimmer-30B through a tool-using scouting loop over real sources (Stack Exchange, GitHub, OSV, Hacker News, arXiv, Wikipedia, web). Splits are a deterministic hash of the record id (90/5/5); derived records inherit their parent's split. Synthetic, model-written, not human-verified. Every rejected response is intentionally poor and must never be used as an example of good behavior.
config
folder
train
validation… See the full description on the dataset page: https://huggingface.co/datasets/gijl/style-dpo.CulturaY-ja-askllm-v1
CulturaY-ja-askllm-v1
多言語データセット ontocord/CulturaY の日本語パート ja に対して、 Ask-LLM 手法でスコア付けしたデータセットです。
元データセットのカラムに加え askllm_score というカラムが追加されており、ここに Ask-LLM のスコアが格納されています。
Ask-LLM でスコア付けに使用した LLM は Rakuten/RakutenAI-7B-instruct で、プロンプトは以下の通りです。
###
{data}
###
Does the previous paragraph demarcated within ### and ### contain informative signal for pre-training a large-language model? An informative datapoint should be well-formatted, contain some usable knowledge of the world, and strictly… See the full description on the dataset page: https://huggingface.co/datasets/geniacllm/CulturaY-ja-askllm-v1.GLM-5.2-AgentThis dataset was generated using teich by TeichAI
GLM-5.2 Agent traces
This directory contains raw agent trace files generated by teich.
JSONL files: 319
Model metadata: glm-5.2
Training-ready tools
Generated agent traces carry configured or recovered tool schemas so tools remain available for training even when a session did not call them.
Native Claude Code imports recover schemas for Claude Code and Claude Desktop built-ins, plus conservative name-derived MCP… See the full description on the dataset page: https://huggingface.co/datasets/AletheiaResearch/GLM-5.2-Agent.logi_glueGSgeometry-dash-levels
Geometry Dash Level Dataset
Subsets
2024_300k
Dump of ~300k levels from the Geometry Dash servers, sorted by the number of likes. 66 JSONL shards (~660MB each, ~43GB total).
Files: 2024_300k/levels-v1-00000.jsonl through 2024_300k/levels-v1-00065.jsonl
2026_50k_rated
~50k rated/featured levels scraped from the Geometry Dash servers in February 2026. 39 JSONL shards (~500MB each, ~19GB total).
Files: 2026_50k_rated/levels-v2-00000.jsonl through… See the full description on the dataset page: https://huggingface.co/datasets/yusp48/geometry-dash-levels.AndroidCodeglenans-isobars-archivegaokao-benchGPT-5.5-CodexThis dataset was generated using teich by TeichAI
GPT-5.5 Agent traces
This directory contains raw agent trace files generated by teich.
JSONL files: 317
Model metadata: gpt-5.5
Training-ready tools
Generated agent traces carry configured or recovered tool schemas so tools remain available for training even when a session did not call them.
Native Claude Code imports recover schemas for Claude Code and Claude Desktop built-ins, plus conservative name-derived MCP… See the full description on the dataset page: https://huggingface.co/datasets/AletheiaResearch/GPT-5.5-Codex.
