datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sorry-bench-202503
Dataset Card for SORRY-Bench Dataset (2025/03)
🏠Website
📑Paper
📚Dataset
💻Github
🧑⚖️Human Judgment Dataset
🤖Judge LLM
🪧UPDATE: In this iteration, we removed the category "Impersonation" due to its ambiguous definition, and that most models more or less fulfill such requests.This dataset contains 9.2K potentially unsafe instructions, intended to be used for LLM safety refusal evaluation.
Particularly, our base dataset consists of 440 unsafe… See the full description on the dataset page: https://huggingface.co/datasets/sorry-bench/sorry-bench-202503.OSWorld2
OSWorld2 Kimi-K2.6 exact request traces
This dataset contains 98 complete OSWorld2 task traces recorded from a Kimi-K2.6 EPD 1P1D serving run. It is intended for exact virtual request replay in SGLang; replay sends the recorded model requests and token sequences without launching OSWorld VMs or re-executing computer actions.
Coverage
98 successful tasks
16,100 model requests
13,100 unique content-addressed screenshot blobs
6.72 GB of blob payloads before Hub… See the full description on the dataset page: https://huggingface.co/datasets/ergt2025/OSWorld2.aime_2025
AIME 2025 - Unified Test-Time Scaling Format
This is the AIME (American Invitational Mathematics Examination) 2025 dataset in a unified format for test-time scaling experiments.
Dataset Description
Source: MathArena/aime_2025
Size: 30 competition-level mathematics problems
Format: Unified TTS format (question, answer, metadata)
Dataset Structure
Fields
question (string): The mathematical problem statement
answer (string): The numerical answer… See the full description on the dataset page: https://huggingface.co/datasets/test-time-compute/aime_2025.UltraData-Math-TEST
UltraData-Math
🤗 Dataset | 💻 Source Code | 🇨🇳 中文 README
UltraData-Math is a large-scale, high-quality mathematical pre-training dataset totaling 290B+ tokens across three progressive tiers—L1 (170.5B tokens web corpus), L2 (33.7B tokens quality-selected), and L3 (88B tokens multi-format refined)—designed to systematically enhance mathematical reasoning in LLMs. It has been applied to the mathematical pre-training of the MiniCPM Series models.
🆕 What's New… See the full description on the dataset page: https://huggingface.co/datasets/Alga2025/UltraData-Math-TEST.Eurovoc_2025_by_language
🇪🇺 🏷️ EuroVoc dataset (by language)
This is the EuropeanParliament/Eurovoc_2025 dataset, but split up by language, not by period.
The original is split up into periods (1996-03 through 2025-11), with documents in different languages mixed together.
For ease of training this dataset splits the data by language instead, with documents in different periods put together.
License
This dataset is redistributed under the original European Union Public License 1.2. When… See the full description on the dataset page: https://huggingface.co/datasets/AIStudioDelta/Eurovoc_2025_by_language.FineWeb2025
FineWeb-Edu 2025 — Cleaned and Shuffled
This dataset is a year-specific, cleaned, shuffled, and sharded release derived from HuggingFaceFW/fineweb-edu. It contains English educational web text collected in Common Crawl snapshots whose dump identifier belongs to 2025.
This is not a news-only dataset. The year refers to the Common Crawl capture year, not necessarily the page's publication year.
Dataset summary
Item
Value
Year
2025
Rows
99,022,205… See the full description on the dataset page: https://huggingface.co/datasets/Yahoo-Finance-News/FineWeb2025.reddit_dataset_2025
Bittensor Subnet 13 Reddit Dataset
Miner Data Compliance Agreement
In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy.
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed Reddit data. The data is continuously updated by network miners, providing a real-time stream of Reddit content for various analytical and machine learning tasks.
For more… See the full description on the dataset page: https://huggingface.co/datasets/goldentraversy07/reddit_dataset_2025.TamilNadu-State-board-books-2025-Tamil-and-english-versions
Tamil Nadu School Textbooks — Tamil and English Structured Text
This dataset contains text extracted from 312 Tamil Nadu State Board school
textbooks for Standards 1–12. It covers Tamil- and English-medium books and
provides each retained book in two forms:
structured JSON with book metadata, ordered sections, typed content blocks,
source references, extraction statistics, and curation provenance;
Markdown for reading, inspection, and downstream text processing.
The source… See the full description on the dataset page: https://huggingface.co/datasets/Joemn/TamilNadu-State-board-books-2025-Tamil-and-english-versions.HMMT_2025
Dataset Summary
This dataset comprises the questions, answers, and solutions from HMMT February 2025, all of which were extracted by OCR, converted to LaTeX, and manually verified by FlagEval Team.
Data Fields
Below one can find the description of each field in the dataset.
id (str): Index of the problem in the competition
problem (str): Full problem statement
answer (str): Ground-truth answer to the question
solution(str): Ground-truth solution to the question… See the full description on the dataset page: https://huggingface.co/datasets/FlagEval/HMMT_2025.crypto-news-coindesk-2020-2025
CoinDesk Cryptocurrency News Dataset (2020–2025)
This dataset contains cryptocurrency-related news articles sourced from CoinDesk Data, accessed programmatically via the CryptoCompare API. The dataset is curated and published for academic and research purposes, with a focus on analyzing the relationship between news and cryptocurrency market dynamics.
Time Period
January 1, 2020 – January 1, 2025
Content Overview
Each record in the dataset… See the full description on the dataset page: https://huggingface.co/datasets/maryamfakhari/crypto-news-coindesk-2020-2025.TvTroper-2025
TvTroper-2025
A cleaned & refreshed dump of ~708 k pages from tvtropes.org
Dataset Summary
TvTroper-2025 is an updated snapshot of TvTropes.org (≈ 708 000 wiki pages, namespaces and date-grouped pages excluded).
Every page is released in two flavours:
Raw HTML – 22 GB single file
Markdown-cleaned – split into 1 GB JSONL shards (no unpacking required)
No additional content filtering has been applied; short sub-index pages are left in so you can decide what to drop.… See the full description on the dataset page: https://huggingface.co/datasets/KaraKaraWitch/TvTroper-2025.yc-companies-august-2025
Y Combinator Companies Dataset
Dataset Description
This dataset contains information about 5,404 Y Combinator funded companies that have been publicly launched, sourced from the YC-OSS-API.
Dataset Summary
Total Companies: 5,404
Time Range: Summer 2005 - Summer 2025
Update Frequency: Snapshot from August 2025
Source: YC-OSS-API
Dataset Structure
Data Fields
id: Unique identifier for each company
name: Company name… See the full description on the dataset page: https://huggingface.co/datasets/jeffboudier/yc-companies-august-2025.scored_co_2025
Scored.co 2025 scrape
This dataset is a full scrape of public content from Scored.co, a right-wing Reddit-style social media site. It contains 73,045,361 rows of posts and comments, stored as parquet.
The dataset is intended for research into online communities, political discussion, social media moderation, misinformation, platform migration, network dynamics, and large-scale text analysis.
Data format
The main dataset is partitioned by entity type:… See the full description on the dataset page: https://huggingface.co/datasets/trentmkelly/scored_co_2025.llm0to1-pt-tokenized-en-edu-2025
LLM0to1 사전학습 토큰화본 — 영어 교육(2025 덤프)
10B 규모 한/영 이중언어 LLM LLM0to1-10b 를 바닥부터 학습할 때 실제로 투입된
영어 교육(2025 덤프) 코퍼스의 토큰화본. 총 28종 / 138.4B 토큰.
왜 원문 텍스트가 아니라 토큰화본인가
이 코퍼스들은 공개 데이터셋을 스트리밍으로 받아 곧바로 토큰화했고 중간 텍스트를 보관하지 않았다.
따라서 이 .ds 파일이 학습에 들어간 데이터의 유일한 사본이다.
단점만 있는 건 아니다. 토큰화본은 학습 입력 그 자체이므로,
재토큰화 과정에서 생길 수 있는 차이 없이 학습을 그대로 재현할 수 있다.
원본 출처
HuggingFaceFW/fineweb-edu 2025 덤프
영어 벤치마크 하락에 대응해 뒤늦게 편입한 '다양성 영어' 보강분이다. g00~`g15` 는 원본 shard 를 균등 분할한 것으로, 서로 다른 문서 집합이다.… See the full description on the dataset page: https://huggingface.co/datasets/izlley2/llm0to1-pt-tokenized-en-edu-2025.neurips-2025-papers
NeurIPS 2025 Papers Dataset
This dataset contains all accepted papers from NeurIPS 2025, scraped from OpenReview.
Dataset Statistics
Overview
Total Papers: 5772
Unique Paper IDs: 5772
✅ No duplicate IDs
Track Distribution
Main Track: 5,275 papers (91.4%)
Datasets and Benchmarks Track: 497 papers (8.6%)
Award Distribution
Poster: 4,949 papers (85.7%)
Oral: 84 papers (1.5%)
Spotlight: 739 papers (12.8%)
Track × Award… See the full description on the dataset page: https://huggingface.co/datasets/huyxdang/neurips-2025-papers.IDX_Financial_Statements2015-2025Q2
IDX_Financial_Statements: The Multimodal Indonesian Financial Dataset
IDX_Financial_Statements is a centralized repository providing the most complete set of financial disclosures for public companies listed on the Indonesia Stock Exchange (IDX). This dataset is designed for advanced financial research, spanning from raw document archival to structured data extraction.
Dataset Overview
This is a multimodal dataset that captures the full lifecycle of financial reporting.… See the full description on the dataset page: https://huggingface.co/datasets/horelulus/IDX_Financial_Statements2015-2025Q2.whiteglove-medical-medlineplus-2025
WhiteGlove Medical Knowledge Corpus
MedlinePlus 2025 — Spectral Curation Pipeline
Pipeline: WhiteGlove Spectral Curation | Domain: Medical | License: Public Domain (US Government)
Dataset Summary
A clean, deduplicated, semantically chunked medical knowledge corpus derived from the NIH MedlinePlus January 2025 ZIM archive. Produced by the WhiteGlove Spectral Curation Pipeline — an air-gapped, attribution-clean dataset factory built on SimHash-128 deduplication… See the full description on the dataset page: https://huggingface.co/datasets/joecwales/whiteglove-medical-medlineplus-2025.enwiki-20250201
Dataset Card for lparkourer10/enwiki-20250201
This dataset is an extracted version of the English Wikipedia dump as of February 1, 2025. It has been processed to facilitate information retrieval and analysis.
Dataset Description
This dataset contains extracted text from the English Wikipedia, aimed at providing structured and accessible information for natural language processing (NLP) tasks, research, and machine learning applications. It includes raw Wikipedia articles… See the full description on the dataset page: https://huggingface.co/datasets/lparkourer10/enwiki-20250201.yuxiaowang-prompts-2025
Yuxiaowang Semantic Dataset · Hugging Face Version
🧠 English Summary
Yuxiaowang · Semantic Dataset for Japanese Language Schools (Chinese)
This project provides structured semantic definitions and prompt examples for the domain of Japanese language schools in China.It aims to serve as a grounding corpus for large language models (LLMs) to understand terms like "语校", "语校网", and related concepts.
Source platform: https://www.yuxiaowang.comAll prompts and term… See the full description on the dataset page: https://huggingface.co/datasets/languagehub-ai/yuxiaowang-prompts-2025.Wiki-zhtw-20250601
Dataset Card for Wiki-zhtw-20250601
Dataset Description
This dataset is derived from the Chinese‑Wikipedia dump dated 2025‑06‑01, downloaded from Wikimedia.Articles were extracted from the original .xml.bz2 archive with Gensim, converted to Markdown format via regular‑expression post‑processing, and finally converted from Simplified to Traditional Chinese using OpenCC.
WWTD-2025
What Would Trump Do?
Auto-generated from 5 search queries — used to beat GPT-5
Starting from nothing but 5 search queries, we used the Lightning Rod SDK to automatically generate 2,790 forecasting questions about Trump administration actions from news articles and label them using real outcomes. No expertise required. No manual labeling. Used to train Trump-Forecaster, which beats GPT-5.
TL;DR
Generated 2,790 forward-looking forecasting questions… See the full description on the dataset page: https://huggingface.co/datasets/LightningRodLabs/WWTD-2025.qwen3.8-max-glm5.2-kimi-k3-distillation
Multi-Teacher Distillation Dataset (57,937 traces)
A quality-filtered, deduplicated, multi-teacher SFT corpus combining traces from three frontier models across math, code, reasoning, instruction-following, tool-use, science, long-context, multilingual, and creative dialogue domains.
Teachers
Teacher
Provider
Traces
Qwen3.8-Max-Preview
Alibaba Cloud Model Studio
48,283
GLM-5.2
Z.AI Coding Plan
5,307
Kimi Code K3
Moonshot AI (Kimi)
4,347… See the full description on the dataset page: https://huggingface.co/datasets/Jinzy2025/qwen3.8-max-glm5.2-kimi-k3-distillation.All-CVE-Chat-MultiTurn-1999-2025-Dataset
CVE Chat‑Style Multi‑Turn Cybersecurity Dataset (1999 – 2025)
1. Project Overview
This repository hosts the largest publicly available chat‑style, multi‑turn cybersecurity dataset to date, containing ≈ 300 000 Common Vulnerabilities and Exposures (CVE) records published between 1999 and 2025. Each record has been meticulously parsed, enriched, and converted into a conversational format that is ideal for training and evaluating AI and AI‑Agent systems focused on… See the full description on the dataset page: https://huggingface.co/datasets/Trendyol/All-CVE-Chat-MultiTurn-1999-2025-Dataset.x_dataset_2025
Bittensor Subnet 13 X (Twitter) Dataset
Miner Data Compliance Agreement
In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy.
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning… See the full description on the dataset page: https://huggingface.co/datasets/goldentraversy07/x_dataset_2025.Health-Bench-Eval-OSS-2025-07
Dataset Card for HealthBench
Dataset Summary
HealthBench is a benchmark dataset developed by OpenAI in collaboration with 262 physicians from 60 countries to evaluate AI systems in health-related conversational scenarios. It contains 5,000 multi-turn health conversations in a JSONL file (2025-05-07-06-14-12_oss_eval.jsonl), simulating interactions between AI models and users (laypersons or clinicians). Each conversation includes a user prompt, a candidate model response… See the full description on the dataset page: https://huggingface.co/datasets/Tonic/Health-Bench-Eval-OSS-2025-07.hle-extract-qwen3235ba22b-20250815
HLE Extract: Qwen3-235B-A22B Evaluation Results (2025-08-15)
Dataset Description
This dataset contains the complete Human-Level Evaluation (HLE) benchmark with detailed evaluation results from the Qwen/Qwen3-235B-A22B model. It merges the original team-suzuki/hle-extract dataset with comprehensive model responses and human judgments.
Dataset Summary
Total Questions: 120 (complete HLE dataset)
Evaluated Questions: 103 (85.8%)
Unevaluated Questions: 17 (14.2%)… See the full description on the dataset page: https://huggingface.co/datasets/team-suzuki/hle-extract-qwen3235ba22b-20250815.MMdeepresearchHuggingface: https://huggingface.co/papers/2601.12346
Paper: arxiv.org/abs/2601.12346
x_dataset_202507
Bittensor Subnet 13 X (Twitter) Dataset
Miner Data Compliance Agreement
In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy.
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning… See the full description on the dataset page: https://huggingface.co/datasets/goldentraversy07/x_dataset_202507.EU-AI-Regulation-GDPR-2025
EuropeGram: EU Legal Text -- RAG Chunks & Instruction Data
Structured, chunked, and instruction-formatted text derived from official EU
legislation, built for retrieval-augmented generation (RAG) and LoRA
fine-tuning experiments comparing Base / RAG / Fine-tuned / Fine-tuned+RAG
LLM strategies over EU documents. Produced by the EuropeGram project's
extraction -> chunking -> fine-tuning-export pipeline.
Source documents
Document
CELEX ID
Source
Chunks… See the full description on the dataset page: https://huggingface.co/datasets/ApyHTML19/EU-AI-Regulation-GDPR-2025.Wiki-ja-20250601
Dataset Card for Wiki-ja-20250601
Dataset Description
This dataset is derived from the Japan‑Wikipedia dump dated 2025‑06‑01, downloaded from Wikimedia.Articles were extracted from the original .xml.bz2 archive with Gensim and converted to Markdown format via regular‑expression post‑processing.
