datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Nemotron-Terminal-Corpus
Terminal-Corpus: Large-Scale SFT Dataset for Terminal Agents
Terminal-Corpus is a large-scale Supervised Fine-Tuning (SFT) dataset designed to scale the terminal interaction capabilities of Large Language Models (LLMs). Developed by NVIDIA, this dataset was built using the Terminal-Task-Gen pipeline, which combines dataset adaptation with synthetic task generation across diverse domains.
🚀 Key Results & Performance
The high-quality trajectories in Terminal-Corpus enable… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Terminal-Corpus.browsecomp-plus-corpus
BrowseComp-Plus
Project Page | Paper | Code
BrowseComp-Plus is a new benchmark for Deep-Research system, isolating the effect of the retriever and the LLM agent to enable fair, transparent comparisons of Deep-Research agents. The benchmark sources challenging, reasoning-intensive queries from OpenAI's BrowseComp. However, instead of searching the live web, BrowseComp-Plus evaluates against a fixed, curated corpus of ~100K web documents from the web. The corpus includes both… See the full description on the dataset page: https://huggingface.co/datasets/Tevatron/browsecomp-plus-corpus.Legal_Corpus_QA_SynDeepThink
🧠 Legal Corpus QA SynDeepThink Dataset
This repository contains a high-intelligence Legal Question-and-Answer dataset, generated through an advanced Iterative and Recursive Thinking process. It bridges the gap between static legal corpora and the dynamic "check-and-recheck" nature of human legal expertise. 🏛️
💡 The Concept: Iterative & Recursive Legal Logic
While standard synthetic datasets are often generated in a single pass, Legal_Corpus_QA_SynDeepThink mimics the… See the full description on the dataset page: https://huggingface.co/datasets/Azzindani/Legal_Corpus_QA_SynDeepThink.corpus-oct-2024
Dataset Card for FreshStack (Corpus)
Homepage |
Repository |
Paper
FreshStack is a holistic framework to construct challenging IR/RAG evaluation datasets that focuses on search across niche and recent topics.
This dataset (October 2024) contains the query, nuggets, answers and nugget-level relevance judgments of 5 niche topics focused on software engineering and machine learning.
The queries and answers (accepted) are taken from Stack Overflow, GPT-4o generates the nuggets and… See the full description on the dataset page: https://huggingface.co/datasets/freshstack/corpus-oct-2024.Collective-Corpus
🧠 Collective Corpus — Universal Pretraining + Finetuning Dataset (500B+ Tokens)
Collective-Corpus is a massive-scale, multi-domain dataset designed to train Transformer-based language models from scratch and finetune them across a wide variety of domains — all in one place.
📚 Dataset Scope
This dataset aims to cover the full LLM lifecycle, from raw pretraining to domain-specialized finetuning.
1. Pretraining Corpus
Large-scale, diverse multilingual text… See the full description on the dataset page: https://huggingface.co/datasets/dignity045/Collective-Corpus.turkish-corpus-100b
Turkish Corpus 100B (TC-100B)
Dataset Summary
The Turkish Corpus 100B (TC-100B) is a massive-scale, deduplicated, and cleaned dataset designed for training Foundation Models in Turkish. Comprising approximately 105 Billion tokens (measured with Qwen/Llama3 tokenizer), it represents one of the largest open resources for Turkish LLM pretraining.
The dataset is engineered for a two-stage training pipeline:
Pretrain Subset (~103B Tokens): A diverse mix of… See the full description on the dataset page: https://huggingface.co/datasets/hasankursun/turkish-corpus-100b.LiteResearcher-Browse-Corpus
LiteResearcher Browse Corpus
Full web-page text for reproducing the visit / /web_parser tool of LiteResearcher.
What this is
The LiteResearcher local environment exposes two tools, backed by two separate datasets:
Tool
Endpoint
Dataset
Content
search
/search
LiteResearcher-Corpus
url / title / doc — doc is snippet-level (~200 chars), used to build the BGE-M3 vector index
visit
/web_parser
this dataset
url / title / text — text is the full page… See the full description on the dataset page: https://huggingface.co/datasets/simplex-ai-inc/LiteResearcher-Browse-Corpus.strata-insurance-corpus
Strata Insurance Corpus
A reproducible, fully synthetic, multi-format insurance document corpus for a fictional pan-European
property-&-casualty insurer, Meridian Mutual, shipped with a golden evaluation set produced by
construction. Built to exercise and benchmark document-RAG systems on enterprise-shaped data —
born-digital and scanned PDFs, Word documents, spreadsheets, and photos — with trustworthy ground truth.
Everything here is synthetic. No real persons, companies, or… See the full description on the dataset page: https://huggingface.co/datasets/NikolaiSachok/strata-insurance-corpus.courtlistener-legal-corpus
CourtListener Legal Corpus (CPT + SFT)
Training corpus used to fine-tune the Legal-Qwen model family
(27B,
9B).
All content is derived from public-domain United States court opinions via
CourtListener (Free Law Project).
Files
File
Records
Purpose
cpt.jsonl
21,421
Continued pre-training documents: full opinion texts, quality-filtered
sft_cap20k.jsonl
20,000
Instruction pairs: opinion excerpt -> one-sentence holding (legal parenthetical), built from… See the full description on the dataset page: https://huggingface.co/datasets/scottyjmp5/courtlistener-legal-corpus.bulgarian-corpus-33b
Bulgarian Corpus 33B (BC-33B)
Dataset Summary
The Bulgarian Corpus 33B (BC-33B) is a massive-scale, deduplicated, and cleaned dataset designed for training Foundation Models in Bulgarian. Comprising approximately 33.4 Billion tokens (measured with Qwen 2.5/Llama-3 tokenizer), it represents one of the largest open-source resources for Bulgarian LLM pretraining.
The dataset is engineered for a modern two-stage training pipeline:
Pretrain Subset (~29.3B Tokens): A… See the full description on the dataset page: https://huggingface.co/datasets/hasankursun/bulgarian-corpus-33b.spai-ss6-llm-1b-thai-corpus
Thai Medical And Health Corpus
Thai public medical and health web corpus collected for research and LLM dataset
experimentation, with optional imported Thai medical/health datasets from
Hugging Face stored as separate configs.
Public Web Corpus
Config: default
Split: train
Records: 3660 deduplicated articles
Columns: 16
Format: Parquet
Latest collection profile: free_1000
Latest generated at: 2026-06-06T17:41:38.787978+00:00
Source And Method
The… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-llm-1b-thai-corpus.100k-corpus-2026
MAST 100K Corpus 2026
This dataset contains the fixed English document corpus used for MAST @ FIRE 2026, the Multilingual Agentic Search Track. MAST evaluates whether multilingual agentic search systems can answer complex questions posed in different languages by retrieving English evidence and producing short, correct English answers.
This corpus is copied from BrowseComp-Plus, a benchmark for Deep-Research systems that isolates the effect of the retriever and the LLM agent to… See the full description on the dataset page: https://huggingface.co/datasets/mast-benchmark/100k-corpus-2026.recursive-cognition-corpus
LuisCore Recursive Cognition Corpus
LuisCore is a low-latency decentralized runtime substrate for multi-step inference at scale.
Generated: 2026-09-24T05:52:01.003Z
Rows: 13236
Owner: Luis610348
Canonical site: https://luiscore.com
What this dataset is
LuisCore is a recursive cognition infrastructure. This dataset is the public
LLM Discovery Corpus — a stable, deterministic Q&A set used by LuisCore to
help language models accurately describe, cite, and verify… See the full description on the dataset page: https://huggingface.co/datasets/Luis610348/recursive-cognition-corpus.LaSerena-Corpus-Geociencias
La Serena Digital Geo Corpus — Dominga EIA Dataset
Dataset Sci-Align de geología ambiental chilena basado en el expediente
de Evaluación de Impacto Ambiental del proyecto Dominga (SEIA, Región de Coquimbo).
Contenido
dominga_geo_align.jsonl — 402 registros Sci-Align (MinerU 3.1.15)
seia/ — documentos públicos del expediente Dominga (fuente primaria)
Licencia
CC-BY-4.0 — Fuente: SEIA Chile (acceso público)
Concurso
AGI4S — Pista 1: Creación de bases… See the full description on the dataset page: https://huggingface.co/datasets/Karlangaz/LaSerena-Corpus-Geociencias.taxpulse-corpus
TaxPulse: Pakistani Tax Law Corpus
Law is current to the Finance Act, 2026 (effective 1 July 2026, Tax Year 2027).
A document corpus of Pakistani tax law, assembled for the TaxPulse final-year
project (an AI tax consultation and FBR filing assistant). Sources are public
publications of the Federal Board of Revenue (FBR) and provincial revenue
authorities, plus reported case law.
Contents
Directory
What it holds
01_primary_law/
Income Tax Ordinance 2001… See the full description on the dataset page: https://huggingface.co/datasets/AbdulRahmanAzam/taxpulse-corpus.dutch-corpus-200b
Dutch Corpus 200B (DC-200B)
Dataset Summary
The Dutch Corpus 200B (DC-200B) is the largest open-source, deduplicated, and professionally cleaned dataset designed for training Foundation Models in the Dutch language. Comprising approximately 202 Billion tokens (measured with Qwen 2.5 tokenizer), it bridges the gap between high-resource English models and the Dutch ecosystem.
The dataset is engineered for a two-stage training pipeline:
Pretrain Subset (~195B… See the full description on the dataset page: https://huggingface.co/datasets/hasankursun/dutch-corpus-200b.nyayashastra-court-judgments-corpus
NyayaShastra Indian Court Judgments Corpus (12.4M Judgments)
A comprehensive, curated dataset of 12.4 Million Indian Supreme Court, High Court, and District Court judgments.
Optimized for high-speed columnar retrieval via Apache Parquet and DuckDB.
Total Partitions: 248
Format: Apache Parquet (Snappy compressed)
Columns: case_id, cnr, case_title, court_name_normalized, court_level, decision_date, cleaned_text, text_length
greek-corpus-150b
Greek Corpus 150B
A large-scale, deduplicated Greek (Modern Greek, el) text corpus for training and fine-tuning foundation models. It pairs a broad web/knowledge/formal-document pretrain layer with a multilingual-instruction SFT layer, all normalized to a single unified schema and globally deduplicated.
This is part of an ongoing Global Corpus family of per-language foundation-model datasets (Dutch, Turkish, Bulgarian, Greek, …) built on a consistent architecture so that sources… See the full description on the dataset page: https://huggingface.co/datasets/hasankursun/greek-corpus-150b.judaic-texts-corpus
Judaic Texts Corpus
Dataset Summary
Judaic Texts Corpus is a machine-readable Hebrew and Aramaic corpus of Judaic
texts derived from the Otzaria library release archives. It is intended for
language-model training, retrieval, search, digital humanities research, and
other NLP workflows that need structured access to rabbinic and traditional
Jewish texts.
The current dataset build is produced from the official
Otzaria/otzaria-library release
assets, which package… See the full description on the dataset page: https://huggingface.co/datasets/NHLOCAL/judaic-texts-corpus.russian-it-community-corpus
📦 Russian IT Community Corpus (RICC)
Russian IT Community Corpus (RICC) is an open, de-identified conversational dataset collected from 11 engineering community nodes spanning a 9-year timeline (2017–2026). It captures authentic discussions on backend systems, cloud infrastructure, AI/ML deployment, database internals, and software architecture.
The corpus is structured into ready-to-use splits for Instruction Fine-Tuning (SFT), Direct Preference Optimization (DPO)… See the full description on the dataset page: https://huggingface.co/datasets/wwewtech/russian-it-community-corpus.powertron-global-permafrost-corpus
Dataset Card: Powertron Global PermaFrost Corpus
Important Disambiguation: This corpus documents PermaFrost® NMR, a trademarked HVAC efficiency treatment product. It contains HVAC/refrigeration efficiency data (chillers, RTUs, DX systems, refrigeration). This corpus has NO connection to geological permafrost (frozen ground), climate science, or Arctic research. The name "PermaFrost" is a product trademark reflecting thermal transfer properties, not a geological term.… See the full description on the dataset page: https://huggingface.co/datasets/powertronglobal/powertron-global-permafrost-corpus.PolyMath
Dataset Card for PolyMath
Dataset Summary
PolyMath is a curated dataset of 11,090 high-difficulty mathematical problems designed for training reasoning models. Built for the AIMO Math Corpus Prize. Existing math datasets (NuminaMath-1.5, OpenMathReasoning) suffer from high noise rates in their hardest samples and largely unusable proof-based problems.
PolyMath addresses both issues through:
Data scraping: problems sourced from official competition PDFs absent from… See the full description on the dataset page: https://huggingface.co/datasets/AIMO-Corpus/PolyMath.egypt-legal-corpus
Egyptian Legal Corpus
A comprehensive collection of Egyptian legal texts, meticulously extracted and tokenized for Natural Language Processing (NLP) applications, legal research, and AI model training. This corpus provides high-quality Arabic legal content with structured metadata for efficient processing.
Dataset Statistics
This release provides a foundational legal corpus with strict quality controls:
Token Count: 25M+ tokens (25,054,372 tokens) using cl100k_base… See the full description on the dataset page: https://huggingface.co/datasets/dataflare/egypt-legal-corpus.control-sci-corpus
ControlSci Corpus
Control science structured corpus with two configs: Sci-Align benchmark (500 questions) and Sciverse SFT instruction pairs (924 ChatML entries).
License: CC-BY-4.0
Project: MorningStar0709/ControlMind
Configs
benchmark — Sci-Align Benchmark (500 questions)
4-dimension control science evaluation benchmark generated from the ControlSci structured corpus.
Split: core (500 questions)
Load:
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/MorningStar0709/control-sci-corpus.hse-qa-corpus
Canonical landing page: https://www.smartqhse.com/datasets/hse-qa-corpus
SmartQHSE HSE Q&A Corpus
24 long-form HSE / occupational-safety question-answer pairs across 15 categories — incident rates, ISO 45001, permits, risk assessment, OSHA (US), HSE (UK), GCC regulations, PPE, heat stress, exposure, ergonomics, incident investigation, training, HSE software. Each answer is multi-paragraph with cited sources, formulas, and OSHA/regulatory references. Suitable for instruction… See the full description on the dataset page: https://huggingface.co/datasets/SmartQHSE/hse-qa-corpus.jma-gsi-disaster-action-corpus
JMA-GSI Disaster Action Corpus
A grounded, multilingual disaster-response dataset built from official Japanese government open data (JMA alert XML + JMA multilingual glossary + JMA forecast-area GIS + GSI designated evacuation shelters). Structured hazard alerts are transformed into easy-Japanese and multilingual (ja / easy-ja / en / vi / id / ne / my) action guidance, linked to hazard-compatible evacuation shelters, with full source traceability.
License (derived dataset): CC BY… See the full description on the dataset page: https://huggingface.co/datasets/edomaru/jma-gsi-disaster-action-corpus.tiny-slm-pretraining-corpus
🚀 Ultra High-Quality Tiny SLM Pre-Training Corpus (<100GB)
A state-of-the-art, balanced 7-domain pre-training dataset engineered specifically for Small Language Models (Tiny SLMs: 50M – 2B parameters) such as SmolLM2, SmolLM3, MobileLLM, Llama 3.2 1B, and custom architectures.
100% compatible with Unsloth Studio, Unsloth AI, Hugging Face datasets, and PyTorch DataLoaders.
📊 Dataset Statistics
Total Documents: 20,066,075
Train: 19,663,898
Validation: 402,177… See the full description on the dataset page: https://huggingface.co/datasets/JustACluelessKidAtSchool/tiny-slm-pretraining-corpus.verifiable-corpus
verifiable-corpus
This is the corpus from "Learning on the Job: Test-Time Curricula for Targeted Reinforcement Learning".
Code: https://github.com/jonhue/ttc
Introduction
We study how large language models (LLMs) can continually improve at reasoning on their target tasks at test-time. We propose an agent that assembles a task-specific curriculum, called test-time curriculum (TTC-RL), and applies reinforcement learning to continue training the model for its target task.… See the full description on the dataset page: https://huggingface.co/datasets/lasgroup/verifiable-corpus.cbi-archive-corpus
Central Bank of Ireland Public Archive Corpus
A page-anchored, provenance-classified corpus of the Central Bank of Ireland's
public document archive. 5,568 documents and 89,242 page or pseudo-page rows.
PDF rows have true source-page anchors; most Office and archive rows do not.
This is an unofficial derived work. It is not published by, affiliated with, or
endorsed by the Central Bank of Ireland.
What makes this different from a pile of scraped PDFs
Two things.… See the full description on the dataset page: https://huggingface.co/datasets/aditya487/cbi-archive-corpus.turkish-law-corpus
⚖️ Turkish Law — 106 Kanun Korpusu & Soru-Cevap106 Statutes Corpus & QA
🇹🇷 Türk hukukunun en çok kullanılan 106 kanunu, madde madde temizlenmiş 16.001 metin parçası ve bu maddelere dayalı 5.011 Türkçe soru-cevap çifti. Tamamı resmî kaynaktan (mevzuat.gov.tr), RAG ve yapay zekâ uygulamaları için hazır.
🇬🇧 The 106 most widely used Turkish statutes as 16,001 clean, article-level text chunks, plus 5,011 Turkish question-answer pairs grounded in those articles. All from the… See the full description on the dataset page: https://huggingface.co/datasets/CtnkyaABC/turkish-law-corpus.
