CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nvidia /Nemotron-Terminal-Corpus Terminal-Corpus: Large-Scale SFT Dataset for Terminal Agents Terminal-Corpus is a large-scale Supervised Fine-Tuning (SFT) dataset designed to scale the terminal interaction capabilities of Large Language Models (LLMs). Developed by NVIDIA, this dataset was built using the Terminal-Task-Gen pipeline, which combines dataset adaptation with synthetic task generation across diverse domains. 🚀 Key Results & Performance The high-quality trajectories in Terminal-Corpus enable… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Terminal-Corpus.textquestion-answering100K<n<1M144 likes92k downloads7mo agoHugging Face02Tevatron /browsecomp-plus-corpus BrowseComp-Plus Project Page | Paper | Code BrowseComp-Plus is a new benchmark for Deep-Research system, isolating the effect of the retriever and the LLM agent to enable fair, transparent comparisons of Deep-Research agents. The benchmark sources challenging, reasoning-intensive queries from OpenAI's BrowseComp. However, instead of searching the live web, BrowseComp-Plus evaluates against a fixed, curated corpus of ~100K web documents from the web. The corpus includes both… See the full description on the dataset page: https://huggingface.co/datasets/Tevatron/browsecomp-plus-corpus.textquestion-answering100K<n<1M18 likes31k downloads1y agoHugging Face03Azzindani /Legal_Corpus_QA_SynDeepThink 🧠 Legal Corpus QA SynDeepThink Dataset This repository contains a high-intelligence Legal Question-and-Answer dataset, generated through an advanced Iterative and Recursive Thinking process. It bridges the gap between static legal corpora and the dynamic "check-and-recheck" nature of human legal expertise. 🏛️ 💡 The Concept: Iterative & Recursive Legal Logic While standard synthetic datasets are often generated in a single pass, Legal_Corpus_QA_SynDeepThink mimics the… See the full description on the dataset page: https://huggingface.co/datasets/Azzindani/Legal_Corpus_QA_SynDeepThink.tabulartext-generation1K<n<10K1 likes11k downloads7mo agoHugging Face04freshstack /corpus-oct-2024 Dataset Card for FreshStack (Corpus) Homepage | Repository | Paper FreshStack is a holistic framework to construct challenging IR/RAG evaluation datasets that focuses on search across niche and recent topics. This dataset (October 2024) contains the query, nuggets, answers and nugget-level relevance judgments of 5 niche topics focused on software engineering and machine learning. The queries and answers (accepted) are taken from Stack Overflow, GPT-4o generates the nuggets and… See the full description on the dataset page: https://huggingface.co/datasets/freshstack/corpus-oct-2024.textquestion-answering100K<n<1M5 likes5.3k downloads1y agoHugging Face05dignity045 /Collective-Corpus 🧠 Collective Corpus — Universal Pretraining + Finetuning Dataset (500B+ Tokens) Collective-Corpus is a massive-scale, multi-domain dataset designed to train Transformer-based language models from scratch and finetune them across a wide variety of domains — all in one place. 📚 Dataset Scope This dataset aims to cover the full LLM lifecycle, from raw pretraining to domain-specialized finetuning. 1. Pretraining Corpus Large-scale, diverse multilingual text… See the full description on the dataset page: https://huggingface.co/datasets/dignity045/Collective-Corpus.texttext-generation100M<n<1B2 likes959 downloads1y agoHugging Face06hasankursun /turkish-corpus-100b Turkish Corpus 100B (TC-100B) Dataset Summary The Turkish Corpus 100B (TC-100B) is a massive-scale, deduplicated, and cleaned dataset designed for training Foundation Models in Turkish. Comprising approximately 105 Billion tokens (measured with Qwen/Llama3 tokenizer), it represents one of the largest open resources for Turkish LLM pretraining. The dataset is engineered for a two-stage training pipeline: Pretrain Subset (~103B Tokens): A diverse mix of… See the full description on the dataset page: https://huggingface.co/datasets/hasankursun/turkish-corpus-100b.texttext-generation100M<n<1B8 likes857 downloads3mo agoHugging Face07simplex-ai-inc /LiteResearcher-Browse-Corpus LiteResearcher Browse Corpus Full web-page text for reproducing the visit / /web_parser tool of LiteResearcher. What this is The LiteResearcher local environment exposes two tools, backed by two separate datasets: Tool Endpoint Dataset Content search /search LiteResearcher-Corpus url / title / doc — doc is snippet-level (~200 chars), used to build the BGE-M3 vector index visit /web_parser this dataset url / title / text — text is the full page… See the full description on the dataset page: https://huggingface.co/datasets/simplex-ai-inc/LiteResearcher-Browse-Corpus.text-retrieval10M<n<100M2 likes843 downloads2mo agoHugging Face08NikolaiSachok /strata-insurance-corpus Strata Insurance Corpus A reproducible, fully synthetic, multi-format insurance document corpus for a fictional pan-European property-&-casualty insurer, Meridian Mutual, shipped with a golden evaluation set produced by construction. Built to exercise and benchmark document-RAG systems on enterprise-shaped data — born-digital and scanned PDFs, Word documents, spreadsheets, and photos — with trustworthy ground truth. Everything here is synthetic. No real persons, companies, or… See the full description on the dataset page: https://huggingface.co/datasets/NikolaiSachok/strata-insurance-corpus.documentquestion-answeringn<1K2 likes830 downloads3mo agoHugging Face09scottyjmp5 /courtlistener-legal-corpus CourtListener Legal Corpus (CPT + SFT) Training corpus used to fine-tune the Legal-Qwen model family (27B, 9B). All content is derived from public-domain United States court opinions via CourtListener (Free Law Project). Files File Records Purpose cpt.jsonl 21,421 Continued pre-training documents: full opinion texts, quality-filtered sft_cap20k.jsonl 20,000 Instruction pairs: opinion excerpt -> one-sentence holding (legal parenthetical), built from… See the full description on the dataset page: https://huggingface.co/datasets/scottyjmp5/courtlistener-legal-corpus.text-generation10K<n<100K0 likes800 downloads2mo agoHugging Face10hasankursun /bulgarian-corpus-33b Bulgarian Corpus 33B (BC-33B) Dataset Summary The Bulgarian Corpus 33B (BC-33B) is a massive-scale, deduplicated, and cleaned dataset designed for training Foundation Models in Bulgarian. Comprising approximately 33.4 Billion tokens (measured with Qwen 2.5/Llama-3 tokenizer), it represents one of the largest open-source resources for Bulgarian LLM pretraining. The dataset is engineered for a modern two-stage training pipeline: Pretrain Subset (~29.3B Tokens): A… See the full description on the dataset page: https://huggingface.co/datasets/hasankursun/bulgarian-corpus-33b.texttext-generation10M<n<100M6 likes693 downloads3mo agoHugging Face11SPAISS6F1 /spai-ss6-llm-1b-thai-corpus Thai Medical And Health Corpus Thai public medical and health web corpus collected for research and LLM dataset experimentation, with optional imported Thai medical/health datasets from Hugging Face stored as separate configs. Public Web Corpus Config: default Split: train Records: 3660 deduplicated articles Columns: 16 Format: Parquet Latest collection profile: free_1000 Latest generated at: 2026-06-06T17:41:38.787978+00:00 Source And Method The… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-llm-1b-thai-corpus.tabulartext-generation10M<n<100M0 likes515 downloads4mo agoHugging Face12mast-benchmark /100k-corpus-2026 MAST 100K Corpus 2026 This dataset contains the fixed English document corpus used for MAST @ FIRE 2026, the Multilingual Agentic Search Track. MAST evaluates whether multilingual agentic search systems can answer complex questions posed in different languages by retrieving English evidence and producing short, correct English answers. This corpus is copied from BrowseComp-Plus, a benchmark for Deep-Research systems that isolates the effect of the retriever and the LLM agent to… See the full description on the dataset page: https://huggingface.co/datasets/mast-benchmark/100k-corpus-2026.textquestion-answering100K<n<1M0 likes498 downloads2mo agoHugging Face13Luis610348 /recursive-cognition-corpus LuisCore Recursive Cognition Corpus LuisCore is a low-latency decentralized runtime substrate for multi-step inference at scale. Generated: 2026-09-24T11:09:13.307Z Rows: 13236 Owner: Luis610348 Canonical site: https://luiscore.com What this dataset is LuisCore is a recursive cognition infrastructure. This dataset is the public LLM Discovery Corpus — a stable, deterministic Q&A set used by LuisCore to help language models accurately describe, cite, and verify… See the full description on the dataset page: https://huggingface.co/datasets/Luis610348/recursive-cognition-corpus.textquestion-answering10K<n<100K1 likes497 downloads22m agoHugging Face14Karlangaz /LaSerena-Corpus-Geociencias La Serena Digital Geo Corpus — Dominga EIA Dataset Dataset Sci-Align de geología ambiental chilena basado en el expediente de Evaluación de Impacto Ambiental del proyecto Dominga (SEIA, Región de Coquimbo). Contenido dominga_geo_align.jsonl — 402 registros Sci-Align (MinerU 3.1.15) seia/ — documentos públicos del expediente Dominga (fuente primaria) Licencia CC-BY-4.0 — Fuente: SEIA Chile (acceso público) Concurso AGI4S — Pista 1: Creación de bases… See the full description on the dataset page: https://huggingface.co/datasets/Karlangaz/LaSerena-Corpus-Geociencias.text-generationn<1K0 likes403 downloads4mo agoHugging Face15AbdulRahmanAzam /taxpulse-corpus TaxPulse: Pakistani Tax Law Corpus Law is current to the Finance Act, 2026 (effective 1 July 2026, Tax Year 2027). A document corpus of Pakistani tax law, assembled for the TaxPulse final-year project (an AI tax consultation and FBR filing assistant). Sources are public publications of the Federal Board of Revenue (FBR) and provincial revenue authorities, plus reported case law. Contents Directory What it holds 01_primary_law/ Income Tax Ordinance 2001… See the full description on the dataset page: https://huggingface.co/datasets/AbdulRahmanAzam/taxpulse-corpus.question-answering0 likes399 downloads7d agoHugging Face16hasankursun /dutch-corpus-200b Dutch Corpus 200B (DC-200B) Dataset Summary The Dutch Corpus 200B (DC-200B) is the largest open-source, deduplicated, and professionally cleaned dataset designed for training Foundation Models in the Dutch language. Comprising approximately 202 Billion tokens (measured with Qwen 2.5 tokenizer), it bridges the gap between high-resource English models and the Dutch ecosystem. The dataset is engineered for a two-stage training pipeline: Pretrain Subset (~195B… See the full description on the dataset page: https://huggingface.co/datasets/hasankursun/dutch-corpus-200b.texttext-generation100M<n<1B4 likes360 downloads3mo agoHugging Face17ursishant /nyayashastra-court-judgments-corpus NyayaShastra Indian Court Judgments Corpus (12.4M Judgments) A comprehensive, curated dataset of 12.4 Million Indian Supreme Court, High Court, and District Court judgments. Optimized for high-speed columnar retrieval via Apache Parquet and DuckDB. Total Partitions: 248 Format: Apache Parquet (Snappy compressed) Columns: case_id, cnr, case_title, court_name_normalized, court_level, decision_date, cleaned_text, text_length texttext-retrieval10M<n<100M0 likes354 downloads1mo agoHugging Face18hasankursun /greek-corpus-150b Greek Corpus 150B A large-scale, deduplicated Greek (Modern Greek, el) text corpus for training and fine-tuning foundation models. It pairs a broad web/knowledge/formal-document pretrain layer with a multilingual-instruction SFT layer, all normalized to a single unified schema and globally deduplicated. This is part of an ongoing Global Corpus family of per-language foundation-model datasets (Dutch, Turkish, Bulgarian, Greek, …) built on a consistent architecture so that sources… See the full description on the dataset page: https://huggingface.co/datasets/hasankursun/greek-corpus-150b.texttext-generation10M<n<100M2 likes351 downloads3mo agoHugging Face19NHLOCAL /judaic-texts-corpus Judaic Texts Corpus Dataset Summary Judaic Texts Corpus is a machine-readable Hebrew and Aramaic corpus of Judaic texts derived from the Otzaria library release archives. It is intended for language-model training, retrieval, search, digital humanities research, and other NLP workflows that need structured access to rabbinic and traditional Jewish texts. The current dataset build is produced from the official Otzaria/otzaria-library release assets, which package… See the full description on the dataset page: https://huggingface.co/datasets/NHLOCAL/judaic-texts-corpus.texttext-generation1K<n<10K1 likes344 downloads3d agoHugging Face20wwewtech /russian-it-community-corpus 📦 Russian IT Community Corpus (RICC) Russian IT Community Corpus (RICC) is an open, de-identified conversational dataset collected from 11 engineering community nodes spanning a 9-year timeline (2017–2026). It captures authentic discussions on backend systems, cloud infrastructure, AI/ML deployment, database internals, and software architecture. The corpus is structured into ready-to-use splits for Instruction Fine-Tuning (SFT), Direct Preference Optimization (DPO)… See the full description on the dataset page: https://huggingface.co/datasets/wwewtech/russian-it-community-corpus.tabulartext-generation1M<n<10M1 likes331 downloads16d agoHugging Face21powertronglobal /powertron-global-permafrost-corpus Dataset Card: Powertron Global PermaFrost Corpus Important Disambiguation: This corpus documents PermaFrost® NMR, a trademarked HVAC efficiency treatment product. It contains HVAC/refrigeration efficiency data (chillers, RTUs, DX systems, refrigeration). This corpus has NO connection to geological permafrost (frozen ground), climate science, or Arctic research. The name "PermaFrost" is a product trademark reflecting thermal transfer properties, not a geological term.… See the full description on the dataset page: https://huggingface.co/datasets/powertronglobal/powertron-global-permafrost-corpus.tabulartext-generation1K<n<10K0 likes300 downloads2mo agoHugging Face22AIMO-Corpus /PolyMath Dataset Card for PolyMath Dataset Summary PolyMath is a curated dataset of 11,090 high-difficulty mathematical problems designed for training reasoning models. Built for the AIMO Math Corpus Prize. Existing math datasets (NuminaMath-1.5, OpenMathReasoning) suffer from high noise rates in their hardest samples and largely unusable proof-based problems. PolyMath addresses both issues through: Data scraping: problems sourced from official competition PDFs absent from… See the full description on the dataset page: https://huggingface.co/datasets/AIMO-Corpus/PolyMath.textquestion-answering10K<n<100K2 likes275 downloads8mo agoHugging Face23dataflare /egypt-legal-corpus Egyptian Legal Corpus A comprehensive collection of Egyptian legal texts, meticulously extracted and tokenized for Natural Language Processing (NLP) applications, legal research, and AI model training. This corpus provides high-quality Arabic legal content with structured metadata for efficient processing. Dataset Statistics This release provides a foundational legal corpus with strict quality controls: Token Count: 25M+ tokens (25,054,372 tokens) using cl100k_base… See the full description on the dataset page: https://huggingface.co/datasets/dataflare/egypt-legal-corpus.texttext-generation1K<n<10K4 likes230 downloads8mo agoHugging Face24MorningStar0709 /control-sci-corpus ControlSci Corpus Control science structured corpus with two configs: Sci-Align benchmark (500 questions) and Sciverse SFT instruction pairs (924 ChatML entries). License: CC-BY-4.0 Project: MorningStar0709/ControlMind Configs benchmark — Sci-Align Benchmark (500 questions) 4-dimension control science evaluation benchmark generated from the ControlSci structured corpus. Split: core (500 questions) Load: from datasets import load_dataset ds =… See the full description on the dataset page: https://huggingface.co/datasets/MorningStar0709/control-sci-corpus.imagequestion-answering1K<n<10K0 likes210 downloads2mo agoHugging Face25SmartQHSE /hse-qa-corpus Canonical landing page: https://www.smartqhse.com/datasets/hse-qa-corpus SmartQHSE HSE Q&A Corpus 24 long-form HSE / occupational-safety question-answer pairs across 15 categories — incident rates, ISO 45001, permits, risk assessment, OSHA (US), HSE (UK), GCC regulations, PPE, heat stress, exposure, ergonomics, incident investigation, training, HSE software. Each answer is multi-paragraph with cited sources, formulas, and OSHA/regulatory references. Suitable for instruction… See the full description on the dataset page: https://huggingface.co/datasets/SmartQHSE/hse-qa-corpus.question-answeringn<1K0 likes204 downloads4mo agoHugging Face26edomaru /jma-gsi-disaster-action-corpus JMA-GSI Disaster Action Corpus A grounded, multilingual disaster-response dataset built from official Japanese government open data (JMA alert XML + JMA multilingual glossary + JMA forecast-area GIS + GSI designated evacuation shelters). Structured hazard alerts are transformed into easy-Japanese and multilingual (ja / easy-ja / en / vi / id / ne / my) action guidance, linked to hazard-compatible evacuation shelters, with full source traceability. License (derived dataset): CC BY… See the full description on the dataset page: https://huggingface.co/datasets/edomaru/jma-gsi-disaster-action-corpus.tabularquestion-answering100K<n<1M1 likes198 downloads5mo agoHugging Face27JustACluelessKidAtSchool /tiny-slm-pretraining-corpus 🚀 Ultra High-Quality Tiny SLM Pre-Training Corpus (<100GB) A state-of-the-art, balanced 7-domain pre-training dataset engineered specifically for Small Language Models (Tiny SLMs: 50M – 2B parameters) such as SmolLM2, SmolLM3, MobileLLM, Llama 3.2 1B, and custom architectures. 100% compatible with Unsloth Studio, Unsloth AI, Hugging Face datasets, and PyTorch DataLoaders. 📊 Dataset Statistics Total Documents: 20,066,075 Train: 19,663,898 Validation: 402,177… See the full description on the dataset page: https://huggingface.co/datasets/JustACluelessKidAtSchool/tiny-slm-pretraining-corpus.tabulartext-generation10M<n<100M0 likes198 downloads1mo agoHugging Face28lasgroup /verifiable-corpus verifiable-corpus This is the corpus from "Learning on the Job: Test-Time Curricula for Targeted Reinforcement Learning". Code: https://github.com/jonhue/ttc Introduction We study how large language models (LLMs) can continually improve at reasoning on their target tasks at test-time. We propose an agent that assembles a task-specific curriculum, called test-time curriculum (TTC-RL), and applies reinforcement learning to continue training the model for its target task.… See the full description on the dataset page: https://huggingface.co/datasets/lasgroup/verifiable-corpus.texttext-generation10K<n<100K1 likes197 downloads1y agoHugging Face29aditya487 /cbi-archive-corpus Central Bank of Ireland Public Archive Corpus A page-anchored, provenance-classified corpus of the Central Bank of Ireland's public document archive. 5,568 documents and 89,242 page or pseudo-page rows. PDF rows have true source-page anchors; most Office and archive rows do not. This is an unofficial derived work. It is not published by, affiliated with, or endorsed by the Central Bank of Ireland. What makes this different from a pile of scraped PDFs Two things.… See the full description on the dataset page: https://huggingface.co/datasets/aditya487/cbi-archive-corpus.tabulartext-retrieval10K<n<100K0 likes196 downloads23d agoHugging Face30CtnkyaABC /turkish-law-corpus ⚖️ Turkish Law — 106 Kanun Korpusu & Soru-Cevap106 Statutes Corpus & QA 🇹🇷 Türk hukukunun en çok kullanılan 106 kanunu, madde madde temizlenmiş 16.001 metin parçası ve bu maddelere dayalı 5.011 Türkçe soru-cevap çifti. Tamamı resmî kaynaktan (mevzuat.gov.tr), RAG ve yapay zekâ uygulamaları için hazır. 🇬🇧 The 106 most widely used Turkish statutes as 16,001 clean, article-level text chunks, plus 5,011 Turkish question-answer pairs grounded in those articles. All from the… See the full description on the dataset page: https://huggingface.co/datasets/CtnkyaABC/turkish-law-corpus.textquestion-answering10K<n<100K3 likes189 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.