CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AlgorithmicResearchGroup /s2orc_full S2ORC Full — Semantic Scholar Open Research Corpus A complete redistribution of the S2ORC dataset in Parquet format on Hugging Face, containing 14.5 million academic papers with full text, structured metadata, and citation information. Dataset Description S2ORC (Semantic Scholar Open Research Corpus) is a general-purpose corpus for NLP and text mining research over scientific papers, originally developed by the Allen Institute for AI. This version provides the full… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/s2orc_full.texttext-generation10M<n<100M2 likes17k downloads5mo agoHugging Face02AuthenticIlm /Shamela4_Full_DB Shamela 4 — Full Islamic Library Corpus A complete extraction of al-Maktaba al-Shamela (الشاملة) v4, containing 8,589 books across 40 categories of classical Islamic sciences. Extracted from the original Lucene + Sqlite Shamela DB on 2026-04-26 with ~7.6 million pages and ~19 GB of Arabic text. Dataset Structure stage0_raw/ ├── _meta/ # Cross-cutting metadata (Parquet + JSONL) │ ├── extraction_manifest.json # Global extraction record │ ├──… See the full description on the dataset page: https://huggingface.co/datasets/AuthenticIlm/Shamela4_Full_DB.text-generation10M<n<100M30 likes9.2k downloads4mo agoHugging Face03AmazonScience /migration-bench-java-full MigrationBench 1. 📖 Overview 🤗 MigrationBench is a large-scale code migration benchmark dataset at the repository level, across multiple programming languages. Current and initial release includes java 8 repositories with the maven build system… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/migration-bench-java-full.tabulartext-generation1K<n<10K4 likes4.8k downloads1y agoHugging Face04OpenDataArena /MMFineReason-Full-2.3M-Qwen3-VL-235B-Thinking MMFineReason-Full-2.3M The Complete Pre-Selection Dataset — Before Quality Filtering 📖 Overview MMFineReason-Full-2.3M is the complete pre-selection dataset containing 2.3M samples and 8.8B solution tokens, generated through our reasoning distillation pipeline before the data selection stage. This dataset includes all samples that passed basic template and length validation, but have not undergone correctness verification filtering. 🎯 Key Characteristics… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/MMFineReason-Full-2.3M-Qwen3-VL-235B-Thinking.imagevisual-question-answering1M<n<10M65 likes4.5k downloads8mo agoHugging Face05jedibear /s2orc_full S2ORC Full — Semantic Scholar Open Research Corpus A complete redistribution of the S2ORC dataset in Parquet format on Hugging Face, containing 14.5 million academic papers with full text, structured metadata, and citation information. Dataset Description S2ORC (Semantic Scholar Open Research Corpus) is a general-purpose corpus for NLP and text mining research over scientific papers, originally developed by the Allen Institute for AI. This version provides the full… See the full description on the dataset page: https://huggingface.co/datasets/jedibear/s2orc_full.texttext-generation10M<n<100M0 likes2.4k downloads5mo agoHugging Face06ericktwo /MMFineReason-Full-2.3M-Qwen3-VL-235B-Thinking MMFineReason-Full-2.3M The Complete Pre-Selection Dataset — Before Quality Filtering 📖 Overview MMFineReason-Full-2.3M is the complete pre-selection dataset containing 2.3M samples and 8.8B solution tokens, generated through our reasoning distillation pipeline before the data selection stage. This dataset includes all samples that passed basic template and length validation, but have not undergone correctness verification filtering. 🎯 Key Characteristics… See the full description on the dataset page: https://huggingface.co/datasets/ericktwo/MMFineReason-Full-2.3M-Qwen3-VL-235B-Thinking.imagevisual-question-answering1M<n<10M1 likes1.4k downloads8mo agoHugging Face07KiteFishAI /arxiv-tex-corpus-fullarxiv-tex-corpus-full (80GB) Large-scale LaTeX corpus from arXiv (math, CS, physics, statistics) 📄 Paper: https://arxiv.org/abs/2602.17288 📚 Overview arxiv-tex-corpus-full (80GB) is a large-scale dataset of LaTeX source content extracted from papers hosted on arXiv. This version contains approximately 80GB of structured JSONL data, restricted to the following arXiv categories: math cs hep-th hep-ph quant-ph stat.ML stat.TH The dataset is designed for research in: Large… See the full description on the dataset page: https://huggingface.co/datasets/KiteFishAI/arxiv-tex-corpus-full.texttext-generation100K<n<1M13 likes1.1k downloads7mo agoHugging Face08faur-ai /fulg ❄️FuLG The FuLG dataset is a comprehensive Romanian language corpus comprising 150 billion tokens, carefully extracted from Common Crawl. This extensive dataset is the result of rigorous filtering and deduplication processes applied to 95 Common Crawl snapshots. The compressed dataset has 289 GB. For more details, check the arXiv preprint. How do I download this? Using 🤗 Datasets from datasets import load_dataset # Full dataset dataset =… See the full description on the dataset page: https://huggingface.co/datasets/faur-ai/fulg.tabulartext-generation10M<n<100M10 likes1k downloads2y agoHugging Face09ulab-ai /AcademicEval_Fullsummarization10K<n<100K1 likes986 downloads2y agoHugging Face10yuntian-deng /WildChat-4.8M-Fullgated Dataset Card for WildChat-4.8M-Full Dataset Description Interactive Search Tool: https://wildvisualizer.com WildChat paper: https://arxiv.org/abs/2405.01470 WildVis paper: https://arxiv.org/abs/2409.03753 Point of Contact: Yuntian Deng Dataset Summary WildChat-4.8M-Full is a collection of 4,743,336 conversations (out of 4,804,190 originally, after removing all conversations flagged with "sexual/minors" by OpenAI Moderation) between human users and… See the full description on the dataset page: https://huggingface.co/datasets/yuntian-deng/WildChat-4.8M-Full.texttext-generation1M<n<10M11 likes862 downloads1y agoHugging Face11MoreThought /DeepSWEGym2-Full Dataset Description This dataset is a merge containing many high quality SWE datasets, it aims to improve benchmark results on DeepSWE-style problems, benchmarks, and general coding skills. It is NOT specifically filtered for rows with complex/long code problems in the original datasets, despite still having an average row size of 169.08kb, a total uncompressed size of 19.80GB, and a total of 122791 examples. Dataset Details Curated by: MoreThought Funded by:… See the full description on the dataset page: https://huggingface.co/datasets/MoreThought/DeepSWEGym2-Full.texttext-generation100K<n<1M1 likes831 downloads18d agoHugging Face12ryankamiri /R2E-Gym-Full R2E-Gym Subset Filtered for MAGRPO Filtered subset of R2E-Gym optimized for 2-agent MAGRPO training with 7B models. Dataset Statistics Total instances: 167 Format: Issue description + Oracle files in prompt Optimized for: 2-agent collaboration, 7B models Filtering Criteria (SWE-bench Lite Style) Problem statement: >40 words (up to 500 for context window) Must have non-empty oracle patch (non-test file changes) File count: Exactly 1 oracle file (single-file… See the full description on the dataset page: https://huggingface.co/datasets/ryankamiri/R2E-Gym-Full.tabulartext-generationn<1K0 likes801 downloads10mo agoHugging Face13bigcode /the-stack-v2-train-full-idsgated The Stack v2 The dataset consists of 4 versions: bigcode/the-stack-v2: the full "The Stack v2" dataset bigcode/the-stack-v2-dedup: based on the bigcode/the-stack-v2 but further near-deduplicated bigcode/the-stack-v2-train-full-ids: based on the bigcode/the-stack-v2-dedup dataset but further filtered with heuristics and spanning 600+ programming languages. The data is grouped into repositories. <-- you are here bigcode/the-stack-v2-train-smol-ids: based on the… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/the-stack-v2-train-full-ids.tabulartext-generation10M<n<100M65 likes800 downloads2mo agoHugging Face14MoreThought /DeepSWEGym-Full Dataset Description This dataset is a merged version of all the SWE-bench/SWE-smith-lang datasets (88k rows total), it aims to improve benchmark results on DeepSWE-style problems, benchmarks, and general coding skills. It is NOT specifically filtered for rows with complex/long code problems in the original datasets, despite still having an average row size of 85.4kb, a total uncompressed size of 7.53GB, and 88130 examples total. Dataset Details Curated by:… See the full description on the dataset page: https://huggingface.co/datasets/MoreThought/DeepSWEGym-Full.texttext-generation10K<n<100K1 likes692 downloads19d agoHugging Face15muhammedturan /turk-ictihat-kararlari-fulltext Türk İçtihat Kararları - Tam Metin (UYAP e-Bedesten) Adalet Bakanlığı UYAP Mevzuat ve İçtihat Programı (bedesten.adalet.gov.tr) üzerinden derlenen Yargıtay kararları veri seti. Kavram bazlı arama sorguları ile toplanmıştır. Boyut Kayıt: 9,899,589 benzersiz karar Terim dağılımı: {"bosanma": 100052, "karar": 9894542, "nafaka": 66241, "tazminat": 100052} Yıl aralığı: 1993-2026 Tam metni gelen karar: 39,301 (artımlı olarak dolduruluyor) Şema… See the full description on the dataset page: https://huggingface.co/datasets/muhammedturan/turk-ictihat-kararlari-fulltext.tabulartext-generation1M<n<10M1 likes592 downloads12d agoHugging Face16akahana /wikipedia-full Dataset Card for Wikimedia Wikipedia Dataset Summary Wikipedia dataset containing cleaned articles of all languages. The dataset is built from the Wikipedia dumps (https://dumps.wikimedia.org/) with one subset per language, each containing a single train split. Each example contains the content of one full Wikipedia article with cleaning to strip markdown and unwanted sections (references, etc.). All language subsets have already been processed for recent dump, and you… See the full description on the dataset page: https://huggingface.co/datasets/akahana/wikipedia-full.texttext-generation10M<n<100M2 likes509 downloads9mo agoHugging Face17hulk10 /conseil-administratives-appel-full-documents Décisions de Justice Administrative Françaises Description du dataset Ce dataset contient un corpus de décisions de justice administrative françaises extraites de la plateforme officielle Open Data de la Justice Administrative. Les documents sont publiés par le Conseil d'État dans le cadre de la politique d'ouverture des données publiques de la justice française. Les décisions sont diffusées au format XML et anonymisées conformément aux exigences légales relatives… See the full description on the dataset page: https://huggingface.co/datasets/hulk10/conseil-administratives-appel-full-documents.texttext-classification1K<n<10K1 likes495 downloads14h agoHugging Face18AdaMLLab /nyu-aco-ocr-full nyu-aco-ocr-full Scanned Arabic books from the Arabic Collections Online (ACO) archive, OCR'd page by page with the dots.ocr vision-language model served through vLLM. The archive spans seven partner collections: NYU, Princeton, Cornell, Columbia, AUB, AUC, and UAE National Archives. Each PDF page is rendered at 200 DPI, OCR'd individually, and the pages of a book are joined into one markdown document separated by \n\n---\n\n. Each row is one complete book. Schema… See the full description on the dataset page: https://huggingface.co/datasets/AdaMLLab/nyu-aco-ocr-full.texttext-generation10K<n<100K0 likes489 downloads1mo agoHugging Face19hulk10 /conseil-detat-full-documents Décisions du Conseil d'État (France) Description Ce dataset contient un corpus de décisions rendues par le Conseil d'État français, la plus haute juridiction de l'ordre administratif. Les décisions proviennent de la plateforme officielle Open Data de la Justice Administrative et sont diffusées au format XML anonymisé. Le corpus rassemble les textes intégraux des décisions ainsi que plusieurs métadonnées permettant leur identification et leur traçabilité. Source… See the full description on the dataset page: https://huggingface.co/datasets/hulk10/conseil-detat-full-documents.tabularquestion-answering1K<n<10K1 likes486 downloads14h agoHugging Face20MohamedRashad /Arabic-VLM-Full-Pearl 💎 The Arabic VLM Dataset (Full Pearl Edition) This repository contains the full, unreviewed dataset comprising 309K multimodal examples. This data was generated automatically using the agentic pipeline developed for the Pearl project, as described in our paper. Disclaimer: This is the raw, synthetic data that has not been subject to human review. It was generated as part of the data creation process and is released for research purposes. It may contain noise, errors, or… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/Arabic-VLM-Full-Pearl.imagequestion-answering100K<n<1M10 likes452 downloads10mo agoHugging Face21amrachraf /arXiv-full-text-chunked Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More Information Needed] Paper [optional]:… See the full description on the dataset page: https://huggingface.co/datasets/amrachraf/arXiv-full-text-chunked.texttext-generation100K<n<1M1 likes443 downloads2y agoHugging Face22festr2 /kimi-k3-full-mxfp4-kld-reference-32x2048 Kimi K3 full-MXFP4 KLD reference logits This dataset contains the canonical full-vocabulary reference logits for quantization comparisons of Kimi K3. The source is the original full MXFP4 checkpoint served as W4A16 on TP16 with vLLM dev/gg-k3, SparkInfer, and InstantTensor. Contents 32 independent 2048-token windows 65,504 scored next-token positions (32 * 2047) vocabulary size 163,840 one [2047, 163840] F32 safetensors tensor per window tensor key: logits total… See the full description on the dataset page: https://huggingface.co/datasets/festr2/kimi-k3-full-mxfp4-kld-reference-32x2048.text-generation10K<n<100K0 likes407 downloads2mo agoHugging Face23laion /llama-nemotron-science-reasoning-on-canonical-think-full Llama-Nemotron science reasoning — Delphi canonical-think (COMPLETE, no length filter) The complete reasoning:on science split of nvidia/Llama-Nemotron-Post-Training-Dataset, converted once into the canonical Delphi chat-template thinking format. 708,920 rows. Unlike the cold-start warmup slice open-athena/llama-nemotron-science-reasoning-on-le3000tok-100k (and its -canonical-think variant), this build applies no length cap and no subsample — every long-CoT science example is… See the full description on the dataset page: https://huggingface.co/datasets/laion/llama-nemotron-science-reasoning-on-canonical-think-full.texttext-generation100K<n<1M0 likes381 downloads21d agoHugging Face24IlyaGusev /librusec_full Ficbook dataset Description A dump of the library, mainly in Russian, including all metadata. Usage of this dataset is possible only for scientific purposes on a non-commercial basis. Script: parse_zip_fb2.py Source: booktracker Point of Contact: Ilya Gusev Languages: Mostly Russian Usage Prerequisites: pip install datasets zstandard jsonlines pysimdjson Dataset iteration: from datasets import load_dataset for example in… See the full description on the dataset page: https://huggingface.co/datasets/IlyaGusev/librusec_full.text-generation100K<n<1M2 likes329 downloads3y agoHugging Face25aslawliet /flan2021-full Task Name FLAN-2021 -> 70 { "ag_news_subset": 108497, "ai2_arc/ARC-Challenge": 829, "ai2_arc/ARC-Easy": 1927, "aeslc": 13187, "anli/r1": 15361, "anli/r2": 41133, "anli/r3": 91048, "bool_q": 8343, "cnn_dailymail": 259607, "coqa": 6456, "cosmos_qa": 22996, "definite_pronoun_resolution": 1079, "drop": 70045, "fix_punct": 25690, "gem/common_gen": 60936, "gem/dart": 56724, "gem/e2e_nlg": 30337, "gem/web_nlg_en": 31899… See the full description on the dataset page: https://huggingface.co/datasets/aslawliet/flan2021-full.texttext-generation10M<n<100M2 likes328 downloads2y agoHugging Face26EPFLiGHT /fully-open-meditron Fully Open Meditron Corpus 👋 Join our LiGHT community. 📖 Check out the MeditronFO blog and MeditronFO preprint. 🔜 If you are a clinician join the MOOVE initiative here. [Hugging Face] [Preprint] [GitHub] [Dataset] License: Apache 2.0 | Authors: LiGHT [!Note] A clinician-vetted training corpus for medical large language models, accompanying the paper Fully Open Meditron: An Auditable Pipeline for Clinical LLMs. The… See the full description on the dataset page: https://huggingface.co/datasets/EPFLiGHT/fully-open-meditron.textquestion-answering100K<n<1M8 likes327 downloads3mo agoHugging Face27asingh15 /tcs-qwen36-27b-direction-rollouts-pilot-50-full-trace TCS Qwen3.6-27B Direction Beam Full-Trace Pilot Verified compact export for tcs_qwen36_27b_direction_beam_pilot50_full_logging_20260814. Source dataset: TCS train-00000-of-00001.parquet Problems: 50 Displayed trajectories: 200 Chunk probes: 1,600 Terminal answers and rubric grades: 6,400 Policy: Qwen/Qwen3.6-27B Judge: openai/gpt-oss-20b (low reasoning) At each displayed chunk, the policy proposes four directions plus a null continuation. A width-four stochastic beam reaches… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/tcs-qwen36-27b-direction-rollouts-pilot-50-full-trace.text-generation0 likes310 downloads1mo agoHugging Face28VmaxRL /SWEUniverse-Repaired-Indist-full-not-SWE-bench-pro-matched VmaxRL/SWEUniverse-Repaired-Indist-full-not-SWE-bench-pro-matched This dataset contains a 350-row subset selected from the Indist SWEUniverse training rows. Selection policy: three-way repo overlap with Bugpilot and LM-Modify, deduped by repo plus introduction patch, then balanced round-robin across overlapping repos. Rows: 350 Selected repos: 19 Deduped overlap capacity: 468 Source dataset: VmaxRL/SWEUniverse-Repaired-Indist-full-not-SWE-bench-pro-matched texttext-generationn<1K0 likes266 downloads4mo agoHugging Face291337xyz1337xyz /codeforces-editorial-full-2026-07-16 Codeforces Editorial Full - 2026-07-16 A deterministic Plan-CRL-compatible materialization of open-r1/codeforces, pinned to revision fbe3f6e903ee854eec2e69e9d96d0306cde59baf. Size train: 9556 problems test: 468 problems total: 10024 unique problems safe fixed-output Plan-CRL evaluation rows: 1479 safe fixed-output rows with an editorial: 646 The upstream snapshot contains 10,024 unique problems, not 11,000+. The larger counts sometimes quoted for this corpus… See the full description on the dataset page: https://huggingface.co/datasets/1337xyz1337xyz/codeforces-editorial-full-2026-07-16.texttext-generation10K<n<100K0 likes201 downloads2mo agoHugging Face30mdonigian /full-structured-instruction-sft-dataset Full Structured + Instruction SFT Corpus Unified SFT training corpus built from Glaive, Hermes, UltraChat, and synthetic structured-output data. Dataset repo mdonigian/full-structured-instruction-sft-datasetRelease date: 2026-03-11 Included files train_full_sft.jsonl: full merged and shuffled SFT dataset source_glaive.jsonl: processed Glaive subset source_hermes.jsonl: processed Hermes subset source_ultrachat.jsonl: processed UltraChat subset… See the full description on the dataset page: https://huggingface.co/datasets/mdonigian/full-structured-instruction-sft-dataset.texttext-generation10K<n<100K0 likes170 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.