CoolFace
18 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AmazonScience /migration-bench-java-full MigrationBench 1. 📖 Overview 🤗 MigrationBench is a large-scale code migration benchmark dataset at the repository level, across multiple programming languages. Current and initial release includes java 8 repositories with the maven build system… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/migration-bench-java-full.tabulartext-generation1K<n<10K4 likes4.8k downloads1y agoHugging Face02faur-ai /fulg ❄️FuLG The FuLG dataset is a comprehensive Romanian language corpus comprising 150 billion tokens, carefully extracted from Common Crawl. This extensive dataset is the result of rigorous filtering and deduplication processes applied to 95 Common Crawl snapshots. The compressed dataset has 289 GB. For more details, check the arXiv preprint. How do I download this? Using 🤗 Datasets from datasets import load_dataset # Full dataset dataset =… See the full description on the dataset page: https://huggingface.co/datasets/faur-ai/fulg.tabulartext-generation10M<n<100M10 likes1k downloads2y agoHugging Face03ryankamiri /R2E-Gym-Full R2E-Gym Subset Filtered for MAGRPO Filtered subset of R2E-Gym optimized for 2-agent MAGRPO training with 7B models. Dataset Statistics Total instances: 167 Format: Issue description + Oracle files in prompt Optimized for: 2-agent collaboration, 7B models Filtering Criteria (SWE-bench Lite Style) Problem statement: >40 words (up to 500 for context window) Must have non-empty oracle patch (non-test file changes) File count: Exactly 1 oracle file (single-file… See the full description on the dataset page: https://huggingface.co/datasets/ryankamiri/R2E-Gym-Full.tabulartext-generationn<1K0 likes801 downloads10mo agoHugging Face04bigcode /the-stack-v2-train-full-idsgated The Stack v2 The dataset consists of 4 versions: bigcode/the-stack-v2: the full "The Stack v2" dataset bigcode/the-stack-v2-dedup: based on the bigcode/the-stack-v2 but further near-deduplicated bigcode/the-stack-v2-train-full-ids: based on the bigcode/the-stack-v2-dedup dataset but further filtered with heuristics and spanning 600+ programming languages. The data is grouped into repositories. <-- you are here bigcode/the-stack-v2-train-smol-ids: based on the… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/the-stack-v2-train-full-ids.tabulartext-generation10M<n<100M65 likes800 downloads2mo agoHugging Face05muhammedturan /turk-ictihat-kararlari-fulltext Türk İçtihat Kararları - Tam Metin (UYAP e-Bedesten) Adalet Bakanlığı UYAP Mevzuat ve İçtihat Programı (bedesten.adalet.gov.tr) üzerinden derlenen Yargıtay kararları veri seti. Kavram bazlı arama sorguları ile toplanmıştır. Boyut Kayıt: 9,899,589 benzersiz karar Terim dağılımı: {"bosanma": 100052, "karar": 9894542, "nafaka": 66241, "tazminat": 100052} Yıl aralığı: 1993-2026 Tam metni gelen karar: 39,301 (artımlı olarak dolduruluyor) Şema… See the full description on the dataset page: https://huggingface.co/datasets/muhammedturan/turk-ictihat-kararlari-fulltext.tabulartext-generation1M<n<10M1 likes592 downloads12d agoHugging Face06hulk10 /conseil-detat-full-documents Décisions du Conseil d'État (France) Description Ce dataset contient un corpus de décisions rendues par le Conseil d'État français, la plus haute juridiction de l'ordre administratif. Les décisions proviennent de la plateforme officielle Open Data de la Justice Administrative et sont diffusées au format XML anonymisé. Le corpus rassemble les textes intégraux des décisions ainsi que plusieurs métadonnées permettant leur identification et leur traçabilité. Source… See the full description on the dataset page: https://huggingface.co/datasets/hulk10/conseil-detat-full-documents.tabularquestion-answering1K<n<10K1 likes486 downloads21h agoHugging Face07bbidpa /flutter-full-examples-v1 Flutter Codegen: Full Examples Synthetic dataset of complete Flutter/Dart widgets, each paired with the goal that describes them and (optionally) starting code. Unlike flutter-codegen-diff-steps, there's no step history or diff structure here -- each row is a single, standalone goal -> complete file example. This is the whole-code counterpart to flutter-diff-steps-v1, intended for training/evaluating a baseline that generates the entire file in one shot, to compare against the… See the full description on the dataset page: https://huggingface.co/datasets/bbidpa/flutter-full-examples-v1.tabulartext-generation10K<n<100K0 likes136 downloads17d agoHugging Face08windchimeran /creativemath_fulltabularquestion-answering1K<n<10K0 likes111 downloads1y agoHugging Face09rescommons /Full-Ecom-Chatbot-Dataset E-commerce Chatbot Training Data A curated, multi-source dataset for training and evaluating e-commerce conversational AI systems. It covers a broad range of customer intents — from product discovery and order management to returns, tool-augmented responses, and RAG-grounded Q&A — across 16+ product domains. Dataset Summary Split Records Train 35,213 Test 8,818 Total 44,031 The train/test split uses prompt-group-level stratified sampling on source ×… See the full description on the dataset page: https://huggingface.co/datasets/rescommons/Full-Ecom-Chatbot-Dataset.tabularquestion-answering10K<n<100K0 likes72 downloads6mo agoHugging Face10Salesforce /lalm-judge-validation-full-duplex LALM Judge Validation on Full-Duplex Voice Agents Companion dataset for the paper A Reliability Assessment of LALM Audio Judges for Full-Duplex Voice Agents. This repository contains the anonymised ratings, adversarial-defect recall tables, JSON schemas, and analysis scripts used to produce every headline number, table, and figure in that paper. Summary 209 rated stereo sessions: 152 full-duplex agent-client conversations across 13 accent-and-condition strata… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/lalm-judge-validation-full-duplex.tabularaudio-classification1K<n<10K2 likes58 downloads2mo agoHugging Face11SkyWhal3 /stxbp1-pubmed-central-fulltext source_datasets: - PubMed Central STXBP1 PubMed Central Full-Text Dataset v2 A comprehensive collection of 31,786 full-text scientific articles from PubMed Central related to STXBP1, synaptic function, and neurological research. 🆕 Version 2 Updates (December 2025) Complete re-extraction with improved HTML parsing Full main text with proper section headers Enhanced metadata extraction 99.7% figure-image matching (see companion multimodal dataset)… See the full description on the dataset page: https://huggingface.co/datasets/SkyWhal3/stxbp1-pubmed-central-fulltext.tabulartext-generation10K<n<100K0 likes46 downloads10mo agoHugging Face12mikeogezi /negopt_fullThis is the dataset constructed in and used to fine-tune the models proposed in our paper Optimizing Negative Prompts for Enhanced Aesthetics and Fidelity in Text-To-Image Generation. The data is gotten from Playground. If you find this dataset useful, please cite us here: @article{ogezi2024optimizing, title={Optimizing Negative Prompts for Enhanced Aesthetics and Fidelity in Text-To-Image Generation}, author={Ogezi, Michael and Shi, Ning}, journal={arXiv preprint arXiv:2403.07605}… See the full description on the dataset page: https://huggingface.co/datasets/mikeogezi/negopt_full.imagetext-to-image100K<n<1M3 likes27 downloads1y agoHugging Face13Mikimi /ru-wikipedia-100k-full-text-daily-stats-10-years 📚 Russian Wikipedia Top 100K: Full Text with Daily Pageviews **Крупнейший открытый датасет русскоязычной Википедии с полными текстами статей и ежедневной статистикой просмотров за 10 лет. МГУ, ОТиПЛ, 2025** 📖 Описание Этот датасет содержит 99,348 самых популярных статей русскоязычной Википедии, отобранных по совокупному количеству просмотров за последние 10 лет. Для каждой статьи собраны полный текст с сохранением структуры, метаданные и детальная ежедневная… See the full description on the dataset page: https://huggingface.co/datasets/Mikimi/ru-wikipedia-100k-full-text-daily-stats-10-years.tabulartext-generation10K<n<100K1 likes27 downloads9mo agoHugging Face14nibzard /narodne-novine-full-text-markdown Narodne Novine Full Text Markdown HTML-to-Markdown text extraction snapshot derived from the NN archive. Coverage Extracted acts: 96855 Failed acts: 157 Missing HTML embodiments: 150 Files texts.parquet failures.parquet metadata.json Notes Extraction prefers /hrv/printhtml, then falls back to /hrv/html. Conversion method: markitdown_html This snapshot does not mirror PDFs. tabulartext-generation10K<n<100K0 likes25 downloads6mo agoHugging Face15taesiri /ArXivSignals-FullText ArXivSignals FullText — arXiv Papers OCR'd to Markdown + Layout A continuously-updated, day-partitioned dataset of arXiv papers converted to clean full text by a vision OCR pipeline: each paper's PDF is rendered to Markdown (headings, paragraphs, tables as HTML, math as LaTeX) plus a structured layout JSON (typed, bounding-boxed blocks). It is the full-text companion to taesiri/ArXivSignals (metadata + LLM signal & summaries) and joins it on paper_id. How it's made… See the full description on the dataset page: https://huggingface.co/datasets/taesiri/ArXivSignals-FullText.tabulartext-generation1K<n<10K0 likes22 downloads2mo agoHugging Face16CLEVDEV /full_before_conv_icmll_pub Iterative vs Recursive Code Pairs Coding problems sourced from LeetCode and Codeforces, each with a verified iterative and recursive Python solution. Test cases are timed and binned by per-problem difficulty (tc_difficulty: easy | medium | hard). Function names in both solutions are deterministically obfuscated (*_obfuscated columns) for benchmarks where lexical signal would leak the paradigm. Columns id, task_id, source, difficulty, title, description, tags, rating… See the full description on the dataset page: https://huggingface.co/datasets/CLEVDEV/full_before_conv_icmll_pub.tabulartext-generation1K<n<10K0 likes19 downloads5mo agoHugging Face17guhhhgu /harmonicbench-planir-main-full HARMONICBench PlanIR Main Full Export This dataset is a normalized Hugging Face export of the local outputs/fixed/main/domain1-domain5 PlanIR artifacts. Coverage Total domains: 5 Total samples: 500 Total condition-plan rows: 7254 Total reference-plan rows: 500 Total domain5 image-description rows: 100 Per-domain coverage: domain1: 100 samples, 1441 condition-plan rows domain2: 100 samples, 1557 condition-plan rows domain3: 100 samples, 1658 condition-plan rows domain4:… See the full description on the dataset page: https://huggingface.co/datasets/guhhhgu/harmonicbench-planir-main-full.tabulartext-generation1K<n<10K0 likes9 downloads5mo agoHugging Face18OptiTransferData /swiss-web-premium-ch-fullgated *.ch Swiss Web Premium (A+) -- Full Dataset Grade A+ -- 112.4M tokens -- 110,491 records -- 29 languages -- 22 production files -- 554.1 MB The complete production release of OptiTransfer's Swiss web corpus. 110,491 documents from the .ch TLD namespace, independently quality-scored (avg 93.3/100, minimum 90), PII-redacted, and SHA256-verified. Delivered in Parquet, JSONL, language splits, and pre-built RAG chunks. This is the full commercial dataset. For evaluation, see the free… See the full description on the dataset page: https://huggingface.co/datasets/OptiTransferData/swiss-web-premium-ch-full.tabulartext-generation100K<n<1M0 likes5 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.