datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
migration-bench-java-full
MigrationBench
1. 📖 Overview
🤗 MigrationBench
is a large-scale code migration benchmark dataset at the repository level,
across multiple programming languages.
Current and initial release includes java 8 repositories with the maven build system… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/migration-bench-java-full.fulg
❄️FuLG
The FuLG dataset is a comprehensive Romanian language corpus comprising 150 billion tokens, carefully
extracted from Common Crawl. This extensive dataset is the result of rigorous filtering and deduplication
processes applied to 95 Common Crawl snapshots. The compressed dataset has 289 GB.
For more details, check the arXiv preprint.
How do I download this?
Using 🤗 Datasets
from datasets import load_dataset
# Full dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/faur-ai/fulg.R2E-Gym-Full
R2E-Gym Subset Filtered for MAGRPO
Filtered subset of R2E-Gym optimized for 2-agent MAGRPO training with 7B models.
Dataset Statistics
Total instances: 167
Format: Issue description + Oracle files in prompt
Optimized for: 2-agent collaboration, 7B models
Filtering Criteria (SWE-bench Lite Style)
Problem statement: >40 words (up to 500 for context window)
Must have non-empty oracle patch (non-test file changes)
File count: Exactly 1 oracle file (single-file… See the full description on the dataset page: https://huggingface.co/datasets/ryankamiri/R2E-Gym-Full.the-stack-v2-train-full-ids
The Stack v2
The dataset consists of 4 versions:
bigcode/the-stack-v2: the full "The Stack v2" dataset
bigcode/the-stack-v2-dedup: based on the bigcode/the-stack-v2 but further near-deduplicated
bigcode/the-stack-v2-train-full-ids: based on the bigcode/the-stack-v2-dedup dataset but further filtered with heuristics and spanning 600+ programming languages. The data is grouped into repositories. <-- you are here
bigcode/the-stack-v2-train-smol-ids: based on the… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/the-stack-v2-train-full-ids.turk-ictihat-kararlari-fulltext
Türk İçtihat Kararları - Tam Metin (UYAP e-Bedesten)
Adalet Bakanlığı UYAP Mevzuat ve İçtihat Programı (bedesten.adalet.gov.tr)
üzerinden derlenen Yargıtay kararları veri seti. Kavram bazlı arama
sorguları ile toplanmıştır.
Boyut
Kayıt: 9,899,589 benzersiz karar
Terim dağılımı: {"bosanma": 100052, "karar": 9894542, "nafaka": 66241, "tazminat": 100052}
Yıl aralığı: 1993-2026
Tam metni gelen karar: 39,301 (artımlı olarak dolduruluyor)
Şema… See the full description on the dataset page: https://huggingface.co/datasets/muhammedturan/turk-ictihat-kararlari-fulltext.conseil-detat-full-documents
Décisions du Conseil d'État (France)
Description
Ce dataset contient un corpus de décisions rendues par le Conseil d'État français, la plus haute juridiction de l'ordre administratif. Les décisions proviennent de la plateforme officielle Open Data de la Justice Administrative et sont diffusées au format XML anonymisé.
Le corpus rassemble les textes intégraux des décisions ainsi que plusieurs métadonnées permettant leur identification et leur traçabilité.
Source… See the full description on the dataset page: https://huggingface.co/datasets/hulk10/conseil-detat-full-documents.flutter-full-examples-v1
Flutter Codegen: Full Examples
Synthetic dataset of complete Flutter/Dart widgets, each paired with the goal
that describes them and (optionally) starting code. Unlike flutter-codegen-diff-steps,
there's no step history or diff structure here -- each row is a single, standalone
goal -> complete file example.
This is the whole-code counterpart to flutter-diff-steps-v1, intended for
training/evaluating a baseline that generates the entire file in one shot, to
compare against the… See the full description on the dataset page: https://huggingface.co/datasets/bbidpa/flutter-full-examples-v1.creativemath_fullFull-Ecom-Chatbot-Dataset
E-commerce Chatbot Training Data
A curated, multi-source dataset for training and evaluating e-commerce conversational AI systems. It covers a broad range of customer intents — from product discovery and order management to returns, tool-augmented responses, and RAG-grounded Q&A — across 16+ product domains.
Dataset Summary
Split
Records
Train
35,213
Test
8,818
Total
44,031
The train/test split uses prompt-group-level stratified sampling on source ×… See the full description on the dataset page: https://huggingface.co/datasets/rescommons/Full-Ecom-Chatbot-Dataset.lalm-judge-validation-full-duplex
LALM Judge Validation on Full-Duplex Voice Agents
Companion dataset for the paper A Reliability Assessment of
LALM Audio Judges for Full-Duplex Voice Agents.
This repository contains the anonymised ratings, adversarial-defect
recall tables, JSON schemas, and analysis scripts used to produce
every headline number, table, and figure in that paper.
Summary
209 rated stereo sessions: 152 full-duplex agent-client
conversations across 13 accent-and-condition strata… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/lalm-judge-validation-full-duplex.stxbp1-pubmed-central-fulltext
source_datasets:
- PubMed Central
STXBP1 PubMed Central Full-Text Dataset v2
A comprehensive collection of 31,786 full-text scientific articles from PubMed Central related to STXBP1, synaptic function, and neurological research.
🆕 Version 2 Updates (December 2025)
Complete re-extraction with improved HTML parsing
Full main text with proper section headers
Enhanced metadata extraction
99.7% figure-image matching (see companion multimodal dataset)… See the full description on the dataset page: https://huggingface.co/datasets/SkyWhal3/stxbp1-pubmed-central-fulltext.negopt_fullThis is the dataset constructed in and used to fine-tune the models proposed in our paper Optimizing Negative Prompts for Enhanced Aesthetics and Fidelity in Text-To-Image Generation. The data is gotten from Playground.
If you find this dataset useful, please cite us here:
@article{ogezi2024optimizing,
title={Optimizing Negative Prompts for Enhanced Aesthetics and Fidelity in Text-To-Image Generation},
author={Ogezi, Michael and Shi, Ning},
journal={arXiv preprint arXiv:2403.07605}… See the full description on the dataset page: https://huggingface.co/datasets/mikeogezi/negopt_full.ru-wikipedia-100k-full-text-daily-stats-10-years
📚 Russian Wikipedia Top 100K: Full Text with Daily Pageviews
**Крупнейший открытый датасет русскоязычной Википедии с полными текстами статей и ежедневной статистикой просмотров за 10 лет. МГУ, ОТиПЛ, 2025**
📖 Описание
Этот датасет содержит 99,348 самых популярных статей русскоязычной Википедии, отобранных по совокупному количеству просмотров за последние 10 лет.
Для каждой статьи собраны полный текст с сохранением структуры, метаданные и детальная ежедневная… See the full description on the dataset page: https://huggingface.co/datasets/Mikimi/ru-wikipedia-100k-full-text-daily-stats-10-years.narodne-novine-full-text-markdown
Narodne Novine Full Text Markdown
HTML-to-Markdown text extraction snapshot derived from the NN archive.
Coverage
Extracted acts: 96855
Failed acts: 157
Missing HTML embodiments: 150
Files
texts.parquet
failures.parquet
metadata.json
Notes
Extraction prefers /hrv/printhtml, then falls back to /hrv/html.
Conversion method: markitdown_html
This snapshot does not mirror PDFs.
ArXivSignals-FullText
ArXivSignals FullText — arXiv Papers OCR'd to Markdown + Layout
A continuously-updated, day-partitioned dataset of arXiv papers converted to
clean full text by a vision OCR pipeline: each paper's PDF is rendered to
Markdown (headings, paragraphs, tables as HTML, math as LaTeX) plus a structured
layout JSON (typed, bounding-boxed blocks). It is the full-text companion to
taesiri/ArXivSignals
(metadata + LLM signal & summaries) and joins it on paper_id.
How it's made… See the full description on the dataset page: https://huggingface.co/datasets/taesiri/ArXivSignals-FullText.full_before_conv_icmll_pub
Iterative vs Recursive Code Pairs
Coding problems sourced from LeetCode and Codeforces, each with a verified
iterative and recursive Python solution. Test cases are timed and binned
by per-problem difficulty (tc_difficulty: easy | medium | hard). Function names
in both solutions are deterministically obfuscated (*_obfuscated columns) for
benchmarks where lexical signal would leak the paradigm.
Columns
id, task_id, source, difficulty, title, description, tags, rating… See the full description on the dataset page: https://huggingface.co/datasets/CLEVDEV/full_before_conv_icmll_pub.harmonicbench-planir-main-full
HARMONICBench PlanIR Main Full Export
This dataset is a normalized Hugging Face export of the local outputs/fixed/main/domain1-domain5 PlanIR artifacts.
Coverage
Total domains: 5
Total samples: 500
Total condition-plan rows: 7254
Total reference-plan rows: 500
Total domain5 image-description rows: 100
Per-domain coverage:
domain1: 100 samples, 1441 condition-plan rows
domain2: 100 samples, 1557 condition-plan rows
domain3: 100 samples, 1658 condition-plan rows
domain4:… See the full description on the dataset page: https://huggingface.co/datasets/guhhhgu/harmonicbench-planir-main-full.swiss-web-premium-ch-full
*.ch Swiss Web Premium (A+) -- Full Dataset
Grade A+ -- 112.4M tokens -- 110,491 records -- 29 languages -- 22 production files -- 554.1 MB
The complete production release of OptiTransfer's Swiss web corpus. 110,491 documents from the .ch TLD namespace, independently quality-scored (avg 93.3/100, minimum 90), PII-redacted, and SHA256-verified. Delivered in Parquet, JSONL, language splits, and pre-built RAG chunks.
This is the full commercial dataset. For evaluation, see the free… See the full description on the dataset page: https://huggingface.co/datasets/OptiTransferData/swiss-web-premium-ch-full.
