datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
apex-r1-real-world-documents
Apex-R1 Real-World Benchmark Documents
This dataset stores real-world document/data assets collected for Apex-R1 synthetic long-horizon agentic RL workspace generation.
The files are intended as seed workspace materials, not as benchmark task labels. They can be injected into APEX-style filesystem/ or .apps_data/ environments to create more realistic and diverse professional-domain tasks.
Contents
benchmark_documents/
EnterpriseBench/ # CRM invoices… See the full description on the dataset page: https://huggingface.co/datasets/mtybilly/apex-r1-real-world-documents.jurisdb-legal-documents
JurisDB - Brazilian Legal Documents Dataset
Dataset Description
This dataset contains a comprehensive collection of Brazilian legal documents, including legislation (federal and state laws) and jurisprudence (court rulings and summaries from TSE, STJ, STF, and TNU).
Dataset Structure
.
├── legislacao_grifada_e_anotada_atualiz_em_01_01_2026/
│ ├── leis_estaduais/
│ ├── leis_federais/
│ └── ...
└── sumulas_tse_stj_stf_e_tnu_atualiz_01_01_2026_2/
├──… See the full description on the dataset page: https://huggingface.co/datasets/SherlockRamos/jurisdb-legal-documents.vietnamese-legal-documents
Vietnamese Legal Documents
A comprehensive collection of Vietnamese legal documents — laws, decrees, circulars, decisions, and other normative acts — sourced from vbpl.vn, the official Government Legal Document Portal operated by the Ministry of Justice. The dataset includes structured metadata for every document, raw HTML full-text content, and a rich graph of cross-document legal relationships (amendments, citations, repeals, etc.).
Curated by: Thịnh Ngô
Source: vbpl.vn… See the full description on the dataset page: https://huggingface.co/datasets/th1nhng0/vietnamese-legal-documents.ifpri-ai-documents-markdown
GAIA / GARDIAN-CIGI Agricultural Research Corpus
This dataset contains 21,726 agricultural research documents extracted from the GARDIAN repository and processed through the CIGI pipeline.
Dataset Overview
Property
Value
Total Documents
21,726
Total Size
623.27 MB
Total Tokens
85,359,442
Total Pages
0
Languages
25
Unique Keywords
7,127
Resource Types
20
Date Generated
2026-07-31 02:55:20
Language Distribution… See the full description on the dataset page: https://huggingface.co/datasets/CGIAR/ifpri-ai-documents-markdown.conseil-detat-full-documents
Décisions du Conseil d'État (France)
Description
Ce dataset contient un corpus de décisions rendues par le Conseil d'État français, la plus haute juridiction de l'ordre administratif. Les décisions proviennent de la plateforme officielle Open Data de la Justice Administrative et sont diffusées au format XML anonymisé.
Le corpus rassemble les textes intégraux des décisions ainsi que plusieurs métadonnées permettant leur identification et leur traçabilité.
Source… See the full description on the dataset page: https://huggingface.co/datasets/hulk10/conseil-detat-full-documents.conseil-administratives-appel-full-documents
Décisions de Justice Administrative Françaises
Description du dataset
Ce dataset contient un corpus de décisions de justice administrative françaises extraites de la plateforme officielle Open Data de la Justice Administrative. Les documents sont publiés par le Conseil d'État dans le cadre de la politique d'ouverture des données publiques de la justice française. Les décisions sont diffusées au format XML et anonymisées conformément aux exigences légales relatives… See the full description on the dataset page: https://huggingface.co/datasets/hulk10/conseil-administratives-appel-full-documents.documents-Egyptian-Arabic
Egyptian Arabic Mega Corpus (EAMC) — 25M Unified Egyptian Dialect Dataset
The Largest Unified Open Corpus for Egyptian Arabic (Masri / arz)
25.5M Samples | 2.66 GB (Parquet) | 9 Configs | Apache 2.0 | Ready-to-train
Comprehensive coverage: Raw Text · Wikipedia · Conversations · Speech (Whisper) · Parallel Translation (EN↔EGY) · Trilingual QA · Wikipedia Quality Classification · Fake Review / Spam Detection
Dataset Summary
Egyptian Arabic Mega Corpus… See the full description on the dataset page: https://huggingface.co/datasets/ISLAM-PO/documents-Egyptian-Arabic.service_public_part-full-documents
🇫🇷 Dataset Service-Public.fr – Fiches administratives structurées
Ce dataset est constitué à partir des contenus officiels publiés sur la plateformeService-Public.fr.Il regroupe des fiches pratiques et ressources administratives à destination des particuliers et des professionnels, couvrant un large éventail de démarches et de thématiques de l’administration française.
La structure et la méthodologie de ce dataset sont fortement inspirées du dataset Service-Public.fr practical… See the full description on the dataset page: https://huggingface.co/datasets/hulk10/service_public_part-full-documents.indian-legal-documents
Indian Legal Documents
Open Indian statutory and legal-document data for AI, search, and legal research.
This dataset is part of the KanoonGPT Open Legal Data Initiative — an effort to make Indian legal data easier to access, structure, and build on for open-source research, legal tech, and production AI systems.
KanoonGPT is building structured Indian legal datasets and data infrastructure for open-source, research, and enterprise AI applications. Learn more at kanoongpt.in.… See the full description on the dataset page: https://huggingface.co/datasets/KanoonGPT/indian-legal-documents.service_public_pro-full-documentstravail_emploi-full-documents
🇫🇷 Dataset Ministère du Travail et de l’Emploi – Fiches structurées
Ce dataset est constitué à partir des contenus publics diffusés sur le site officiel duMinistère du Travail et de l’Emploi :https://travail-emploi.gouv.fr/
Les données sources proviennent du dépôt GitHub officiel de l’administration française :https://github.com/SocialGouv/fiches-travail-data
La structure et la logique générale de ce dataset sont inspirées du dataset Travail Emploi website Dataset, publié sur… See the full description on the dataset page: https://huggingface.co/datasets/hulk10/travail_emploi-full-documents.vietnamese-legal-documents
Vietnamese Legal Documents
A comprehensive dataset of 518,255 Vietnamese legal documents sourced from
thuvienphapluat.vn — the largest Vietnamese legal
document repository. The dataset covers laws, decrees, circulars, decisions, and
other official documents issued by Vietnamese government bodies, spanning from
1924 to 2026.
At a Glance
🗂️ Total documents
518,255
📅 Date range
1924 – 2026
🏛️ Issuing authorities
2,393 unique bodies
📋 Document types
36… See the full description on the dataset page: https://huggingface.co/datasets/vohuutridung/vietnamese-legal-documents.tribunal-administratif-full-documents
Tribunal Administratif - Full Documents
Description
Ce dataset contient un corpus de décisions issues des Tribunaux Administratifs français, converties en documents textuels exploitables pour les applications d'intelligence artificielle.
L'objectif est de fournir un corpus prêt à l'emploi pour :
le Retrieval-Augmented Generation (RAG) ;
la recherche juridique ;
la question-réponse ;
la classification documentaire ;
le fine-tuning de modèles de langage spécialisés… See the full description on the dataset page: https://huggingface.co/datasets/hulk10/tribunal-administratif-full-documents.worldbank-project-documents
Dataset Card for World Bank Project Documents
Dataset Summary
This is a dataset of documents related to World Bank development projects in the period 1947-2020. The dataset includes
the documents used to propose or describe projects when they are launched, and those in the review. The documents are indexed
by the World Bank project ID, which can be used to obtain features from multiple publicly available tabular datasets.
Supported Tasks and Leaderboards
No… See the full description on the dataset page: https://huggingface.co/datasets/lukesjordan/worldbank-project-documents.Vietnamese-Legal-Documents
Vietnamese Legal Documents Dataset
1. Dataset Summary
Raw data: tmnam20/BKAI-Legal-Retrieval
The Vietnamese Legal Documents Dataset is a benchmark dataset designed for legal information retrieval in the Vietnamese language. It consists of:
A corpus of legal documents.
Train/test splits containing natural language queries and their corresponding relevant documents.
This dataset is intended to support research and development in:
Information Retrieval (IR)… See the full description on the dataset page: https://huggingface.co/datasets/YuITC/Vietnamese-Legal-Documents.vietnamese-legal-documents
Vietnamese Legal Documents
A comprehensive dataset of 518,255 Vietnamese legal documents sourced from
thuvienphapluat.vn — the largest Vietnamese legal
document repository. The dataset covers laws, decrees, circulars, decisions, and
other official documents issued by Vietnamese government bodies, spanning from
1924 to 2026.
At a Glance
🗂️ Total documents
518,255
📅 Date range
1924 – 2026
🏛️ Issuing authorities
2,393 unique bodies
📋 Document types
36… See the full description on the dataset page: https://huggingface.co/datasets/minhnguyent546/vietnamese-legal-documents.vietnamese-legal-documents
Vietnamese Legal Documents
A comprehensive collection of Vietnamese legal documents — laws, decrees, circulars, decisions, and other normative acts — sourced from vbpl.vn, the official Government Legal Document Portal operated by the Ministry of Justice. The dataset includes structured metadata for every document, raw HTML full-text content, and a rich graph of cross-document legal relationships (amendments, citations, repeals, etc.).
Curated by: Thịnh Ngô
Source: vbpl.vn
Language:… See the full description on the dataset page: https://huggingface.co/datasets/anhquan12/vietnamese-legal-documents.vietnamese-legal-documents
Vietnamese Legal Documents
A comprehensive dataset of 518,255 Vietnamese legal documents sourced from
thuvienphapluat.vn — the largest Vietnamese legal
document repository. The dataset covers laws, decrees, circulars, decisions, and
other official documents issued by Vietnamese government bodies, spanning from
1924 to 2026.
At a Glance
🗂️ Total documents
518,255
📅 Date range
1924 – 2026
🏛️ Issuing authorities
1,335 unique bodies
📋 Document types
36… See the full description on the dataset page: https://huggingface.co/datasets/pdt590/vietnamese-legal-documents.gardian-cigi-ai-documents
GAIA / GARDIAN-CIGI Agricultural Research Corpus
This dataset contains 85,782 agricultural research documents extracted from the GARDIAN repository and processed through the CIGI pipeline.
Dataset Overview
Property
Value
Total Documents
85,782
Total Size
1.47 GB
Total Tokens
199,872,861
Total Pages
0
Languages
59
Unique Keywords
76,373
Resource Types
32
Date Generated
2026-07-16 08:22:33
Language Distribution… See the full description on the dataset page: https://huggingface.co/datasets/CGIAR/gardian-cigi-ai-documents.vietnamese-legal-documents
Vietnamese Legal Documents
A comprehensive collection of Vietnamese legal documents — laws, decrees, circulars, decisions, and other normative acts — sourced from vbpl.vn, the official Government Legal Document Portal operated by the Ministry of Justice. The dataset includes structured metadata for every document, raw HTML full-text content, and a rich graph of cross-document legal relationships (amendments, citations, repeals, etc.).
Curated by: Thịnh Ngô
Source: vbpl.vn
Language:… See the full description on the dataset page: https://huggingface.co/datasets/vanhthefirst/vietnamese-legal-documents.embrapa-ai-documents
GAIA / GARDIAN-CIGI Agricultural Research Corpus
This dataset contains 130,922 agricultural research documents extracted from the GARDIAN repository and processed through the CIGI pipeline.
Dataset Overview
Property
Value
Total Documents
130,922
Total Size
256.66 MB
Total Tokens
11,827,230
Total Pages
0
Languages
8
Unique Keywords
106,492
Resource Types
11
Date Generated
2026-07-16 08:37:30
Language Distribution… See the full description on the dataset page: https://huggingface.co/datasets/CGIAR/embrapa-ai-documents.vietnamese-legal-documents
Vietnamese Legal Documents
A comprehensive dataset of 518,255 Vietnamese legal documents sourced from
thuvienphapluat.vn — the largest Vietnamese legal
document repository. The dataset covers laws, decrees, circulars, decisions, and
other official documents issued by Vietnamese government bodies, spanning from
1924 to 2026.
At a Glance
🗂️ Total documents
518,255
📅 Date range
1924 – 2026
🏛️ Issuing authorities
1,335 unique bodies
📋 Document types
36… See the full description on the dataset page: https://huggingface.co/datasets/lovienal/vietnamese-legal-documents.vietnamese-legal-documents
Vietnamese Legal Documents
A comprehensive dataset of 518,255 Vietnamese legal documents sourced from
thuvienphapluat.vn — the largest Vietnamese legal
document repository. The dataset covers laws, decrees, circulars, decisions, and
other official documents issued by Vietnamese government bodies, spanning from
1924 to 2026.
At a Glance
🗂️ Total documents
518,255
📅 Date range
1924 – 2026
🏛️ Issuing authorities
2,393 unique bodies
📋 Document types
36… See the full description on the dataset page: https://huggingface.co/datasets/minhdoan17/vietnamese-legal-documents.indian-legal-documents
Indian Legal Documents
Open Indian statutory and legal-document data for AI, search, and legal research.
This dataset is part of the KanoonGPT Open Legal Data Initiative — an effort to make Indian legal data easier to access, structure, and build on for open-source research, legal tech, and production AI systems.
KanoonGPT is building structured Indian legal datasets and data infrastructure for open-source, research, and enterprise AI applications. Learn more at kanoongpt.in.… See the full description on the dataset page: https://huggingface.co/datasets/abby2231/indian-legal-documents.ifpri-ai-documents
GAIA / GARDIAN-CIGI Agricultural Research Corpus
This dataset contains 21,726 agricultural research documents extracted from the GARDIAN repository and processed through the CIGI pipeline.
Dataset Overview
Property
Value
Total Documents
21,726
Total Size
623.27 MB
Total Tokens
85,359,442
Total Pages
0
Languages
25
Unique Keywords
7,127
Resource Types
20
Date Generated
2026-07-31 02:55:20
Language Distribution… See the full description on the dataset page: https://huggingface.co/datasets/CGIAR/ifpri-ai-documents.cirad-ai-documents
GAIA / CIRAD Agricultural Documents (English)
A curated, machine-readable corpus of 4,209 agricultural research
publications sourced from
CIRAD (the French Agricultural Research
Centre for International Development), produced by the
Generative AI for Agriculture (GAIA)
project. Documents are indexed through
GARDIAN — CGIAR's agri-food
research index — and converted from PDF to structured JSON via the
GAIA-CIGI pipeline using GROBID.
This is an independent CIRAD corpus in the… See the full description on the dataset page: https://huggingface.co/datasets/CGIAR/cirad-ai-documents.vietnamese-legal-documents
Vietnamese Legal Documents
A comprehensive dataset of 518,255 Vietnamese legal documents sourced from
thuvienphapluat.vn — the largest Vietnamese legal
document repository. The dataset covers laws, decrees, circulars, decisions, and
other official documents issued by Vietnamese government bodies, spanning from
1924 to 2026.
At a Glance
🗂️ Total documents
518,255
📅 Date range
1924 – 2026
🏛️ Issuing authorities
1,335 unique bodies
📋 Document types
36… See the full description on the dataset page: https://huggingface.co/datasets/dienmoc/vietnamese-legal-documents.usda-nal-ai-documents-en
GAIA / GARDIAN-CIGI Agricultural Research Corpus
This dataset contains 22,529 agricultural research documents extracted from the GARDIAN repository and processed through the CIGI pipeline.
Dataset Overview
Property
Value
Total Documents
22,529
Total Size
89.42 MB
Total Tokens
5,227,798
Total Pages
0
Languages
1
Unique Keywords
46,728
Resource Types
1
Date Generated
2026-07-16 08:32:49
Language Distribution… See the full description on the dataset page: https://huggingface.co/datasets/CGIAR/usda-nal-ai-documents-en.
