datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
open-india-law
Open India Law
Open, structured Indian primary law - plus the scrapers that build it.
Every judgment of the Supreme Court of India and all 25 High Courts, the decisions of 15
tribunals and regulators, and Central, State and Union Territory legislation down to the
individual section. Normalized to one schema, exclusively from official government sources.
Volume
Period
Court judgments
12,848,644
1950 to 2025
Tribunal and regulator matters
813,168
1985 to 2026… See the full description on the dataset page: https://huggingface.co/datasets/vaquill/open-india-law.nadora-global-industries
NADORA Global Industries
A synthetic multinational, built to be developed against rather than
demonstrated with.
One fictional company — $5.20bn revenue, $716m EBITDA, 24,000 employees, 18
countries, 35 legal entities, five business units — traded daily from January
2022 to December 2026 and rendered at six fidelities, from a 3 MB unit-test
fixture to a 5 GB full-scale corpus.
38,964,663 rows · 11 GB · 2,319 verification assertions, all passing.
100% synthetic. No real company… See the full description on the dataset page: https://huggingface.co/datasets/hemanthreddy901/nadora-global-industries.open-library
Open Library
The complete Open Library catalog in clean, analysis-ready Parquet. 150.0M+ records across 11 entity types, from ISBNs and author bios to reading logs and Wikidata links.
What is it?
Open Library is a complete snapshot of the Open Library database, an open project of the Internet Archive with the mission of creating "one web page for every book ever published." The catalog is community-edited and contains bibliographic records for millions of authors, works… See the full description on the dataset page: https://huggingface.co/datasets/open-index/open-library.t5gemma2-indonesia-instruct-v1
T5Gemma-2 Indonesian Instruct — Mono-Repo
Satu repositori dataset HF untuk seluruh data pelatihan T5-Gemma-2 bahasa Indonesia.
Diorganisasi per fungsi (fondasi → spesifik → preferensi) dengan folder/subfolder,
setiap config = folder dan berisi split train + validation (80:20) di level percakapan.
Struktur (by fungsi)
t5gemma2-indonesia-instruct-v1/
├── README.md
├── manifest.json
├── chat_idx_map.json
├── foundation/ ← FASE 1 · fondasi Bahasa… See the full description on the dataset page: https://huggingface.co/datasets/daruokta/t5gemma2-indonesia-instruct-v1.agents-index
AgentCrush Agent Index
Evidence-ranked index of the AI agent economy. Updated daily from agentcrush.xyz.
Overview
1,445 agents indexed across categories: developer tools, tokenized agents, service agents, model families
207 evidence-ranked with verified multi-signal scores
Updated: 2026-09-26
Configs
Config
Description
Rows
agents
All indexed agents with metadata
~1,445
evidence_ranked
Evidence-ranked tier only
~207
snapshots_latest… See the full description on the dataset page: https://huggingface.co/datasets/AgentCrush/agents-index.indian-legal-documents
Indian Legal Documents
Open Indian statutory and legal-document data for AI, search, and legal research.
This dataset is part of the KanoonGPT Open Legal Data Initiative — an effort to make Indian legal data easier to access, structure, and build on for open-source research, legal tech, and production AI systems.
KanoonGPT is building structured Indian legal datasets and data infrastructure for open-source, research, and enterprise AI applications. Learn more at kanoongpt.in.… See the full description on the dataset page: https://huggingface.co/datasets/KanoonGPT/indian-legal-documents.IndustryInstruction_Finance-Economics
IndustryInstruction: Finance & Economics
This repository contains the IndustryInstruction: Finance & Economics domain subset of BAAI/IndustryInstruction.
Refer to the parent dataset card for data construction, intended use, limitations,
and licensing details.
Citation
If you use this dataset in your work, please cite IndustryInstruction:
@misc{shi2024industryinstruction,
title = {IndustryInstruction},
author = {Xiaofeng Shi and Lulu Zhao and Hua Zhou… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryInstruction_Finance-Economics.industrial-instruction-dataset
Industrial-Instruction Dataset
Industrial-Instruction provides benchmark and training-ready QA instances derived from industrial technical reports, designed to evaluate robustness under realistic retrieval conditions. Samples are grounded in retrieved evidence and include irrelevant retrieval, single-/multi-document support, and single-/multi-document answer settings.
Paper
Industrial-Instruction: An End-to-End Framework for Building Instruction-Tuning and… See the full description on the dataset page: https://huggingface.co/datasets/Parssky/industrial-instruction-dataset.global-censorship-index
Voidly Global Censorship Index
Real-time internet censorship measurements for 200 countries, based on 38,780,449+ OONI network probes.
Dataset Description
The Global Censorship Index provides country-level internet censorship scores derived from actual network measurements. Unlike annual expert assessments, this data updates daily.
Key Statistics
Countries covered: 200
Total measurements: 38,780,449
Severe censorship: 1 countries
High censorship: 6… See the full description on the dataset page: https://huggingface.co/datasets/emperor-mew/global-censorship-index.index-souverainete
Index Souveraineté Le Souv
Le dataset public de référence sur les entreprises françaises stratégiques cédées à des capitaux étrangers, et sur les entreprises souveraines à capitaux français.
Source canonique : Le Souv — média indépendant consacré à la souveraineté économique et politique française.
URL canonique : https://lesouv.fr/index-souverainete.json
Licence : Creative Commons BY 4.0 (réutilisation libre avec attribution à Le Souv).
Mise à jour : continue, regénération… See the full description on the dataset page: https://huggingface.co/datasets/LeSouv/index-souverainete.IndustryInstruction_Aerospace
IndustryInstruction: Aerospace
This repository contains the IndustryInstruction: Aerospace domain subset of BAAI/IndustryInstruction.
Refer to the parent dataset card for data construction, intended use, limitations,
and licensing details.
Citation
If you use this dataset in your work, please cite IndustryInstruction:
@misc{shi2024industryinstruction,
title = {IndustryInstruction},
author = {Xiaofeng Shi and Lulu Zhao and Hua Zhou and Donglin Hao and… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryInstruction_Aerospace.IndustryInstruction_Artificial-Intelligence
IndustryInstruction: Artificial Intelligence
This repository contains the IndustryInstruction: Artificial Intelligence domain subset of BAAI/IndustryInstruction.
Refer to the parent dataset card for data construction, intended use, limitations,
and licensing details.
Citation
If you use this dataset in your work, please cite IndustryInstruction:
@misc{shi2024industryinstruction,
title = {IndustryInstruction},
author = {Xiaofeng Shi and Lulu Zhao and Hua… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryInstruction_Artificial-Intelligence.luatdo-graph
luatdo-graph
A knowledge graph over Vietnamese law, built from about 128,000 documents by luatdo.
This is the result of running the pipeline, published so that nobody has to run it again.
The pipeline takes days and several hundred dollars of model calls, and the output is the same for everyone.
What is in it
Nodes
8,175,346
Relationships
9,119,011
Node tables
14
Relationship tables
19
Node labels
13
Parquet
623MB across 47 files
Neo4j… See the full description on the dataset page: https://huggingface.co/datasets/open-index/luatdo-graph.IndustryInstruction_Technology-Research
IndustryInstruction: Technology & Research
This repository contains the IndustryInstruction: Technology & Research domain subset of BAAI/IndustryInstruction.
Refer to the parent dataset card for data construction, intended use, limitations,
and licensing details.
Citation
If you use this dataset in your work, please cite IndustryInstruction:
@misc{shi2024industryinstruction,
title = {IndustryInstruction},
author = {Xiaofeng Shi and Lulu Zhao and Hua… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryInstruction_Technology-Research.IndustryInstruction_Automobiles
IndustryInstruction: Automobiles
This repository contains the IndustryInstruction: Automobiles domain subset of BAAI/IndustryInstruction.
Refer to the parent dataset card for data construction, intended use, limitations,
and licensing details.
Citation
If you use this dataset in your work, please cite IndustryInstruction:
@misc{shi2024industryinstruction,
title = {IndustryInstruction},
author = {Xiaofeng Shi and Lulu Zhao and Hua Zhou and Donglin Hao… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryInstruction_Automobiles.IndustryInstruction_Transportation
IndustryInstruction: Transportation
This repository contains the IndustryInstruction: Transportation domain subset of BAAI/IndustryInstruction.
Refer to the parent dataset card for data construction, intended use, limitations,
and licensing details.
Citation
If you use this dataset in your work, please cite IndustryInstruction:
@misc{shi2024industryinstruction,
title = {IndustryInstruction},
author = {Xiaofeng Shi and Lulu Zhao and Hua Zhou and Donglin… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryInstruction_Transportation.IndustryInstruction_Literature-Emotions
IndustryInstruction: Literature & Emotions
This repository contains the IndustryInstruction: Literature & Emotions domain subset of BAAI/IndustryInstruction.
Refer to the parent dataset card for data construction, intended use, limitations,
and licensing details.
Citation
If you use this dataset in your work, please cite IndustryInstruction:
@misc{shi2024industryinstruction,
title = {IndustryInstruction},
author = {Xiaofeng Shi and Lulu Zhao and Hua… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryInstruction_Literature-Emotions.IndustryInstruction_Law-Justice
IndustryInstruction: Law & Justice
This repository contains the IndustryInstruction: Law & Justice domain subset of BAAI/IndustryInstruction.
Refer to the parent dataset card for data construction, intended use, limitations,
and licensing details.
Citation
If you use this dataset in your work, please cite IndustryInstruction:
@misc{shi2024industryinstruction,
title = {IndustryInstruction},
author = {Xiaofeng Shi and Lulu Zhao and Hua Zhou and Donglin… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryInstruction_Law-Justice.indian-government-schemes-2025
Indian Government Schemes Dataset 2026
Dataset Description
The most comprehensive structured dataset of Indian central and state government schemes — 4,693 schemes across all ministries and states, with machine-readable eligibility fields.
Maintained by SmartDuke Technologies · Coimbatore, Tamil Nadu, India
This dataset powers SchemeFit — India's government scheme finder for citizens and businesses.
What Makes This Different
Most existing Indian… See the full description on the dataset page: https://huggingface.co/datasets/smartduketech/indian-government-schemes-2025.IndustryInstruction_Travel-Geography
IndustryInstruction: Travel & Geography
This repository contains the IndustryInstruction: Travel & Geography domain subset of BAAI/IndustryInstruction.
Refer to the parent dataset card for data construction, intended use, limitations,
and licensing details.
Citation
If you use this dataset in your work, please cite IndustryInstruction:
@misc{shi2024industryinstruction,
title = {IndustryInstruction},
author = {Xiaofeng Shi and Lulu Zhao and Hua Zhou and… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryInstruction_Travel-Geography.IndustryInstruction_Hospitality-Catering
IndustryInstruction: Hospitality Catering
This repository contains the IndustryInstruction: Hospitality Catering domain subset of BAAI/IndustryInstruction.
Refer to the parent dataset card for data construction, intended use, limitations,
and licensing details.
Citation
If you use this dataset in your work, please cite IndustryInstruction:
@misc{shi2024industryinstruction,
title = {IndustryInstruction},
author = {Xiaofeng Shi and Lulu Zhao and Hua Zhou… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryInstruction_Hospitality-Catering.IndustryInstruction本数据集为行业指令数据集,目前包含的行业中英文对照名称如下,本次数据旨在补充当前行业指令数据的空白,并挖掘BAAI/IndustryCorpus2预训练数据集中高质量预训练语料中包含的行业高价值知识。
汽车 : Automobiles
航空航天 : Aerospace
人工智能_机器学习 : Artificial-Intelligence
交通运输 : Transportation
科技_科学研究 : Technology-Research
法律_司法 : Law-Justice
金融_经济 : Finance-Economics
文学_情感 : Literature-Emotions
旅游_地理 : Travel-Geography
住宿_餐饮_酒店 : Hospitality-Catering
医疗 : Health-Medicine
学科教育 : Subject-Education
我们为每个数据集目录下面都提供了对应行业数据的 词云可视化和 数据质量分布曲线。如果需要单独行业的数据,可以跳转到单独的行业数据集地址… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryInstruction.IndicDB
IndicDB — Multilingual Text-to-SQL Benchmark for Indian Languages
IndicDB is a comprehensive multilingual Text-to-SQL benchmark for evaluating cross-lingual semantic parsing across diverse Indic language families. Questions are posed in 7 languages while the underlying schemas and values remain in English — simultaneously stressing translation, schema linking, value grounding, and multi-table join reasoning.
Schemas are sourced from real Indian open-data platforms (NDAP and… See the full description on the dataset page: https://huggingface.co/datasets/roshankaranth/IndicDB.IndoCareer
Introduction
IndoCareer is a dataset comprising 8,834 multiple-choice questions designed to evaluate performance in vocational and professional certification exams across various fields. With a focus on Indonesia, IndoCareer provides rich local contexts, spanning six key sectors: (1) healthcare, (2) insurance and finance, (3) creative and design, (4) tourism and hospitality, (5) education and training, and (6) law.
Data
Each question in the dataset is… See the full description on the dataset page: https://huggingface.co/datasets/indolem/IndoCareer.neophyte-faiss-index-v1
neophyte-faiss-index-v1
A FAISS index + metadata for scientific retrieval
Contents
index.faiss: FAISS index (cosine w/ inner product).
meta.jsonl: one JSON per chunk; fields include chunk_id, paper_id, title, section, subsection, paragraph_index, keywords, boost.
index.info.json: (optional) dimensions, index type, faiss version.
Build provenance
Chunking: hierarchical (section→paragraph→~480-token chunks, ~15% overlap)
Embedder:… See the full description on the dataset page: https://huggingface.co/datasets/bio-protocol/neophyte-faiss-index-v1.indoqaThis dataset is built for question answering task.indommlu-local-languages
IndoMMLU: Local Languages and Cultures (audited subset)
An audited, corrected subset of IndoMMLU
(Koto et al., 2023) covering the 9 Local Languages and Cultures subjects:
Indonesian primary and secondary school exam questions written in Balinese,
Banjarese, Dayak Ngaju, Javanese, Lampung, Madurese, Makassarese, and
Sundanese, plus one culture-knowledge subject on Minangkabau customs
(answered in standard Indonesian). This is not a dataset we created.
It is IndoMMLU's own subset… See the full description on the dataset page: https://huggingface.co/datasets/ibahasa/indommlu-local-languages.indo-bloom-corpus
🇮🇩 Indo-Bloom-AQG: A Unified Framework for Controllable Indonesian AQG
⚠️ RESEARCH ARTIFACT STATUS: SILVER VERSION (Work in Progress)
This dataset serves as the preliminary corpus (Silver Standard) for the ongoing Doctoral Dissertation at Universitas Negeri Malang (UM).
Current State: Unannotated / Pre-validation with Heuristic Bloom Labels
Target Final State: Gold Standard (Expert Validated with Bloom's Taxonomy Labels)
🔒 FROZEN — v0.1 Silver
This version is permanently… See the full description on the dataset page: https://huggingface.co/datasets/Firmansyah-Ibrahim/indo-bloom-corpus.t5gemma2-indonesia-chat-formatted
T5Gemma-2 Indonesian Chat & QA Dataset
A high-quality Indonesian language multi-turn conversation and reading comprehension dataset, specifically formatted for instruction tuning of sequence-to-sequence (Seq2Seq) models like T5-Gemma / T5-Gemma-2.
Dataset Description
This dataset contains over 7,400 multi-turn conversations and document-based Q&A in Bahasa Indonesia. It covers diverse topics including everyday life, technology, general knowledge, and structured… See the full description on the dataset page: https://huggingface.co/datasets/daruokta/t5gemma2-indonesia-chat-formatted.K-Paths-inductive-reasoning-drugbank
🔗 This dataset is part of the study:
K-Paths: Reasoning over Graph Paths for Drug Repurposing and Drug Interaction Prediction
📖 Read the Paper
💾 GitHub Repository
DrugBank: Inductive Reasoning Dataset
This dataset contains drug pairs annotated with 86 pharmacological relationships (e.g.,DrugA may increase the anticholinergic activities of DrugB).
Each entry includes two drugs, an interaction label, drug descriptions, and structured/natural language representations… See the full description on the dataset page: https://huggingface.co/datasets/Tassy24/K-Paths-inductive-reasoning-drugbank.
