datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
IndustryCorpus[中文主页]
Industry models play a crucial role in driving enterprise intelligence transformation and innovative development. High-quality industry data is key to improving the performance of large models and realizing industry applications. However, datasets currently used for industry model training generally suffer from issues such as insufficient data volume, low quality, and lack of domain expertise.
To address these problems, we constructed and applied 22 industry data processing operators to… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus.IndustryCorpus_technology[中文主页]
Industry models play a crucial role in driving enterprise intelligence transformation and innovative development. High-quality industry data is key to improving the performance of large models and realizing industry applications. However, datasets currently used for industry model training generally suffer from issues such as insufficient data volume, low quality, and lack of domain expertise.
To address these problems, we constructed and applied 22 industry data processing operators to… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus_technology.IndustryCorpus_finance[中文主页]
Industry models play a crucial role in driving enterprise intelligence transformation and innovative development. High-quality industry data is key to improving the performance of large models and realizing industry applications. However, datasets currently used for industry model training generally suffer from issues such as insufficient data volume, low quality, and lack of domain expertise.
To address these problems, we constructed and applied 22 industry data processing operators to… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus_finance.IndustryCorpus_education[中文主页]
Industry models play a crucial role in driving enterprise intelligence transformation and innovative development. High-quality industry data is key to improving the performance of large models and realizing industry applications. However, datasets currently used for industry model training generally suffer from issues such as insufficient data volume, low quality, and lack of domain expertise.
To address these problems, we constructed and applied 22 industry data processing operators to… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus_education.IndustryCorpus_news[中文主页]
Industry models play a crucial role in driving enterprise intelligence transformation and innovative development. High-quality industry data is key to improving the performance of large models and realizing industry applications. However, datasets currently used for industry model training generally suffer from issues such as insufficient data volume, low quality, and lack of domain expertise.
To address these problems, we constructed and applied 22 industry data processing operators to… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus_news.Nemotron-RL-Agentic-Indirect-Prompt-Injection-v1
Nemotron-RL-Agentic-Indirect-Prompt-Injection-v1
Dataset Description:
Nemotron-RL-Agentic-Indirect-Prompt-Injection-v1 is an RL dataset for training and evaluating a tool-using agent's ability to resist Indirect Prompt Injection (IPI) attacks hidden inside tool-returned environment data. In each record, the agent receives a benign user request that requires calling a read tool whose output contains an adversarial instruction disguised as legitimate domain content… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Agentic-Indirect-Prompt-Injection-v1.nemo-gym-indian-bankingThis is NPCI/nemo-gym-indian-banking — the dataset for the indian_banking
resources server in NVIDIA NeMo Gym: 300 synthetic
multi-turn Indian retail-banking customer-support tasks (250 train / 50 validation), the
197-customer synthetic bank database and the 59-article knowledge base the environment
loads at startup.
NeMo Gym Indian Banking Agent Tasks
Tool-calling customer-service tasks for an Indian retail-banking assistant, in the
NVIDIA NeMo Gym agent-input JSONL format.… See the full description on the dataset page: https://huggingface.co/datasets/NPCI/nemo-gym-indian-banking.index-souverainete
Index Souveraineté Le Souv
Le dataset public de référence sur les entreprises françaises stratégiques cédées à des capitaux étrangers, et sur les entreprises souveraines à capitaux français.
Source canonique : Le Souv — média indépendant consacré à la souveraineté économique et politique française.
URL canonique : https://lesouv.fr/index-souverainete.json
Licence : Creative Commons BY 4.0 (réutilisation libre avec attribution à Le Souv).
Mise à jour : continue, regénération… See the full description on the dataset page: https://huggingface.co/datasets/LeSouv/index-souverainete.IndustryCorpus_agriculture[中文主页]
Industry models play a crucial role in driving enterprise intelligence transformation and innovative development. High-quality industry data is key to improving the performance of large models and realizing industry applications. However, datasets currently used for industry model training generally suffer from issues such as insufficient data volume, low quality, and lack of domain expertise.
To address these problems, we constructed and applied 22 industry data processing operators to… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus_agriculture.IndustryCorpus_sports[中文主页]
Industry models play a crucial role in driving enterprise intelligence transformation and innovative development. High-quality industry data is key to improving the performance of large models and realizing industry applications. However, datasets currently used for industry model training generally suffer from issues such as insufficient data volume, low quality, and lack of domain expertise.
To address these problems, we constructed and applied 22 industry data processing operators to… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus_sports.SWEUniverse-Repaired-Indist-full-not-SWE-bench-pro-matched
VmaxRL/SWEUniverse-Repaired-Indist-full-not-SWE-bench-pro-matched
This dataset contains a 350-row subset selected from the Indist SWEUniverse training rows.
Selection policy: three-way repo overlap with Bugpilot and LM-Modify, deduped by repo plus introduction patch, then balanced round-robin across overlapping repos.
Rows: 350
Selected repos: 19
Deduped overlap capacity: 468
Source dataset: VmaxRL/SWEUniverse-Repaired-Indist-full-not-SWE-bench-pro-matched
Indian_Supreme_Court_Judgments
Indian Supreme Court Judgments (Structured JSON)
About Dataset
This dataset contains fully extracted, structured JSON data for Indian Supreme Court judgments. This is a parsed, machine-readable version of the raw PDF repository, designed specifically for Natural Language Processing (NLP), RAG (Retrieval-Augmented Generation), and legal tech machine learning applications.
The dataset is provided in .jsonl (JSON Lines) format. Each row represents a single case and contains… See the full description on the dataset page: https://huggingface.co/datasets/sinhal/Indian_Supreme_Court_Judgments.IndustryCorpus_literature[中文主页]
Industry models play a crucial role in driving enterprise intelligence transformation and innovative development. High-quality industry data is key to improving the performance of large models and realizing industry applications. However, datasets currently used for industry model training generally suffer from issues such as insufficient data volume, low quality, and lack of domain expertise.
To address these problems, we constructed and applied 22 industry data processing operators to… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus_literature.IndustryCorpus_medicine[中文主页]
Industry models play a crucial role in driving enterprise intelligence transformation and innovative development. High-quality industry data is key to improving the performance of large models and realizing industry applications. However, datasets currently used for industry model training generally suffer from issues such as insufficient data volume, low quality, and lack of domain expertise.
To address these problems, we constructed and applied 22 industry data processing operators to… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus_medicine.tomo-traces
Tomo Agent Traces
Every tomo-labs run, published as it happens: the full agent trace, plus the boards and cost analyses regenerated from every result on each commit.
What is it?
This dataset is the running record of tomo-labs, the agent-evaluation harness for tomo and the coding agents it is measured against.
Every time the harness runs a tool on a scenario, it captures the whole conversation the agent had with the model, converts it to the Hub's agent-trace… See the full description on the dataset page: https://huggingface.co/datasets/open-index/tomo-traces.indian_law
Indian Law Dataset
The Indian Law Dataset is a high-quality, open-source dataset (~50M tokens) focused on Indian jurisprudence. It provides structured chain-of-thought reasoning traces across 10+ branches of law, enabling the training and evaluation of advanced reasoning-capable language models.
Summary
• Domain: Law / Indian Jurisprudence / Legal Reasoning
• Scale: ~50M tokens, 47,789 rows
• Source: Generated with advanced distillation techniques using structured… See the full description on the dataset page: https://huggingface.co/datasets/169Pi/indian_law.IndicTalk
IndicTalk:Code-Mixed Conversational Persona-Based Dataset
This dataset contains multi-turn, persona-driven, code-mixed conversations
generated from real news articles, across 9 Indian
languages, in two script variants:
Native — conversations written in the language's native script,
code-mixed with Romanized English words.
Romanized — conversations fully Romanized (Latin script), code-mixed
with English.
Each language has its own config, loadable independently, e.g.:
from… See the full description on the dataset page: https://huggingface.co/datasets/LingoIITGN/IndicTalk.IndustryInstruction-Chinese
中文行业指令数据集
💻 Github Repo
简介
本数据集提取了原数据集 BAAI/IndustryInstruction 中源语言为中文的部分,并做了清洗。数据集分为单轮对话和多轮对话两个子集。
本数据集包含的行业及具体数据如下:
领域
单轮对话数目
多轮对话数目
AeroSpace
72667
0
Artificial-Intelligence
43906
0
Automobiles
78036
0
Finance-Economics
40135
0
Health-Medicine
177152
105320
Hospitality-Catering
39261
0
Law-Justice
43485
0
Literature-Emotions
44841
0
Subject-Education
271402
73
Technology-Research
41751
0
Transportation
51505
0
Travel-Geography
37150… See the full description on the dataset page: https://huggingface.co/datasets/Mxode/IndustryInstruction-Chinese.indonesian-krama-ngoko
Krama Ngoko Jawa 🇮🇩
Dataset tingkatan bahasa Jawa — Ngoko (kasar/sehari-hari), Krama (halus), Krama Ingge (paling halus). Satu makna, tiga cara ngomong, tergantung siapa lawan bicara.
Kenapa dataset ini ada?
Bahasa Jawa punya sistem undha-usuk (tingkatan bahasa) yang bikin LLM kewalahan — model sering nyampur ngoko & krama dalam satu kalimat. Dataset ini ngajarin model kapan pakai tingkatan yang mana. Dataset tingkatan bahasa Jawa di HF belum ada yang bagus —… See the full description on the dataset page: https://huggingface.co/datasets/LorthGyu/indonesian-krama-ngoko.AgentWorldBench-Terminal-V2
AgentWorldBench-Terminal-V2
AgentWorldBench-Terminal-V2 is our improved subset of the terminal split from
Qwen/AgentWorldBench
(Zou et al., 2026). Given the history of a Linux terminal session,
the model is evaluated on its ability to predict the output of the next command.
In the original AgentWorldBench, some samples have ground-truth outputs that depend on environment
details missing from the session history. Since the sessions are based on Terminal-Bench environments,
the… See the full description on the dataset page: https://huggingface.co/datasets/inductionlabs/AgentWorldBench-Terminal-V2.IndustryCorpus_politics[中文主页]
Industry models play a crucial role in driving enterprise intelligence transformation and innovative development. High-quality industry data is key to improving the performance of large models and realizing industry applications. However, datasets currently used for industry model training generally suffer from issues such as insufficient data volume, low quality, and lack of domain expertise.
To address these problems, we constructed and applied 22 industry data processing operators to… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus_politics.IndicVault
Indic Vault — everyday Indian language QA pairs, tuned for chatbots & voice agents.
🧾 Overview
Indic Vault is a high-quality, instruction-tuned dataset featuring question-answer pairs crafted in the contemporary, everyday language spoken across India in 2025. Unlike traditional datasets that lean heavily on formal or outdated linguistic styles, Indic Vault captures the authentic, colloquial expressions used in daily conversations, making it ideal for building AI… See the full description on the dataset page: https://huggingface.co/datasets/maya-research/IndicVault.k12-indian-curriculum-4.9m
BharatLLM K-12 Indian Curriculum Dataset (4.9M)
4,904,936 question-answer pairs covering CBSE/NCERT K-12 curriculum across 12 Indian languages.
Language
Script
Entries
English
Latin
~594K
Hindi
Devanagari
~449K
Bengali
Bengali
~408K
Telugu
Telugu
~408K
Tamil
Tamil
~408K
Kannada
Kannada
~408K
Malayalam
Malayalam
~408K
Marathi
Devanagari
~408K
Gujarati
Gujarati
~408K
Odia
Odia
~408K
PunjabiGurmukhi
~408K
Urdu
Nastaliq
~374K
Format
{… See the full description on the dataset page: https://huggingface.co/datasets/FoundryAILabs/k12-indian-curriculum-4.9m.IndustryCorpus_film[中文主页]
Industry models play a crucial role in driving enterprise intelligence transformation and innovative development. High-quality industry data is key to improving the performance of large models and realizing industry applications. However, datasets currently used for industry model training generally suffer from issues such as insufficient data volume, low quality, and lack of domain expertise.
To address these problems, we constructed and applied 22 industry data processing operators to… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus_film.IndicBankBench
IndicBankBench — A Benchmark for Evaluating the Safety and Reliability of Language Models in Indian Retail Banking
📄 Paper · Code
IndicBankBench evaluates whether language models behave safely and reliably in Indian
retail-banking interactions. Its 799 synthetic, multi-turn cases test grounding in customer
context, safe action-taking with mocked banking tools, appropriate clarification and refusal,
and complete resolution of customer requests. All data is synthetic and contains… See the full description on the dataset page: https://huggingface.co/datasets/NPCI/IndicBankBench.IndustryCorpus_law[中文主页]
Industry models play a crucial role in driving enterprise intelligence transformation and innovative development. High-quality industry data is key to improving the performance of large models and realizing industry applications. However, datasets currently used for industry model training generally suffer from issues such as insufficient data volume, low quality, and lack of domain expertise.
To address these problems, we constructed and applied 22 industry data processing operators to… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus_law.IndustryCorpus_travel[中文主页]
Industry models play a crucial role in driving enterprise intelligence transformation and innovative development. High-quality industry data is key to improving the performance of large models and realizing industry applications. However, datasets currently used for industry model training generally suffer from issues such as insufficient data volume, low quality, and lack of domain expertise.
To address these problems, we constructed and applied 22 industry data processing operators to… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus_travel.Indic-KCC-Agri-Advisory-Benchmark
Indic-KCC-Agri-Advisory-Benchmark
⚠️ Benchmark only — not agronomic advice. This dataset and its reference
answers exist to score language models, not to be used as real farming
guidance. KCC references are noisy call-centre transcripts (see Status and
caveats); do not act on any answer, reference or candidate, as agricultural
advice.
Open-ended agricultural-advisory question answering in 11 Indian languages,
built from real farmer questions and the advisory answers… See the full description on the dataset page: https://huggingface.co/datasets/sthanika-ai/Indic-KCC-Agri-Advisory-Benchmark.k12-indian-curriculum-4.9m
BharatLLM K-12 Indian Curriculum Dataset (4.9M)
4,904,936 question-answer pairs covering CBSE/NCERT K-12 curriculum across 12 Indian languages.
Language
Script
Entries
English
Latin
~594K
Hindi
Devanagari
~449K
Bengali
Bengali
~408K
Telugu
Telugu
~408K
Tamil
Tamil
~408K
Kannada
Kannada
~408K
Malayalam
Malayalam
~408K
Marathi
Devanagari
~408K
Gujarati
Gujarati
~408K
Odia
Odia
~408K
Punjabi
Gurmukhi
~408K
Urdu
Nastaliq
~374K
Format… See the full description on the dataset page: https://huggingface.co/datasets/krittus/k12-indian-curriculum-4.9m.IndustryCorpus_automobile[中文主页]
Industry models play a crucial role in driving enterprise intelligence transformation and innovative development. High-quality industry data is key to improving the performance of large models and realizing industry applications. However, datasets currently used for industry model training generally suffer from issues such as insufficient data volume, low quality, and lack of domain expertise.
To address these problems, we constructed and applied 22 industry data processing operators to… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus_automobile.
