CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01BAAI /IndustryCorpus[中文主页] Industry models play a crucial role in driving enterprise intelligence transformation and innovative development. High-quality industry data is key to improving the performance of large models and realizing industry applications. However, datasets currently used for industry model training generally suffer from issues such as insufficient data volume, low quality, and lack of domain expertise. To address these problems, we constructed and applied 22 industry data processing operators to… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus.texttext-generation100M<n<1B61 likes7.8k downloads1mo agoHugging Face02BAAI /IndustryCorpus_technology[中文主页] Industry models play a crucial role in driving enterprise intelligence transformation and innovative development. High-quality industry data is key to improving the performance of large models and realizing industry applications. However, datasets currently used for industry model training generally suffer from issues such as insufficient data volume, low quality, and lack of domain expertise. To address these problems, we constructed and applied 22 industry data processing operators to… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus_technology.texttext-generation10M<n<100M4 likes3.8k downloads1mo agoHugging Face03BAAI /IndustryCorpus_finance[中文主页] Industry models play a crucial role in driving enterprise intelligence transformation and innovative development. High-quality industry data is key to improving the performance of large models and realizing industry applications. However, datasets currently used for industry model training generally suffer from issues such as insufficient data volume, low quality, and lack of domain expertise. To address these problems, we constructed and applied 22 industry data processing operators to… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus_finance.texttext-generation10M<n<100M19 likes2.3k downloads1mo agoHugging Face04BAAI /IndustryCorpus_education[中文主页] Industry models play a crucial role in driving enterprise intelligence transformation and innovative development. High-quality industry data is key to improving the performance of large models and realizing industry applications. However, datasets currently used for industry model training generally suffer from issues such as insufficient data volume, low quality, and lack of domain expertise. To address these problems, we constructed and applied 22 industry data processing operators to… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus_education.texttext-generation10M<n<100M4 likes2.1k downloads1mo agoHugging Face05BAAI /IndustryCorpus_news[中文主页] Industry models play a crucial role in driving enterprise intelligence transformation and innovative development. High-quality industry data is key to improving the performance of large models and realizing industry applications. However, datasets currently used for industry model training generally suffer from issues such as insufficient data volume, low quality, and lack of domain expertise. To address these problems, we constructed and applied 22 industry data processing operators to… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus_news.texttext-generation100M<n<1B5 likes1.8k downloads1mo agoHugging Face06nvidia /Nemotron-RL-Agentic-Indirect-Prompt-Injection-v1 Nemotron-RL-Agentic-Indirect-Prompt-Injection-v1 Dataset Description: Nemotron-RL-Agentic-Indirect-Prompt-Injection-v1 is an RL dataset for training and evaluating a tool-using agent's ability to resist Indirect Prompt Injection (IPI) attacks hidden inside tool-returned environment data. In each record, the agent receives a benign user request that requires calling a read tool whose output contains an adversarial instruction disguised as legitimate domain content… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Agentic-Indirect-Prompt-Injection-v1.textreinforcement-learning1K<n<10K8 likes1.4k downloads4mo agoHugging Face07NPCI /nemo-gym-indian-bankingThis is NPCI/nemo-gym-indian-banking — the dataset for the indian_banking resources server in NVIDIA NeMo Gym: 300 synthetic multi-turn Indian retail-banking customer-support tasks (250 train / 50 validation), the 197-customer synthetic bank database and the 59-article knowledge base the environment loads at startup. NeMo Gym Indian Banking Agent Tasks Tool-calling customer-service tasks for an Indian retail-banking assistant, in the NVIDIA NeMo Gym agent-input JSONL format.… See the full description on the dataset page: https://huggingface.co/datasets/NPCI/nemo-gym-indian-banking.texttext-generationn<1K3 likes422 downloads20d agoHugging Face08LeSouv /index-souverainete Index Souveraineté Le Souv Le dataset public de référence sur les entreprises françaises stratégiques cédées à des capitaux étrangers, et sur les entreprises souveraines à capitaux français. Source canonique : Le Souv — média indépendant consacré à la souveraineté économique et politique française. URL canonique : https://lesouv.fr/index-souverainete.json Licence : Creative Commons BY 4.0 (réutilisation libre avec attribution à Le Souv). Mise à jour : continue, regénération… See the full description on the dataset page: https://huggingface.co/datasets/LeSouv/index-souverainete.tabulartext-classificationn<1K1 likes339 downloads2d agoHugging Face09BAAI /IndustryCorpus_agriculture[中文主页] Industry models play a crucial role in driving enterprise intelligence transformation and innovative development. High-quality industry data is key to improving the performance of large models and realizing industry applications. However, datasets currently used for industry model training generally suffer from issues such as insufficient data volume, low quality, and lack of domain expertise. To address these problems, we constructed and applied 22 industry data processing operators to… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus_agriculture.texttext-generation1M<n<10M5 likes313 downloads1mo agoHugging Face10BAAI /IndustryCorpus_sports[中文主页] Industry models play a crucial role in driving enterprise intelligence transformation and innovative development. High-quality industry data is key to improving the performance of large models and realizing industry applications. However, datasets currently used for industry model training generally suffer from issues such as insufficient data volume, low quality, and lack of domain expertise. To address these problems, we constructed and applied 22 industry data processing operators to… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus_sports.texttext-generation10M<n<100M2 likes289 downloads1mo agoHugging Face11VmaxRL /SWEUniverse-Repaired-Indist-full-not-SWE-bench-pro-matched VmaxRL/SWEUniverse-Repaired-Indist-full-not-SWE-bench-pro-matched This dataset contains a 350-row subset selected from the Indist SWEUniverse training rows. Selection policy: three-way repo overlap with Bugpilot and LM-Modify, deduped by repo plus introduction patch, then balanced round-robin across overlapping repos. Rows: 350 Selected repos: 19 Deduped overlap capacity: 468 Source dataset: VmaxRL/SWEUniverse-Repaired-Indist-full-not-SWE-bench-pro-matched texttext-generationn<1K0 likes266 downloads4mo agoHugging Face12sinhal /Indian_Supreme_Court_Judgments Indian Supreme Court Judgments (Structured JSON) About Dataset This dataset contains fully extracted, structured JSON data for Indian Supreme Court judgments. This is a parsed, machine-readable version of the raw PDF repository, designed specifically for Natural Language Processing (NLP), RAG (Retrieval-Augmented Generation), and legal tech machine learning applications. The dataset is provided in .jsonl (JSON Lines) format. Each row represents a single case and contains… See the full description on the dataset page: https://huggingface.co/datasets/sinhal/Indian_Supreme_Court_Judgments.texttext-classification10K<n<100K1 likes178 downloads5mo agoHugging Face13BAAI /IndustryCorpus_literature[中文主页] Industry models play a crucial role in driving enterprise intelligence transformation and innovative development. High-quality industry data is key to improving the performance of large models and realizing industry applications. However, datasets currently used for industry model training generally suffer from issues such as insufficient data volume, low quality, and lack of domain expertise. To address these problems, we constructed and applied 22 industry data processing operators to… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus_literature.texttext-generation10M<n<100M2 likes172 downloads1mo agoHugging Face14BAAI /IndustryCorpus_medicine[中文主页] Industry models play a crucial role in driving enterprise intelligence transformation and innovative development. High-quality industry data is key to improving the performance of large models and realizing industry applications. However, datasets currently used for industry model training generally suffer from issues such as insufficient data volume, low quality, and lack of domain expertise. To address these problems, we constructed and applied 22 industry data processing operators to… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus_medicine.texttext-generation10M<n<100M8 likes169 downloads1mo agoHugging Face15open-index /tomo-traces Tomo Agent Traces Every tomo-labs run, published as it happens: the full agent trace, plus the boards and cost analyses regenerated from every result on each commit. What is it? This dataset is the running record of tomo-labs, the agent-evaluation harness for tomo and the coding agents it is measured against. Every time the harness runs a tool on a scenario, it captures the whole conversation the agent had with the model, converts it to the Hub's agent-trace… See the full description on the dataset page: https://huggingface.co/datasets/open-index/tomo-traces.tabulartext-generationn<1K0 likes156 downloads1mo agoHugging Face16169Pi /indian_law Indian Law Dataset The Indian Law Dataset is a high-quality, open-source dataset (~50M tokens) focused on Indian jurisprudence. It provides structured chain-of-thought reasoning traces across 10+ branches of law, enabling the training and evaluation of advanced reasoning-capable language models. Summary • Domain: Law / Indian Jurisprudence / Legal Reasoning • Scale: ~50M tokens, 47,789 rows • Source: Generated with advanced distillation techniques using structured… See the full description on the dataset page: https://huggingface.co/datasets/169Pi/indian_law.texttext-generation10K<n<100K4 likes151 downloads7mo agoHugging Face17LingoIITGN /IndicTalk IndicTalk:Code-Mixed Conversational Persona-Based Dataset This dataset contains multi-turn, persona-driven, code-mixed conversations generated from real news articles, across 9 Indian languages, in two script variants: Native — conversations written in the language's native script, code-mixed with Romanized English words. Romanized — conversations fully Romanized (Latin script), code-mixed with English. Each language has its own config, loadable independently, e.g.: from… See the full description on the dataset page: https://huggingface.co/datasets/LingoIITGN/IndicTalk.texttext-generation1M<n<10M2 likes146 downloads2mo agoHugging Face18Mxode /IndustryInstruction-Chinese 中文行业指令数据集 💻 Github Repo 简介 本数据集提取了原数据集 BAAI/IndustryInstruction 中源语言为中文的部分,并做了清洗。数据集分为单轮对话和多轮对话两个子集。 本数据集包含的行业及具体数据如下: 领域 单轮对话数目 多轮对话数目 AeroSpace 72667 0 Artificial-Intelligence 43906 0 Automobiles 78036 0 Finance-Economics 40135 0 Health-Medicine 177152 105320 Hospitality-Catering 39261 0 Law-Justice 43485 0 Literature-Emotions 44841 0 Subject-Education 271402 73 Technology-Research 41751 0 Transportation 51505 0 Travel-Geography 37150… See the full description on the dataset page: https://huggingface.co/datasets/Mxode/IndustryInstruction-Chinese.tabulartext-generation1M<n<10M2 likes135 downloads1y agoHugging Face19LorthGyu /indonesian-krama-ngoko Krama Ngoko Jawa 🇮🇩 Dataset tingkatan bahasa Jawa — Ngoko (kasar/sehari-hari), Krama (halus), Krama Ingge (paling halus). Satu makna, tiga cara ngomong, tergantung siapa lawan bicara. Kenapa dataset ini ada? Bahasa Jawa punya sistem undha-usuk (tingkatan bahasa) yang bikin LLM kewalahan — model sering nyampur ngoko & krama dalam satu kalimat. Dataset ini ngajarin model kapan pakai tingkatan yang mana. Dataset tingkatan bahasa Jawa di HF belum ada yang bagus —… See the full description on the dataset page: https://huggingface.co/datasets/LorthGyu/indonesian-krama-ngoko.texttext-generationn<1K0 likes123 downloads2mo agoHugging Face20inductionlabs /AgentWorldBench-Terminal-V2 AgentWorldBench-Terminal-V2 AgentWorldBench-Terminal-V2 is our improved subset of the terminal split from Qwen/AgentWorldBench (Zou et al., 2026). Given the history of a Linux terminal session, the model is evaluated on its ability to predict the output of the next command. In the original AgentWorldBench, some samples have ground-truth outputs that depend on environment details missing from the session history. Since the sessions are based on Terminal-Bench environments, the… See the full description on the dataset page: https://huggingface.co/datasets/inductionlabs/AgentWorldBench-Terminal-V2.tabulartext-generationn<1K0 likes119 downloads1mo agoHugging Face21BAAI /IndustryCorpus_politics[中文主页] Industry models play a crucial role in driving enterprise intelligence transformation and innovative development. High-quality industry data is key to improving the performance of large models and realizing industry applications. However, datasets currently used for industry model training generally suffer from issues such as insufficient data volume, low quality, and lack of domain expertise. To address these problems, we constructed and applied 22 industry data processing operators to… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus_politics.texttext-generation10M<n<100M5 likes114 downloads1mo agoHugging Face22maya-research /IndicVault Indic Vault — everyday Indian language QA pairs, tuned for chatbots & voice agents. 🧾 Overview Indic Vault is a high-quality, instruction-tuned dataset featuring question-answer pairs crafted in the contemporary, everyday language spoken across India in 2025. Unlike traditional datasets that lean heavily on formal or outdated linguistic styles, Indic Vault captures the authentic, colloquial expressions used in daily conversations, making it ideal for building AI… See the full description on the dataset page: https://huggingface.co/datasets/maya-research/IndicVault.textquestion-answering100K<n<1M70 likes112 downloads1y agoHugging Face23FoundryAILabs /k12-indian-curriculum-4.9m BharatLLM K-12 Indian Curriculum Dataset (4.9M) 4,904,936 question-answer pairs covering CBSE/NCERT K-12 curriculum across 12 Indian languages. Language Script Entries English Latin ~594K Hindi Devanagari ~449K Bengali Bengali ~408K Telugu Telugu ~408K Tamil Tamil ~408K Kannada Kannada ~408K Malayalam Malayalam ~408K Marathi Devanagari ~408K Gujarati Gujarati ~408K Odia Odia ~408K PunjabiGurmukhi ~408K Urdu Nastaliq ~374K Format {… See the full description on the dataset page: https://huggingface.co/datasets/FoundryAILabs/k12-indian-curriculum-4.9m.textquestion-answering1M<n<10M1 likes110 downloads6mo agoHugging Face24BAAI /IndustryCorpus_film[中文主页] Industry models play a crucial role in driving enterprise intelligence transformation and innovative development. High-quality industry data is key to improving the performance of large models and realizing industry applications. However, datasets currently used for industry model training generally suffer from issues such as insufficient data volume, low quality, and lack of domain expertise. To address these problems, we constructed and applied 22 industry data processing operators to… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus_film.texttext-generation10M<n<100M3 likes109 downloads1mo agoHugging Face25NPCI /IndicBankBench IndicBankBench — A Benchmark for Evaluating the Safety and Reliability of Language Models in Indian Retail Banking 📄 Paper · Code IndicBankBench evaluates whether language models behave safely and reliably in Indian retail-banking interactions. Its 799 synthetic, multi-turn cases test grounding in customer context, safe action-taking with mocked banking tools, appropriate clarification and refusal, and complete resolution of customer requests. All data is synthetic and contains… See the full description on the dataset page: https://huggingface.co/datasets/NPCI/IndicBankBench.texttext-generationn<1K0 likes109 downloads16h agoHugging Face26BAAI /IndustryCorpus_law[中文主页] Industry models play a crucial role in driving enterprise intelligence transformation and innovative development. High-quality industry data is key to improving the performance of large models and realizing industry applications. However, datasets currently used for industry model training generally suffer from issues such as insufficient data volume, low quality, and lack of domain expertise. To address these problems, we constructed and applied 22 industry data processing operators to… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus_law.texttext-generation10M<n<100M3 likes102 downloads1mo agoHugging Face27BAAI /IndustryCorpus_travel[中文主页] Industry models play a crucial role in driving enterprise intelligence transformation and innovative development. High-quality industry data is key to improving the performance of large models and realizing industry applications. However, datasets currently used for industry model training generally suffer from issues such as insufficient data volume, low quality, and lack of domain expertise. To address these problems, we constructed and applied 22 industry data processing operators to… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus_travel.texttext-generation10M<n<100M3 likes101 downloads1mo agoHugging Face28sthanika-ai /Indic-KCC-Agri-Advisory-Benchmarkgated Indic-KCC-Agri-Advisory-Benchmark ⚠️ Benchmark only — not agronomic advice. This dataset and its reference answers exist to score language models, not to be used as real farming guidance. KCC references are noisy call-centre transcripts (see Status and caveats); do not act on any answer, reference or candidate, as agricultural advice. Open-ended agricultural-advisory question answering in 11 Indian languages, built from real farmer questions and the advisory answers… See the full description on the dataset page: https://huggingface.co/datasets/sthanika-ai/Indic-KCC-Agri-Advisory-Benchmark.texttext-generation1K<n<10K4 likes100 downloads4d agoHugging Face29krittus /k12-indian-curriculum-4.9m BharatLLM K-12 Indian Curriculum Dataset (4.9M) 4,904,936 question-answer pairs covering CBSE/NCERT K-12 curriculum across 12 Indian languages. Language Script Entries English Latin ~594K Hindi Devanagari ~449K Bengali Bengali ~408K Telugu Telugu ~408K Tamil Tamil ~408K Kannada Kannada ~408K Malayalam Malayalam ~408K Marathi Devanagari ~408K Gujarati Gujarati ~408K Odia Odia ~408K Punjabi Gurmukhi ~408K Urdu Nastaliq ~374K Format… See the full description on the dataset page: https://huggingface.co/datasets/krittus/k12-indian-curriculum-4.9m.textquestion-answering1M<n<10M1 likes95 downloads6mo agoHugging Face30BAAI /IndustryCorpus_automobile[中文主页] Industry models play a crucial role in driving enterprise intelligence transformation and innovative development. High-quality industry data is key to improving the performance of large models and realizing industry applications. However, datasets currently used for industry model training generally suffer from issues such as insufficient data volume, low quality, and lack of domain expertise. To address these problems, we constructed and applied 22 industry data processing operators to… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus_automobile.texttext-generation1M<n<10M6 likes94 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.