CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01open-index /hacker-news Hacker News - Complete Archive Every Hacker News item since 2006, live-updated every 5 minutes What is it? This dataset contains the complete Hacker News archive: every story, comment, Ask HN, Show HN, job posting, and poll ever submitted to the site. Hacker News is one of the longest-running and most influential technology communities on the internet, operated by Y Combinator since 2007. It has become the de facto gathering place for founders, engineers… See the full description on the dataset page: https://huggingface.co/datasets/open-index/hacker-news.texttext-generation10M<n<100M343 likes57k downloads1mo agoHugging Face02BAAI /IndustryCorpus[中文主页] Industry models play a crucial role in driving enterprise intelligence transformation and innovative development. High-quality industry data is key to improving the performance of large models and realizing industry applications. However, datasets currently used for industry model training generally suffer from issues such as insufficient data volume, low quality, and lack of domain expertise. To address these problems, we constructed and applied 22 industry data processing operators to… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus.texttext-generation100M<n<1B61 likes7.8k downloads1mo agoHugging Face03commoncrawl /host-index-testing-v2 Common Crawl Host Index v2 GitHub: https://github.com/commoncrawl/cc-host-index Each crawl, we generate a Host Index, which aggregates information about each web hosted visited during the crawl. The information is aggregated from the Common Crawl columnar index, web graph, and raw crawler logs. Quickstart The dataset is Hive-partitioned on crawl (data/crawl=CC-MAIN-2025-18/*.parquet). Open the whole dataset once, then filter with WHERE crawl = '...': because… See the full description on the dataset page: https://huggingface.co/datasets/commoncrawl/host-index-testing-v2.tabulartext-generation1B<n<10B0 likes7.2k downloads11d agoHugging Face04BAAI /IndustryCorpus_technology[中文主页] Industry models play a crucial role in driving enterprise intelligence transformation and innovative development. High-quality industry data is key to improving the performance of large models and realizing industry applications. However, datasets currently used for industry model training generally suffer from issues such as insufficient data volume, low quality, and lack of domain expertise. To address these problems, we constructed and applied 22 industry data processing operators to… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus_technology.texttext-generation10M<n<100M4 likes3.8k downloads1mo agoHugging Face05BAAI /IndustryCorpus_finance[中文主页] Industry models play a crucial role in driving enterprise intelligence transformation and innovative development. High-quality industry data is key to improving the performance of large models and realizing industry applications. However, datasets currently used for industry model training generally suffer from issues such as insufficient data volume, low quality, and lack of domain expertise. To address these problems, we constructed and applied 22 industry data processing operators to… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus_finance.texttext-generation10M<n<100M19 likes2.3k downloads1mo agoHugging Face06BAAI /IndustryCorpus_education[中文主页] Industry models play a crucial role in driving enterprise intelligence transformation and innovative development. High-quality industry data is key to improving the performance of large models and realizing industry applications. However, datasets currently used for industry model training generally suffer from issues such as insufficient data volume, low quality, and lack of domain expertise. To address these problems, we constructed and applied 22 industry data processing operators to… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus_education.texttext-generation10M<n<100M4 likes2.1k downloads1mo agoHugging Face07ai4bharat /indic-align IndicAlign A diverse collection of Instruction and Toxic alignment datasets for 14 Indic Languages. The collection comprises of: IndicAlign - Instruct Indic-ShareLlama Dolly-T OpenAssistant-T WikiHow IndoWordNet Anudesh Wiki-Conv Wiki-Chat IndicAlign - Toxic HHRLHF-T Toxic-Matrix We use IndicTrans2 (Gala et al., 2023) for the translation of the datasets. We recommend the readers to check out our paper on Arxiv for detailed information on the curation process of these… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/indic-align.tabulartext-generation10M<n<100M21 likes2.1k downloads2y agoHugging Face08BAAI /IndustryCorpus_news[中文主页] Industry models play a crucial role in driving enterprise intelligence transformation and innovative development. High-quality industry data is key to improving the performance of large models and realizing industry applications. However, datasets currently used for industry model training generally suffer from issues such as insufficient data volume, low quality, and lack of domain expertise. To address these problems, we constructed and applied 22 industry data processing operators to… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus_news.texttext-generation100M<n<1B5 likes1.8k downloads1mo agoHugging Face09nvidia /Nemotron-Personas-India Nemotron-Personas-India A compound AI approach to personas grounded in real-world distributions वास्तविक दुनिया के वितरण पर आधारित व्यक्तित्वों के लिए एक मिश्रित AI दृष्टिकोण Dataset Overview (डेटासेट अवलोकन) Nemotron-Personas-India is an open-source (CC BY 4.0) dataset of synthetically-generated personas. This dataset is grounded in real-world demographic, geographic and personality trait distributions in India to capture the diversity and richness of… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Personas-India.imagetext-generation1M<n<10M56 likes1.7k downloads9mo agoHugging Face10open-index /open-arxiv Open arXiv Every arXiv paper's metadata in one place: search, filter, and explore 40 years of science What is it? Open arXiv is the complete arXiv metadata dataset, covering titles, abstracts, authors, categories, DOIs, version history, and more. It is converted from the Cornell University Kaggle dataset into Parquet format for efficient querying and streaming. The dataset contains 2.99M papers spanning from 1991 to 2026, packaged into 417 Parquet shards (Zstd… See the full description on the dataset page: https://huggingface.co/datasets/open-index/open-arxiv.texttext-classification1M<n<10M18 likes1.7k downloads6mo agoHugging Face11nvidia /Nemotron-RL-Agentic-Indirect-Prompt-Injection-v1 Nemotron-RL-Agentic-Indirect-Prompt-Injection-v1 Dataset Description: Nemotron-RL-Agentic-Indirect-Prompt-Injection-v1 is an RL dataset for training and evaluating a tool-using agent's ability to resist Indirect Prompt Injection (IPI) attacks hidden inside tool-returned environment data. In each record, the agent receives a benign user request that requires calling a read tool whose output contains an adversarial instruction disguised as legitimate domain content… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Agentic-Indirect-Prompt-Injection-v1.textreinforcement-learning1K<n<10K8 likes1.4k downloads4mo agoHugging Face12open-index /ccrawl-recrawl-domains Common Crawl Domain Recrawl Live fetches of the home page of every ranked domain in Common Crawl's web graph, rendered to Markdown as they are fetched What is it? Common Crawl's web graph ranks domains by how central they are, but it does not tell you what those domains actually serve today. This dataset walks that ranking from the top and fetches each domain's home page now, storing the response as one Parquet row with the body, the headers, the timing and the… See the full description on the dataset page: https://huggingface.co/datasets/open-index/ccrawl-recrawl-domains.tabulartext-generation1M<n<10M0 likes1.2k downloads1mo agoHugging Face13open-index /open-github OpenGitHub What is it? This dataset contains every public event on GitHub: every push, pull request, issue, star, fork, code review, release, and discussion across all public repositories. GitHub is the world's largest software development platform, home to over 200 million repositories and the daily work of tens of millions of developers, from individual open-source contributors to the engineering teams behind the most widely used software on earth. The archive currently… See the full description on the dataset page: https://huggingface.co/datasets/open-index/open-github.tabulartext-generation100K<n<1M9 likes953 downloads6mo agoHugging Face14indikamk /misconceptionstexttext-generationn<1K1 likes904 downloads3y agoHugging Face15ashtok897 /indic-hplt-v2 Indic HPLT v2 A multilingual pretraining corpus of 34,605,630 documents (~25.5B estimated tokens, ~218 GB raw JSONL) across 13 Indic languages and English, built from HPLT Monolingual v3 high-quality web crawl data. This is the larger successor to Indic HPLT v1 (9.8M docs, 11 languages). Compared to v1, this release adds 3 new Indic languages (Nepali, Odia, Assamese) and ~3.5× more documents overall. Quick Start from datasets import load_dataset # Full training… See the full description on the dataset page: https://huggingface.co/datasets/ashtok897/indic-hplt-v2.tabulartext-generation10M<n<100M3 likes827 downloads4mo agoHugging Face16ai4bharat /IndicIFEval IndicIFEval Paper | GitHub Instruction-following benchmarks remain predominantly English-centric, leaving a critical evaluation gap for the hundreds of millions of Indic language speakers. We introduce IndicIFEval, a benchmark evaluating constrained generation of LLMs across 14 Indic languages using automatically verifiable, rule-based instructions. It combines two complementary tracks: IndicIFEval-Trans, translated prompts from IFEval carefully localized for Indic… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/IndicIFEval.texttext-generation10K<n<100K4 likes774 downloads10d agoHugging Face17ashtok897 /indic-hplt-v1 Indic HPLT v1 A multilingual pretraining corpus of 9,836,075 documents (~8.4B estimated tokens) across 10 Indic languages and English, built from HPLT Monolingual v3 high-quality web crawl data. Quick Start from datasets import load_dataset # Full training split ds = load_dataset("ashtok897/indic-hplt-v1", split="train") # Filter by language hi_ds = ds.filter(lambda x: x["lang"] == "hi") # Streaming (recommended for large-scale use) ds =… See the full description on the dataset page: https://huggingface.co/datasets/ashtok897/indic-hplt-v1.tabulartext-generation1M<n<10M4 likes742 downloads4mo agoHugging Face18AdaMLLab /IndMix IndMix (https://arxiv.org/abs/2512.18834) is an Indonesian pretraining corpus built by combining six publicly available Indonesian datasets, applying Indonesian-specific quality filtering, and performing cross-dataset deduplication. Subsets Subset Description quality_filtered Quality-filtered data before deduplication minhash_deduped Document-level MinHash deduplication matched Documents appearing in 2+ source datasets The matched subset uses… See the full description on the dataset page: https://huggingface.co/datasets/AdaMLLab/IndMix.texttext-generation100M<n<1B1 likes567 downloads5mo agoHugging Face19indonesian-nlp /mc4-idA thoroughly cleaned version of the Italian portion of the multilingual colossal, cleaned version of Common Crawl's web crawl corpus (mC4) by AllenAI. Based on Common Crawl dataset: "https://commoncrawl.org". This is the processed version of Google's mC4 dataset by AllenAI, with further cleaning detailed in the repository README file.texttext-generation1M<n<10M14 likes556 downloads4y agoHugging Face20open-index /ccrawl-recrawl-urls Common Crawl URL Recrawl Live refetches of pages from Common Crawl's URL index, with the body inline and the text already extracted What is it? Common Crawl publishes which URLs it saw and when, but the page bodies live in WARC archives that are awkward to query and are as old as the crawl that made them. This dataset takes the URL index for a single monthly crawl and fetches the pages again now, storing each response as one Parquet row with the body, the… See the full description on the dataset page: https://huggingface.co/datasets/open-index/ccrawl-recrawl-urls.tabulartext-generation1M<n<10M0 likes553 downloads1mo agoHugging Face21daruokta /t5gemma2-indonesia-instruct-v1 T5Gemma-2 Indonesian Instruct — Mono-Repo Satu repositori dataset HF untuk seluruh data pelatihan T5-Gemma-2 bahasa Indonesia. Diorganisasi per fungsi (fondasi → spesifik → preferensi) dengan folder/subfolder, setiap config = folder dan berisi split train + validation (80:20) di level percakapan. Struktur (by fungsi) t5gemma2-indonesia-instruct-v1/ ├── README.md ├── manifest.json ├── chat_idx_map.json ├── foundation/ ← FASE 1 · fondasi Bahasa… See the full description on the dataset page: https://huggingface.co/datasets/daruokta/t5gemma2-indonesia-instruct-v1.imagetext-generation100K<n<1M0 likes513 downloads2d agoHugging Face22open-index /open-library Open Library The complete Open Library catalog in clean, analysis-ready Parquet. 150.0M+ records across 11 entity types, from ISBNs and author bios to reading logs and Wikidata links. What is it? Open Library is a complete snapshot of the Open Library database, an open project of the Internet Archive with the mission of creating "one web page for every book ever published." The catalog is community-edited and contains bibliographic records for millions of authors, works… See the full description on the dataset page: https://huggingface.co/datasets/open-index/open-library.tabulartext-generation100M<n<1B9 likes490 downloads6mo agoHugging Face23kshitijthakkar /Nemotron-Personas-India Nemotron-Personas-India A compound AI approach to personas grounded in real-world distributions वास्तविक दुनिया के वितरण पर आधारित व्यक्तित्वों के लिए एक मिश्रित AI दृष्टिकोण Dataset Overview (डेटासेट अवलोकन) Nemotron-Personas-India is an open-source (CC BY 4.0) dataset of synthetically-generated personas. This dataset is grounded in real-world demographic, geographic and personality trait distributions in India to capture the diversity and… See the full description on the dataset page: https://huggingface.co/datasets/kshitijthakkar/Nemotron-Personas-India.imagetext-generation1M<n<10M0 likes461 downloads3mo agoHugging Face24LinguaLift /IndicMMLU-Pro IndicMMLU Dataset This dataset contains the following languages: punjabi hindi urdu telugu gujrati kannada tamil marathi bengali UPLOAD Cite our work. This dataset is also described in IndicMMLU-Pro: Benchmarking Indic Large Language Models on Multi-Task Language Understanding. @dataset{kj2024indicmmlupro, author = {Kj, Sankalp and Kumar, Ashutosh and Balaji, Laxmaan and Kotecha, Nikunj and Jain, Vinija and Chadha, Aman and Bhaduri, Sreyoshi}, title =… See the full description on the dataset page: https://huggingface.co/datasets/LinguaLift/IndicMMLU-Pro.tabulartext-generation100K<n<1M4 likes438 downloads2y agoHugging Face25NPCI /nemo-gym-indian-bankingThis is NPCI/nemo-gym-indian-banking — the dataset for the indian_banking resources server in NVIDIA NeMo Gym: 300 synthetic multi-turn Indian retail-banking customer-support tasks (250 train / 50 validation), the 197-customer synthetic bank database and the 59-article knowledge base the environment loads at startup. NeMo Gym Indian Banking Agent Tasks Tool-calling customer-service tasks for an Indian retail-banking assistant, in the NVIDIA NeMo Gym agent-input JSONL format.… See the full description on the dataset page: https://huggingface.co/datasets/NPCI/nemo-gym-indian-banking.texttext-generationn<1K3 likes422 downloads20d agoHugging Face26KanoonGPT /indian-legal-documents Indian Legal Documents Open Indian statutory and legal-document data for AI, search, and legal research. This dataset is part of the KanoonGPT Open Legal Data Initiative — an effort to make Indian legal data easier to access, structure, and build on for open-source research, legal tech, and production AI systems. KanoonGPT is building structured Indian legal datasets and data infrastructure for open-source, research, and enterprise AI applications. Learn more at kanoongpt.in.… See the full description on the dataset page: https://huggingface.co/datasets/KanoonGPT/indian-legal-documents.texttext-generation10K<n<100K0 likes389 downloads4mo agoHugging Face27Parssky /industrial-instruction-dataset Industrial-Instruction Dataset Industrial-Instruction provides benchmark and training-ready QA instances derived from industrial technical reports, designed to evaluate robustness under realistic retrieval conditions. Samples are grounded in retrieved evidence and include irrelevant retrieval, single-/multi-document support, and single-/multi-document answer settings. Paper Industrial-Instruction: An End-to-End Framework for Building Instruction-Tuning and… See the full description on the dataset page: https://huggingface.co/datasets/Parssky/industrial-instruction-dataset.tabularquestion-answering10K<n<100K0 likes378 downloads1mo agoHugging Face28overthelex /indian-court-decisions Indian Court Decisions A large-scale dataset of Indian court decisions with full text, metadata, and outcome labels covering the Supreme Court of India and 25 High Courts (1950–2026). Dataset Summary Config Train Validation Test Total high_courts 11,682,776 1,459,319 1,457,934 14,600,029 supreme_court 40,044 4,990 5,019 50,053 Total 14,650,082 This is one of the largest publicly available legal NLP datasets, containing over 14.6 million… See the full description on the dataset page: https://huggingface.co/datasets/overthelex/indian-court-decisions.tabulartext-classification10M<n<100M1 likes356 downloads4mo agoHugging Face29LeSouv /index-souverainete Index Souveraineté Le Souv Le dataset public de référence sur les entreprises françaises stratégiques cédées à des capitaux étrangers, et sur les entreprises souveraines à capitaux français. Source canonique : Le Souv — média indépendant consacré à la souveraineté économique et politique française. URL canonique : https://lesouv.fr/index-souverainete.json Licence : Creative Commons BY 4.0 (réutilisation libre avec attribution à Le Souv). Mise à jour : continue, regénération… See the full description on the dataset page: https://huggingface.co/datasets/LeSouv/index-souverainete.tabulartext-classificationn<1K1 likes339 downloads2d agoHugging Face30BAAI /IndustryCorpus_agriculture[中文主页] Industry models play a crucial role in driving enterprise intelligence transformation and innovative development. High-quality industry data is key to improving the performance of large models and realizing industry applications. However, datasets currently used for industry model training generally suffer from issues such as insufficient data volume, low quality, and lack of domain expertise. To address these problems, we constructed and applied 22 industry data processing operators to… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus_agriculture.texttext-generation1M<n<10M5 likes313 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.