CoolFace
10 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Podtech /llm-jp-corpus-v4-ja_sip_comprehensive_html llm-jp-corpus-v4 — ja_sip_comprehensive_html Mirror of the ja/ja_sip_comprehensive_html sub-corpus of LLM-jp Corpus v4, built by the LLM-jp Corpus Building WG (NII). Source: https://gitlab.llm-jp.nii.ac.jp/datasets/llm-jp-corpus-v4 Sub-corpus: ja_sip_comprehensive_html Files: 181 × jsonl.gz (23.4 GB compressed) Format: one JSON object per line, with a text key and a meta key (document id, URL, and other provenance fields). Directory layout mirrors the upstream repository.… See the full description on the dataset page: https://huggingface.co/datasets/Podtech/llm-jp-corpus-v4-ja_sip_comprehensive_html.texttext-generation1M<n<10M0 likes645 downloads2mo agoHugging Face02Podtech /llm-jp-corpus-v4-ja_sip_comprehensive_pdf llm-jp-corpus-v4 — ja_sip_comprehensive_pdf Mirror of the ja/ja_sip_comprehensive_pdf sub-corpus of LLM-jp Corpus v4, built by the LLM-jp Corpus Building WG (NII). Source: https://gitlab.llm-jp.nii.ac.jp/datasets/llm-jp-corpus-v4 Sub-corpus: ja_sip_comprehensive_pdf Files: 156 × jsonl.gz (39.1 GB compressed) Format: one JSON object per line, with a text key and a meta key (document id, URL, and other provenance fields). Directory layout mirrors the upstream repository.… See the full description on the dataset page: https://huggingface.co/datasets/Podtech/llm-jp-corpus-v4-ja_sip_comprehensive_pdf.texttext-generation1M<n<10M0 likes494 downloads2mo agoHugging Face033amthoughts /hsc-zoology-bangla-comprehensive-dataset 🧬 HSC Zoology Bangla Comprehensive Dataset A Diverse Multi-Chapter Academic Dataset This dataset contains 15,000 high-quality instruction-response pairs designed for Supervised Fine-Tuning (SFT). Unlike single-topic datasets, this collection spans several critical chapters of the HSC Zoology curriculum. 📚 Chapters Covered Human Physiology (মানুষের শারীরতত্ত্ব): Detailed Q&A on Digestion (পরিপাক) and Blood Circulation (রক্ত ও সঞ্চালন).… See the full description on the dataset page: https://huggingface.co/datasets/3amthoughts/hsc-zoology-bangla-comprehensive-dataset.textquestion-answering10K<n<100K1 likes54 downloads3mo agoHugging Face04Amvhunt /celestial-comprehensive-dataset-v2 CELESTIAL Comprehensive Spiritual AI Dataset v2.0 🌟 Overview The most comprehensive dataset for training spiritual AI assistants, featuring 9,000+ high-quality examples across all major spiritual and astrological domains. 📊 Dataset Statistics Total Examples: 9,000 Training Split: 7,200 examples Validation Split: 900 examples Test Split: 900 examples Categories: 4 categories Languages: English, Hindi (transliterated) 🎯 Categories Included… See the full description on the dataset page: https://huggingface.co/datasets/Amvhunt/celestial-comprehensive-dataset-v2.texttext-generation10K<n<100K0 likes50 downloads7mo agoHugging Face05ethanolivertroy /cmmc-training-comprehensive CMMC Training Dataset - Comprehensive Variant Dataset Description This is the Comprehensive variant of the CMMC (Cybersecurity Maturity Model Certification) training dataset, containing 11,279 high-quality training examples from the complete NIST CMMC publication library. Dataset Characteristics Total Examples: 11,279 (9,023 train / 2,256 validation) Source Documents: 381 NIST publications CMMC Levels Covered: Level 1, Level 2, Level 3 CMMC Domains: All 17… See the full description on the dataset page: https://huggingface.co/datasets/ethanolivertroy/cmmc-training-comprehensive.texttext-generation10K<n<100K0 likes45 downloads11mo agoHugging Face06pkchwy /turkish-comprehensive-movie-series-dataset Beyazperde Film & Series Dataset This dataset contains a comprehensive collection of Turkish films and TV series from Beyazperde.com, including detailed information about movies, series, cast, reviews, and ratings. Dataset Summary Total Movies: 27,227 Total Series: 11,240 Total Entries: 38,467 File Size: ~222 MB Format: JSONL (JSON Lines) Language: Turkish Source: Beyazperde.com Data Structure Each line in the JSONL file contains a JSON object… See the full description on the dataset page: https://huggingface.co/datasets/pkchwy/turkish-comprehensive-movie-series-dataset.imagetext-classification10K<n<100K4 likes22 downloads1y agoHugging Face07Adilbai /kz-rus-articles-comprehensive 🇰🇿🇷🇺 Kazakh-Russian Articles Comprehensive Dataset A high-quality bilingual corpus for cross-lingual NLP research 📋 Dataset Overview The Kazakh-Russian Articles Comprehensive Dataset is a meticulously curated bilingual corpus designed to advance natural language processing research for Kazakh and Russian languages. This dataset addresses the critical need for high-quality parallel and comparable text resources in Central Asian language pairs, particularly… See the full description on the dataset page: https://huggingface.co/datasets/Adilbai/kz-rus-articles-comprehensive.tabulartranslationn<1K1 likes18 downloads1y agoHugging Face08dp1812 /celestial-comprehensive-dataset-v2 CELESTIAL Comprehensive Spiritual AI Dataset v2.0 🌟 Overview The most comprehensive dataset for training spiritual AI assistants, featuring 9,000+ high-quality examples across all major spiritual and astrological domains. 📊 Dataset Statistics Total Examples: 9,000 Training Split: 7,200 examples Validation Split: 900 examples Test Split: 900 examples Categories: 4 categories Languages: English, Hindi (transliterated) 🎯 Categories Included… See the full description on the dataset page: https://huggingface.co/datasets/dp1812/celestial-comprehensive-dataset-v2.texttext-generation1K<n<10K0 likes15 downloads1y agoHugging Face09Nathan-Maine /cmmc-benchmark-v3-comprehensive-2026-q2gated CMMC Benchmark v3 Comprehensive — Q2 2026 Version: 2026-q2 Tier: v3 Comprehensive (1,273 questions, 15 evaluation dimensions) Purpose: The full, authoritative evaluation for compliance AI Valid through: June 30, 2026 Next release: July 1, 2026 (Q3 2026) License: CC-BY-4.0 Author: Nathan Maine What This Is This is the comprehensive tier of the CMMC Compliance Benchmark suite: 1,273 questions across 15 evaluation dimensions, covering the full scope of CMMC 2.0 /… See the full description on the dataset page: https://huggingface.co/datasets/Nathan-Maine/cmmc-benchmark-v3-comprehensive-2026-q2.textquestion-answering1K<n<10K0 likes5 downloads4mo agoHugging Face10memoriant /cmmc-benchmark-v3-comprehensive-2026-q2gated CMMC Benchmark v3 Comprehensive — Q2 2026 Version: 2026-q2 Tier: v3 Comprehensive (1,273 questions, 15 evaluation dimensions) Purpose: The authoritative evaluation for compliance AI Valid through: June 30, 2026 Next release: July 1, 2026 (Q3 2026) License: CC-BY-4.0 Publisher: Memoriant, Inc. What This Is The Memoriant Industrial Benchmark v3 — the comprehensive evaluation framework for compliance AI systems. 1,273 questions across 15 evaluation dimensions… See the full description on the dataset page: https://huggingface.co/datasets/memoriant/cmmc-benchmark-v3-comprehensive-2026-q2.textquestion-answering1K<n<10K0 likes4 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.