datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
comprehensive-arithmetic-problemscomprehensive-arithmetic-problems-carriesllm-jp-corpus-v4-ja_sip_comprehensive_html
llm-jp-corpus-v4 — ja_sip_comprehensive_html
Mirror of the ja/ja_sip_comprehensive_html sub-corpus of LLM-jp Corpus v4,
built by the LLM-jp Corpus Building WG (NII).
Source: https://gitlab.llm-jp.nii.ac.jp/datasets/llm-jp-corpus-v4
Sub-corpus: ja_sip_comprehensive_html
Files: 181 × jsonl.gz (23.4 GB compressed)
Format: one JSON object per line, with a text key and a meta key
(document id, URL, and other provenance fields).
Directory layout mirrors the upstream repository.… See the full description on the dataset page: https://huggingface.co/datasets/Podtech/llm-jp-corpus-v4-ja_sip_comprehensive_html.llm-jp-corpus-v4-ja_sip_comprehensive_pdf
llm-jp-corpus-v4 — ja_sip_comprehensive_pdf
Mirror of the ja/ja_sip_comprehensive_pdf sub-corpus of LLM-jp Corpus v4,
built by the LLM-jp Corpus Building WG (NII).
Source: https://gitlab.llm-jp.nii.ac.jp/datasets/llm-jp-corpus-v4
Sub-corpus: ja_sip_comprehensive_pdf
Files: 156 × jsonl.gz (39.1 GB compressed)
Format: one JSON object per line, with a text key and a meta key
(document id, URL, and other provenance fields).
Directory layout mirrors the upstream repository.… See the full description on the dataset page: https://huggingface.co/datasets/Podtech/llm-jp-corpus-v4-ja_sip_comprehensive_pdf.Comprehensive-Antiquarian-and-Rare-Books-Archive
Comprehensive Antiquarian & Rare Books Archive
Dataset Description
This dataset contains pristine, commerce-free bibliographical metadata extracted from the Govi Rare Books Archive. It is engineered to provide high-fidelity, structured historical data for Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) pipelines. By supplying ground-truth bibliographical metadata, this repository aims to reduce AI hallucinations and improve semantic reasoning… See the full description on the dataset page: https://huggingface.co/datasets/Govi-Rare-Books-Archive/Comprehensive-Antiquarian-and-Rare-Books-Archive.Comprehensive-English-Premier-League-Match-Dataset
Comprehensive English Premier League Match Dataset (2000–2026)
A match-level dataset covering 26 English Premier League seasons, from 2000/2001 through 2025/2026, combining classic scoreline data with in-game statistics, Expected Goals (xG), end-of-season standings, managers, geography, historical club form, and head-to-head form — all in a single flat CSV, ready for machine learning and analysis.
📦 GitHub: RezaGooner/english-premier-league-match-dataset
📚 Zenodo… See the full description on the dataset page: https://huggingface.co/datasets/Rezagooner/Comprehensive-English-Premier-League-Match-Dataset.Arabic_Poem_Comprehensive_Dataset_APCDpitvqa-comprehensive-spatial
PitVQA Comprehensive Spatial Dataset
High-fidelity surgical spatial localization dataset for training vision-language models on pituitary surgery instrument and anatomy detection.
🔗 GitHub: https://github.com/matheus-rech/pit_project
🤖 Trained Model: mmrech/pitvqa-qwen2vl-spatial
📄 Original Dataset: UCL Research Data Repository
Dataset Description
This dataset contains 10,139 surgical frames with precise spatial annotations for instrument localization and anatomy… See the full description on the dataset page: https://huggingface.co/datasets/mmrech/pitvqa-comprehensive-spatial.reward-bench-mistral-7b-sft-beta-comprehensiverag-comprehensive-triplets
RAG Comprehensive Triplets Dataset
Dataset Description
This dataset, "rag-comprehensive-triplets", is a comprehensive collection of query-positive-negative triplets designed for training and evaluating Retrieval-Augmented Generation (RAG) models. It is derived from the "baconnier/RAG_sparse_dataset" and includes various query types paired with positive and negative responses.
Key Features:
Triplet Structure: Each entry consists of a query, a positive response… See the full description on the dataset page: https://huggingface.co/datasets/baconnier/rag-comprehensive-triplets.hsc-zoology-bangla-comprehensive-dataset
🧬 HSC Zoology Bangla Comprehensive Dataset
A Diverse Multi-Chapter Academic Dataset
This dataset contains 15,000 high-quality instruction-response pairs designed for Supervised Fine-Tuning (SFT). Unlike single-topic datasets, this collection spans several critical chapters of the HSC Zoology curriculum.
📚 Chapters Covered
Human Physiology (মানুষের শারীরতত্ত্ব): Detailed Q&A on Digestion (পরিপাক) and Blood Circulation (রক্ত ও সঞ্চালন).… See the full description on the dataset page: https://huggingface.co/datasets/3amthoughts/hsc-zoology-bangla-comprehensive-dataset.bazi_comprehensive_dataset
AstroAlchemy BaZi Dataset Documentation
Overview
This documentation describes the comprehensive BaZi dataset created for the AstroAlchemy Web3 dApp project. The dataset is designed for fine-tuning a Mistral B instruct model to generate hyper-personalized, BaZi-powered "spiritual strategies" across multiple domains.
Dataset Structure
The dataset is provided in JSONL (JSON Lines) format, with each line containing a complete JSON object with two fields:
input: A… See the full description on the dataset page: https://huggingface.co/datasets/viveriveniversumvivusvici/bazi_comprehensive_dataset.celestial-comprehensive-dataset-v2
CELESTIAL Comprehensive Spiritual AI Dataset v2.0
🌟 Overview
The most comprehensive dataset for training spiritual AI assistants, featuring 9,000+ high-quality examples across all major spiritual and astrological domains.
📊 Dataset Statistics
Total Examples: 9,000
Training Split: 7,200 examples
Validation Split: 900 examples
Test Split: 900 examples
Categories: 4 categories
Languages: English, Hindi (transliterated)
🎯 Categories Included… See the full description on the dataset page: https://huggingface.co/datasets/Amvhunt/celestial-comprehensive-dataset-v2.cmmc-training-comprehensive
CMMC Training Dataset - Comprehensive Variant
Dataset Description
This is the Comprehensive variant of the CMMC (Cybersecurity Maturity Model Certification) training dataset, containing 11,279 high-quality training examples from the complete NIST CMMC publication library.
Dataset Characteristics
Total Examples: 11,279 (9,023 train / 2,256 validation)
Source Documents: 381 NIST publications
CMMC Levels Covered: Level 1, Level 2, Level 3
CMMC Domains: All 17… See the full description on the dataset page: https://huggingface.co/datasets/ethanolivertroy/cmmc-training-comprehensive.comprehensive-hair-37799e
comprehensive-hair-37799e
Synthetic sensors test data: 47 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/Velvet-Bito/comprehensive-hair-37799e.comprehensive-quality-5e384e
comprehensive-quality-5e384e
Synthetic products test data: 36 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting… See the full description on the dataset page: https://huggingface.co/datasets/boyerdanielle/comprehensive-quality-5e384e.comprehensive-healthbench-v2Comprehensive_VQA_MMEEvery item has 2 T/F questions (one True and one False) - only if these two are both correct, acc_score += 1.
Here is the function that you can use:
from sklearn.metrics import accuracy_score, precision_score, recall_score, confusion_matrix
def process_result(preds, gts):
"""
This func. is working for only one task (items with same labels)
"""
cnt = 0
acc_plus_correct_num = 0
for pred, gt in zip(preds, gts):
if pred == gt:
cnt += 1
if cnt == 2:… See the full description on the dataset page: https://huggingface.co/datasets/Holmes377/Comprehensive_VQA_MME.turkish-comprehensive-movie-series-dataset
Beyazperde Film & Series Dataset
This dataset contains a comprehensive collection of Turkish films and TV series from Beyazperde.com, including detailed information about movies, series, cast, reviews, and ratings.
Dataset Summary
Total Movies: 27,227
Total Series: 11,240
Total Entries: 38,467
File Size: ~222 MB
Format: JSONL (JSON Lines)
Language: Turkish
Source: Beyazperde.com
Data Structure
Each line in the JSONL file contains a JSON object… See the full description on the dataset page: https://huggingface.co/datasets/pkchwy/turkish-comprehensive-movie-series-dataset.comprehensive-gradio-coding-datasetswe-comprehensive-datasetkz-rus-articles-comprehensive
🇰🇿🇷🇺 Kazakh-Russian Articles Comprehensive Dataset
A high-quality bilingual corpus for cross-lingual NLP research
📋 Dataset Overview
The Kazakh-Russian Articles Comprehensive Dataset is a meticulously curated bilingual corpus designed to advance natural language processing research for Kazakh and Russian languages. This dataset addresses the critical need for high-quality parallel and comparable text resources in Central Asian language pairs, particularly… See the full description on the dataset page: https://huggingface.co/datasets/Adilbai/kz-rus-articles-comprehensive.europe-owid-comprehensive-nuclear-test-ban-treaty
Comprehensive Nuclear Test Ban Treaty | Europe (Our World in Data)
🇪🇺 1,320 observations · 44 Europe countries · 1996–2025 · Repackaged by Electric Sheep Europe
TL;DR
This dataset contains 1,320 observations of Comprehensive Nuclear Test Ban Treaty data across 44 Europe countries, spanning 1996–2025.
About the source
Source: Our World in Data
Publisher: Our World in Data
License: cc-by-4.0
Topic: Comprehensive Nuclear Test Ban Treaty… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepeurope/europe-owid-comprehensive-nuclear-test-ban-treaty.mew1a-v4-pokemon-tcg-comprehensivecelestial-comprehensive-dataset-v2
CELESTIAL Comprehensive Spiritual AI Dataset v2.0
🌟 Overview
The most comprehensive dataset for training spiritual AI assistants, featuring 9,000+ high-quality examples across all major spiritual and astrological domains.
📊 Dataset Statistics
Total Examples: 9,000
Training Split: 7,200 examples
Validation Split: 900 examples
Test Split: 900 examples
Categories: 4 categories
Languages: English, Hindi (transliterated)
🎯 Categories Included… See the full description on the dataset page: https://huggingface.co/datasets/dp1812/celestial-comprehensive-dataset-v2.ramanv-image-real-comprehensiveppo-ltr-comprehensive-ranking-datasetafrica-worldbank-wbl-supportive-framework-entrepreneurship-there-is-a-comprehensive-framework-to
WBL: Supportive Framework, Entrepreneurship, There is a comprehensive framework to support women entrepreneurs, women-owned or women-led businesses | Africa (World Bank — Gender Statistics) | Africa (World Bank)
Size category: n<1K - Formats: parquet - Sector: demographics_social - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-worldbank-wbl-supportive-framework-entrepreneurship-there-is-a-comprehensive-framework-to.Comprehensive-Medical-QA-Datasetcomprehensive-765-annotated
