datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
comprehensive-arithmetic-problemscomprehensive-arithmetic-problems-carriesindian-stocks-comprehensive-fundamentals-dataset
Indian Stocks Comprehensive Fundamentals Dataset
From screener.in | 5701 Stocks | 375.20 MB+ Data | Weekly Updates
Highlights :
Total Number of stocks : 5701
Dataset Size : 375.20 MB
Status :
last updated on Friday, 25 Sep 2026 00:39:40 +0000
Usage Notes :
They are stored in 5701 individual files.
[stockname].json means all data related to that stock.
For example, titan.json contains all available fundamental data… See the full description on the dataset page: https://huggingface.co/datasets/AYUSHKHAIRE/indian-stocks-comprehensive-fundamentals-dataset.llm-jp-corpus-v4-ja_sip_comprehensive_html
llm-jp-corpus-v4 — ja_sip_comprehensive_html
Mirror of the ja/ja_sip_comprehensive_html sub-corpus of LLM-jp Corpus v4,
built by the LLM-jp Corpus Building WG (NII).
Source: https://gitlab.llm-jp.nii.ac.jp/datasets/llm-jp-corpus-v4
Sub-corpus: ja_sip_comprehensive_html
Files: 181 × jsonl.gz (23.4 GB compressed)
Format: one JSON object per line, with a text key and a meta key
(document id, URL, and other provenance fields).
Directory layout mirrors the upstream repository.… See the full description on the dataset page: https://huggingface.co/datasets/Podtech/llm-jp-corpus-v4-ja_sip_comprehensive_html.llm-jp-corpus-v4-ja_sip_comprehensive_pdf
llm-jp-corpus-v4 — ja_sip_comprehensive_pdf
Mirror of the ja/ja_sip_comprehensive_pdf sub-corpus of LLM-jp Corpus v4,
built by the LLM-jp Corpus Building WG (NII).
Source: https://gitlab.llm-jp.nii.ac.jp/datasets/llm-jp-corpus-v4
Sub-corpus: ja_sip_comprehensive_pdf
Files: 156 × jsonl.gz (39.1 GB compressed)
Format: one JSON object per line, with a text key and a meta key
(document id, URL, and other provenance fields).
Directory layout mirrors the upstream repository.… See the full description on the dataset page: https://huggingface.co/datasets/Podtech/llm-jp-corpus-v4-ja_sip_comprehensive_pdf.Comprehensive-Antiquarian-and-Rare-Books-Archive
Comprehensive Antiquarian & Rare Books Archive
Dataset Description
This dataset contains pristine, commerce-free bibliographical metadata extracted from the Govi Rare Books Archive. It is engineered to provide high-fidelity, structured historical data for Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) pipelines. By supplying ground-truth bibliographical metadata, this repository aims to reduce AI hallucinations and improve semantic reasoning… See the full description on the dataset page: https://huggingface.co/datasets/Govi-Rare-Books-Archive/Comprehensive-Antiquarian-and-Rare-Books-Archive.comprehensive-car-damage
Car Front and Rear Damage Detection Dataset
Dataset Summary
This dataset is designed for training and evaluating machine learning models for car damage detection, specifically focusing on front and rear vehicle damages.
It includes high-quality labeled images categorized into six distinct classes:
R_Normal: Rear view of undamaged cars
R_Crushed: Rear view of cars with crushed damage
R_Breakage: Rear view of cars with visible breakage
F_Normal: Front view of… See the full description on the dataset page: https://huggingface.co/datasets/DrBimmer/comprehensive-car-damage.Comprehensive-English-Premier-League-Match-Dataset
Comprehensive English Premier League Match Dataset (2000–2026)
A match-level dataset covering 26 English Premier League seasons, from 2000/2001 through 2025/2026, combining classic scoreline data with in-game statistics, Expected Goals (xG), end-of-season standings, managers, geography, historical club form, and head-to-head form — all in a single flat CSV, ready for machine learning and analysis.
📦 GitHub: RezaGooner/english-premier-league-match-dataset
📚 Zenodo… See the full description on the dataset page: https://huggingface.co/datasets/Rezagooner/Comprehensive-English-Premier-League-Match-Dataset.comprehensive-car-damage
Car Front and Rear Damage Detection Dataset
Dataset Summary
This dataset is designed for training and evaluating machine learning models for car damage detection, specifically focusing on front and rear vehicle damages.
It includes high-quality labeled images categorized into six distinct classes:
R_Normal: Rear view of undamaged cars
R_Crushed: Rear view of cars with crushed damage
R_Breakage: Rear view of cars with visible breakage
F_Normal: Front view of… See the full description on the dataset page: https://huggingface.co/datasets/SaiVaibhavS/comprehensive-car-damage.Arabic_Poem_Comprehensive_Dataset_APCDComprehensive-Exoplanet-Dataset
Comprehensive Exoplanet Dataset
This dataset was published on Kaggle by Samyakraj Bayar and mirrored here.
Download
The dataset is available as a ZIP archive: Comprehensive Exoplanet Dataset.zip
License
MIT
pitvqa-comprehensive-spatial
PitVQA Comprehensive Spatial Dataset
High-fidelity surgical spatial localization dataset for training vision-language models on pituitary surgery instrument and anatomy detection.
🔗 GitHub: https://github.com/matheus-rech/pit_project
🤖 Trained Model: mmrech/pitvqa-qwen2vl-spatial
📄 Original Dataset: UCL Research Data Repository
Dataset Description
This dataset contains 10,139 surgical frames with precise spatial annotations for instrument localization and anatomy… See the full description on the dataset page: https://huggingface.co/datasets/mmrech/pitvqa-comprehensive-spatial.comprehensive-qa-dataset
Comprehensive Question Answering Dataset
A large-scale, diverse collection of question answering datasets combined into a unified format for training and evaluating QA models. This dataset contains over 160,000 question-answer pairs from three popular QA benchmarks.
Dataset Summary
This comprehensive dataset combines three popular question answering datasets into a single, unified format:
SQuAD 2.0 (Stanford Question Answering Dataset) - Context passages from Wikipedia… See the full description on the dataset page: https://huggingface.co/datasets/Successmove/comprehensive-qa-dataset.celestial-comprehensive-spiritual-ai
🌟 CELESTIAL Comprehensive Spiritual AI Dataset
🚀 SPEED-OPTIMIZED TRAINING - 45-90 MINUTES!
Latest Update: Added speed-optimized training notebook that reduces training time from 21+ hours to 45-90 minutes (15-20x faster!)
📊 Dataset Overview
Comprehensive spiritual AI training dataset with 3000+ conversations covering all 50+ CELESTIAL spiritual systems including the newly integrated Sanjay Jumaani numerology method.
🎯 Key Features:
⚡… See the full description on the dataset page: https://huggingface.co/datasets/dp1812/celestial-comprehensive-spiritual-ai.reward-bench-mistral-7b-sft-beta-comprehensiverag-comprehensive-triplets
RAG Comprehensive Triplets Dataset
Dataset Description
This dataset, "rag-comprehensive-triplets", is a comprehensive collection of query-positive-negative triplets designed for training and evaluating Retrieval-Augmented Generation (RAG) models. It is derived from the "baconnier/RAG_sparse_dataset" and includes various query types paired with positive and negative responses.
Key Features:
Triplet Structure: Each entry consists of a query, a positive response… See the full description on the dataset page: https://huggingface.co/datasets/baconnier/rag-comprehensive-triplets.hsc-zoology-bangla-comprehensive-dataset
🧬 HSC Zoology Bangla Comprehensive Dataset
A Diverse Multi-Chapter Academic Dataset
This dataset contains 15,000 high-quality instruction-response pairs designed for Supervised Fine-Tuning (SFT). Unlike single-topic datasets, this collection spans several critical chapters of the HSC Zoology curriculum.
📚 Chapters Covered
Human Physiology (মানুষের শারীরতত্ত্ব): Detailed Q&A on Digestion (পরিপাক) and Blood Circulation (রক্ত ও সঞ্চালন).… See the full description on the dataset page: https://huggingface.co/datasets/3amthoughts/hsc-zoology-bangla-comprehensive-dataset.bazi_comprehensive_dataset
AstroAlchemy BaZi Dataset Documentation
Overview
This documentation describes the comprehensive BaZi dataset created for the AstroAlchemy Web3 dApp project. The dataset is designed for fine-tuning a Mistral B instruct model to generate hyper-personalized, BaZi-powered "spiritual strategies" across multiple domains.
Dataset Structure
The dataset is provided in JSONL (JSON Lines) format, with each line containing a complete JSON object with two fields:
input: A… See the full description on the dataset page: https://huggingface.co/datasets/viveriveniversumvivusvici/bazi_comprehensive_dataset.celestial-comprehensive-dataset-v2
CELESTIAL Comprehensive Spiritual AI Dataset v2.0
🌟 Overview
The most comprehensive dataset for training spiritual AI assistants, featuring 9,000+ high-quality examples across all major spiritual and astrological domains.
📊 Dataset Statistics
Total Examples: 9,000
Training Split: 7,200 examples
Validation Split: 900 examples
Test Split: 900 examples
Categories: 4 categories
Languages: English, Hindi (transliterated)
🎯 Categories Included… See the full description on the dataset page: https://huggingface.co/datasets/Amvhunt/celestial-comprehensive-dataset-v2.cmmc-training-comprehensive
CMMC Training Dataset - Comprehensive Variant
Dataset Description
This is the Comprehensive variant of the CMMC (Cybersecurity Maturity Model Certification) training dataset, containing 11,279 high-quality training examples from the complete NIST CMMC publication library.
Dataset Characteristics
Total Examples: 11,279 (9,023 train / 2,256 validation)
Source Documents: 381 NIST publications
CMMC Levels Covered: Level 1, Level 2, Level 3
CMMC Domains: All 17… See the full description on the dataset page: https://huggingface.co/datasets/ethanolivertroy/cmmc-training-comprehensive.Comprehensive_Feature_Extraction_DDoS_Datasets
Dataset Card for Comprehensive_Feature_Extraction_DDoS_Datasets
This dataset card aims to be provided preprocessed five published DDoS datasets based on three feature extracted methods, including correlation, IM, and UFS.
Dataset Description
The imperative for robust detection mechanisms has grown in the face of increasingly sophisticated Distributed Denial of Service (DDoS) attacks. This paper introduces DDoSBERT, an innovative approach harnessing transformer text… See the full description on the dataset page: https://huggingface.co/datasets/Thi-Thu-Huong/Comprehensive_Feature_Extraction_DDoS_Datasets.comprehensive-hair-37799e
comprehensive-hair-37799e
Synthetic sensors test data: 47 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/Velvet-Bito/comprehensive-hair-37799e.adamvakar_apple-comprehensive-financial-dataset-1980-2026
Apple Financial Dataset (1980-2026)
All-in-One Apple Stock Dataset: Prices, Financials, and ML Features (1980-2026)
Dataset Info
Source: Kaggle
Original Size: 8.11 MB
Kaggle Downloads: 1,540
Files: 4
Files
aapl_master_enriched.csv
aapl_quarterly_master.csv
aapl_quarterly_summary.csv
aapl_stock_ml_features.csv
Mirrored from Kaggle
comprehensive-quality-5e384e
comprehensive-quality-5e384e
Synthetic products test data: 36 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting… See the full description on the dataset page: https://huggingface.co/datasets/boyerdanielle/comprehensive-quality-5e384e.AyAI-Turk-Comprehensive-Corpuscomprehensive-healthbench-v2Comprehensive_VQA_MMEEvery item has 2 T/F questions (one True and one False) - only if these two are both correct, acc_score += 1.
Here is the function that you can use:
from sklearn.metrics import accuracy_score, precision_score, recall_score, confusion_matrix
def process_result(preds, gts):
"""
This func. is working for only one task (items with same labels)
"""
cnt = 0
acc_plus_correct_num = 0
for pred, gt in zip(preds, gts):
if pred == gt:
cnt += 1
if cnt == 2:… See the full description on the dataset page: https://huggingface.co/datasets/Holmes377/Comprehensive_VQA_MME.Heart-Disease-Dataset-_Comprehensiveturkish-comprehensive-movie-series-dataset
Beyazperde Film & Series Dataset
This dataset contains a comprehensive collection of Turkish films and TV series from Beyazperde.com, including detailed information about movies, series, cast, reviews, and ratings.
Dataset Summary
Total Movies: 27,227
Total Series: 11,240
Total Entries: 38,467
File Size: ~222 MB
Format: JSONL (JSON Lines)
Language: Turkish
Source: Beyazperde.com
Data Structure
Each line in the JSONL file contains a JSON object… See the full description on the dataset page: https://huggingface.co/datasets/pkchwy/turkish-comprehensive-movie-series-dataset.comprehensive-gradio-coding-dataset
