datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
llm-jp-corpus-v4-ja_sip_comprehensive_html
llm-jp-corpus-v4 — ja_sip_comprehensive_html
Mirror of the ja/ja_sip_comprehensive_html sub-corpus of LLM-jp Corpus v4,
built by the LLM-jp Corpus Building WG (NII).
Source: https://gitlab.llm-jp.nii.ac.jp/datasets/llm-jp-corpus-v4
Sub-corpus: ja_sip_comprehensive_html
Files: 181 × jsonl.gz (23.4 GB compressed)
Format: one JSON object per line, with a text key and a meta key
(document id, URL, and other provenance fields).
Directory layout mirrors the upstream repository.… See the full description on the dataset page: https://huggingface.co/datasets/Podtech/llm-jp-corpus-v4-ja_sip_comprehensive_html.llm-jp-corpus-v4-ja_sip_comprehensive_pdf
llm-jp-corpus-v4 — ja_sip_comprehensive_pdf
Mirror of the ja/ja_sip_comprehensive_pdf sub-corpus of LLM-jp Corpus v4,
built by the LLM-jp Corpus Building WG (NII).
Source: https://gitlab.llm-jp.nii.ac.jp/datasets/llm-jp-corpus-v4
Sub-corpus: ja_sip_comprehensive_pdf
Files: 156 × jsonl.gz (39.1 GB compressed)
Format: one JSON object per line, with a text key and a meta key
(document id, URL, and other provenance fields).
Directory layout mirrors the upstream repository.… See the full description on the dataset page: https://huggingface.co/datasets/Podtech/llm-jp-corpus-v4-ja_sip_comprehensive_pdf.hsc-zoology-bangla-comprehensive-dataset
🧬 HSC Zoology Bangla Comprehensive Dataset
A Diverse Multi-Chapter Academic Dataset
This dataset contains 15,000 high-quality instruction-response pairs designed for Supervised Fine-Tuning (SFT). Unlike single-topic datasets, this collection spans several critical chapters of the HSC Zoology curriculum.
📚 Chapters Covered
Human Physiology (মানুষের শারীরতত্ত্ব): Detailed Q&A on Digestion (পরিপাক) and Blood Circulation (রক্ত ও সঞ্চালন).… See the full description on the dataset page: https://huggingface.co/datasets/3amthoughts/hsc-zoology-bangla-comprehensive-dataset.turkish-comprehensive-movie-series-dataset
Beyazperde Film & Series Dataset
This dataset contains a comprehensive collection of Turkish films and TV series from Beyazperde.com, including detailed information about movies, series, cast, reviews, and ratings.
Dataset Summary
Total Movies: 27,227
Total Series: 11,240
Total Entries: 38,467
File Size: ~222 MB
Format: JSONL (JSON Lines)
Language: Turkish
Source: Beyazperde.com
Data Structure
Each line in the JSONL file contains a JSON object… See the full description on the dataset page: https://huggingface.co/datasets/pkchwy/turkish-comprehensive-movie-series-dataset.cmmc-benchmark-v3-comprehensive-2026-q2
CMMC Benchmark v3 Comprehensive — Q2 2026
Version: 2026-q2
Tier: v3 Comprehensive (1,273 questions, 15 evaluation dimensions)
Purpose: The full, authoritative evaluation for compliance AI
Valid through: June 30, 2026
Next release: July 1, 2026 (Q3 2026)
License: CC-BY-4.0
Author: Nathan Maine
What This Is
This is the comprehensive tier of the CMMC Compliance Benchmark suite: 1,273 questions across 15 evaluation dimensions, covering the full scope of CMMC 2.0 /… See the full description on the dataset page: https://huggingface.co/datasets/Nathan-Maine/cmmc-benchmark-v3-comprehensive-2026-q2.cmmc-benchmark-v3-comprehensive-2026-q2
CMMC Benchmark v3 Comprehensive — Q2 2026
Version: 2026-q2
Tier: v3 Comprehensive (1,273 questions, 15 evaluation dimensions)
Purpose: The authoritative evaluation for compliance AI
Valid through: June 30, 2026
Next release: July 1, 2026 (Q3 2026)
License: CC-BY-4.0
Publisher: Memoriant, Inc.
What This Is
The Memoriant Industrial Benchmark v3 — the comprehensive evaluation framework for compliance AI systems. 1,273 questions across 15 evaluation dimensions… See the full description on the dataset page: https://huggingface.co/datasets/memoriant/cmmc-benchmark-v3-comprehensive-2026-q2.
