CoolFace
14 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Podtech /llm-jp-corpus-v4-ja_sip_comprehensive_html llm-jp-corpus-v4 — ja_sip_comprehensive_html Mirror of the ja/ja_sip_comprehensive_html sub-corpus of LLM-jp Corpus v4, built by the LLM-jp Corpus Building WG (NII). Source: https://gitlab.llm-jp.nii.ac.jp/datasets/llm-jp-corpus-v4 Sub-corpus: ja_sip_comprehensive_html Files: 181 × jsonl.gz (23.4 GB compressed) Format: one JSON object per line, with a text key and a meta key (document id, URL, and other provenance fields). Directory layout mirrors the upstream repository.… See the full description on the dataset page: https://huggingface.co/datasets/Podtech/llm-jp-corpus-v4-ja_sip_comprehensive_html.texttext-generation1M<n<10M0 likes645 downloads2mo agoHugging Face02Podtech /llm-jp-corpus-v4-ja_sip_comprehensive_pdf llm-jp-corpus-v4 — ja_sip_comprehensive_pdf Mirror of the ja/ja_sip_comprehensive_pdf sub-corpus of LLM-jp Corpus v4, built by the LLM-jp Corpus Building WG (NII). Source: https://gitlab.llm-jp.nii.ac.jp/datasets/llm-jp-corpus-v4 Sub-corpus: ja_sip_comprehensive_pdf Files: 156 × jsonl.gz (39.1 GB compressed) Format: one JSON object per line, with a text key and a meta key (document id, URL, and other provenance fields). Directory layout mirrors the upstream repository.… See the full description on the dataset page: https://huggingface.co/datasets/Podtech/llm-jp-corpus-v4-ja_sip_comprehensive_pdf.texttext-generation1M<n<10M0 likes494 downloads2mo agoHugging Face03Govi-Rare-Books-Archive /Comprehensive-Antiquarian-and-Rare-Books-Archive Comprehensive Antiquarian & Rare Books Archive Dataset Description This dataset contains pristine, commerce-free bibliographical metadata extracted from the Govi Rare Books Archive. It is engineered to provide high-fidelity, structured historical data for Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) pipelines. By supplying ground-truth bibliographical metadata, this repository aims to reduce AI hallucinations and improve semantic reasoning… See the full description on the dataset page: https://huggingface.co/datasets/Govi-Rare-Books-Archive/Comprehensive-Antiquarian-and-Rare-Books-Archive.textn<1K0 likes231 downloads3h agoHugging Face04mmrech /pitvqa-comprehensive-spatial PitVQA Comprehensive Spatial Dataset High-fidelity surgical spatial localization dataset for training vision-language models on pituitary surgery instrument and anatomy detection. 🔗 GitHub: https://github.com/matheus-rech/pit_project 🤖 Trained Model: mmrech/pitvqa-qwen2vl-spatial 📄 Original Dataset: UCL Research Data Repository Dataset Description This dataset contains 10,139 surgical frames with precise spatial annotations for instrument localization and anatomy… See the full description on the dataset page: https://huggingface.co/datasets/mmrech/pitvqa-comprehensive-spatial.tabularvisual-question-answering10K<n<100K1 likes84 downloads9mo agoHugging Face053amthoughts /hsc-zoology-bangla-comprehensive-dataset 🧬 HSC Zoology Bangla Comprehensive Dataset A Diverse Multi-Chapter Academic Dataset This dataset contains 15,000 high-quality instruction-response pairs designed for Supervised Fine-Tuning (SFT). Unlike single-topic datasets, this collection spans several critical chapters of the HSC Zoology curriculum. 📚 Chapters Covered Human Physiology (মানুষের শারীরতত্ত্ব): Detailed Q&A on Digestion (পরিপাক) and Blood Circulation (রক্ত ও সঞ্চালন).… See the full description on the dataset page: https://huggingface.co/datasets/3amthoughts/hsc-zoology-bangla-comprehensive-dataset.textquestion-answering10K<n<100K1 likes54 downloads3mo agoHugging Face06viveriveniversumvivusvici /bazi_comprehensive_dataset AstroAlchemy BaZi Dataset Documentation Overview This documentation describes the comprehensive BaZi dataset created for the AstroAlchemy Web3 dApp project. The dataset is designed for fine-tuning a Mistral B instruct model to generate hyper-personalized, BaZi-powered "spiritual strategies" across multiple domains. Dataset Structure The dataset is provided in JSONL (JSON Lines) format, with each line containing a complete JSON object with two fields: input: A… See the full description on the dataset page: https://huggingface.co/datasets/viveriveniversumvivusvici/bazi_comprehensive_dataset.text1K<n<10K2 likes51 downloads1y agoHugging Face07pkchwy /turkish-comprehensive-movie-series-dataset Beyazperde Film & Series Dataset This dataset contains a comprehensive collection of Turkish films and TV series from Beyazperde.com, including detailed information about movies, series, cast, reviews, and ratings. Dataset Summary Total Movies: 27,227 Total Series: 11,240 Total Entries: 38,467 File Size: ~222 MB Format: JSONL (JSON Lines) Language: Turkish Source: Beyazperde.com Data Structure Each line in the JSONL file contains a JSON object… See the full description on the dataset page: https://huggingface.co/datasets/pkchwy/turkish-comprehensive-movie-series-dataset.imagetext-classification10K<n<100K4 likes22 downloads1y agoHugging Face08iengpho /suno-model-comprehensive-trainingtextn<1K0 likes10 downloads3mo agoHugging Face09CristiD7 /Comprehensive_7Day_Workout_Plans_100textn<1K1 likes7 downloads2y agoHugging Face10Nathan-Maine /cmmc-benchmark-v3-comprehensive-2026-q2gated CMMC Benchmark v3 Comprehensive — Q2 2026 Version: 2026-q2 Tier: v3 Comprehensive (1,273 questions, 15 evaluation dimensions) Purpose: The full, authoritative evaluation for compliance AI Valid through: June 30, 2026 Next release: July 1, 2026 (Q3 2026) License: CC-BY-4.0 Author: Nathan Maine What This Is This is the comprehensive tier of the CMMC Compliance Benchmark suite: 1,273 questions across 15 evaluation dimensions, covering the full scope of CMMC 2.0 /… See the full description on the dataset page: https://huggingface.co/datasets/Nathan-Maine/cmmc-benchmark-v3-comprehensive-2026-q2.textquestion-answering1K<n<10K0 likes5 downloads4mo agoHugging Face11garrykuwanto /fireball_comprehensive_pairstext1M<n<10M0 likes4 downloads1y agoHugging Face12memoriant /cmmc-benchmark-v3-comprehensive-2026-q2gated CMMC Benchmark v3 Comprehensive — Q2 2026 Version: 2026-q2 Tier: v3 Comprehensive (1,273 questions, 15 evaluation dimensions) Purpose: The authoritative evaluation for compliance AI Valid through: June 30, 2026 Next release: July 1, 2026 (Q3 2026) License: CC-BY-4.0 Publisher: Memoriant, Inc. What This Is The Memoriant Industrial Benchmark v3 — the comprehensive evaluation framework for compliance AI systems. 1,273 questions across 15 evaluation dimensions… See the full description on the dataset page: https://huggingface.co/datasets/memoriant/cmmc-benchmark-v3-comprehensive-2026-q2.textquestion-answering1K<n<10K0 likes4 downloads6mo agoHugging Face13agrpankaj /comprehensive_COT_dataset_finaltextn<1K0 likes3 downloads2y agoHugging Face14CristiD7 /Comprehensive_7Day_Workout_Planstextn<1K2 likes1 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.