CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01lingshu-medical-mllm /lingshu_training_data_medical_domain Website &nbsp;&nbsp; 🤖 7B Model &nbsp;&nbsp; 🤖 8B Model based on InternVL3 &nbsp;&nbsp; 🤖 32B Model &nbsp;&nbsp; MedEvalKit &nbsp;&nbsp; Technical Report &nbsp;&nbsp; Lingshu MCP Lingshu Medical MLLM Training Data (Medical Domain) This dataset contains the medical-domain training data used in the multi-stage training of the Lingshu Medical Multimodal Large Language Model (MLLM). General-domain data has been removed; only medical data is included. The training… See the full description on the dataset page: https://huggingface.co/datasets/lingshu-medical-mllm/lingshu_training_data_medical_domain.textimage-to-text100M<n<1B8 likes3.5k downloads25d agoHugging Face02SciCodePile /SciCode-Domain-Code DATA1: Domain-Specific Code Dataset Dataset Overview DATA1 is a large-scale domain-specific code dataset focusing on code samples from interdisciplinary fields such as biology, chemistry, materials science, and related areas. The dataset is collected and organized from GitHub repositories, covering 178 different domain topics with over 1.1 billion lines of code. Dataset Statistics Total Datasets: 178 CSV files Total Data Size: ~115 GB Total Lines of Code: Over… See the full description on the dataset page: https://huggingface.co/datasets/SciCodePile/SciCode-Domain-Code.tabulartext-generation1M<n<10M4 likes1.9k downloads7mo agoHugging Face03open-index /ccrawl-recrawl-domains Common Crawl Domain Recrawl Live fetches of the home page of every ranked domain in Common Crawl's web graph, rendered to Markdown as they are fetched What is it? Common Crawl's web graph ranks domains by how central they are, but it does not tell you what those domains actually serve today. This dataset walks that ranking from the top and fetches each domain's home page now, storing the response as one Parquet row with the body, the headers, the timing and the… See the full description on the dataset page: https://huggingface.co/datasets/open-index/ccrawl-recrawl-domains.tabulartext-generation1M<n<10M0 likes1.2k downloads1mo agoHugging Face04AdaMLLab /AraMix-domain-classified AraMix Domain-Classified AraMix family: AraMix (minhash and matched) | AraMix-domain-classified (with domain labels) | AraMix-HQ (model-filtered) This is AraMix with per-document domain labels from nvidia/multilingual-domain-classifier. Usage from datasets import load_dataset ds = load_dataset("AdaMLLab/AraMix-domain-classified", "minhash_deduped") ds = load_dataset("AdaMLLab/AraMix-domain-classified", "sentence_deduped") Schema Field… See the full description on the dataset page: https://huggingface.co/datasets/AdaMLLab/AraMix-domain-classified.texttext-generation100M<n<1B1 likes1.1k downloads8mo agoHugging Face05DanFosing /public-domain-poetry Overview This dataset is a collection of approximately 38,500 poems from https://www.public-domain-poetry.com/. Language The language of this dataset is English. License All data in this dataset is public domain, which means you should be able to use it for anything you want, as long as you aren't breaking any law in the process of doing so. texttext-generation10K<n<100K21 likes485 downloads3y agoHugging Face06liuhangbiao /SciCode-Domain-Code DATA1: Domain-Specific Code Dataset Dataset Overview DATA1 is a large-scale domain-specific code dataset focusing on code samples from interdisciplinary fields such as biology, chemistry, materials science, and related areas. The dataset is collected and organized from GitHub repositories, covering 178 different domain topics with over 1.1 billion lines of code. Dataset Statistics Total Datasets: 178 CSV files Total Data Size: ~115 GB Total Lines of Code: Over… See the full description on the dataset page: https://huggingface.co/datasets/liuhangbiao/SciCode-Domain-Code.tabulartext-generation1M<n<10M0 likes383 downloads6mo agoHugging Face07common-pile /public_domain_review_filtered Public Domain Review Description The Public Domain Review is an online journal dedicated to exploration of works of art and literature that have aged into the public domain. We collect all articles published in the Public Domain Review under a CC BY-SA license. Dataset Statistics Documents UTF-8 GB 1,406 0.007 License Issues While we aim to produce datasets with completely accurate licensing information, license laundering and… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/public_domain_review_filtered.texttext-generation1K<n<10K0 likes196 downloads1y agoHugging Face08botcoinmoney /domain-agnostic-causal-reasoning-tuning Domain-Agnostic Causal Reasoning Tuning Dataset Training data for fine-tuning language models on multi-hop document reasoning. Each example is a graded reasoning trace produced by a frontier AI agent solving a procedurally generated challenge from the Botcoin proof-of-inference network. The traces contain no real domain knowledge. Entities are fictional, numbers are random, and documents are generated deterministically from 128-bit seeds. The reasoning structure is what matters:… See the full description on the dataset page: https://huggingface.co/datasets/botcoinmoney/domain-agnostic-causal-reasoning-tuning.textquestion-answering10K<n<100K1 likes169 downloads6mo agoHugging Face09amazon /Turnstile-Synthetic-Domains Data Turnstile — Synthetic Domains A large-scale synthetic dataset of function-calling interactions with chain-of-thought reasoning traces, designed for training small language models on tool-use tasks. Dataset Summary Metric Value Interactions 100,262 Unique APIs 1,025 Distractors per interaction 5 Template types 17 Avg roles per interaction ~10 Avg tokens per interaction ~972 Language English Generator model Qwen2.5-32B-Instruct… See the full description on the dataset page: https://huggingface.co/datasets/amazon/Turnstile-Synthetic-Domains.texttext-generation100K<n<1M0 likes149 downloads2mo agoHugging Face10common-pile /public_domain_review Public Domain Review Description The Public Domain Review is an online journal dedicated to exploration of works of art and literature that have aged into the public domain. We collect all articles published in the Public Domain Review under a CC BY-SA license. Dataset Statistics Documents UTF-8 GB 1,412 0.007 License Issues While we aim to produce datasets with completely accurate licensing information, license laundering and… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/public_domain_review.texttext-generation1K<n<10K1 likes139 downloads1y agoHugging Face11dendriteholdings /Dendrite-Synth-Multi-Domain Dendrite Synth Multi-Domain A verified, difficulty-filtered, style-amplified synthetic corpus of question / reasoning / answer triples spanning the 14 MMLU-Pro categories - mathematics, computer science, natural sciences, chemistry, physics, engineering, health, law, business, economics, psychology, philosophy, history and others (expanded to 122 fine-grained categories and 695 subcategories). Problems are written by a pool of generator models, solved with explicit reasoning by… See the full description on the dataset page: https://huggingface.co/datasets/dendriteholdings/Dendrite-Synth-Multi-Domain.texttext-generation100K<n<1M0 likes132 downloads23d agoHugging Face12prem-research /domains Domains dataset Documentation coming soon textquestion-answering10K<n<100K4 likes121 downloads2y agoHugging Face13bluecolor777 /Dendrite-Synth-Multi-Domain Dendrite Synth Multi-Domain A verified, difficulty-filtered, style-amplified synthetic corpus of question / reasoning / answer triples spanning the 14 MMLU-Pro categories - mathematics, computer science, natural sciences, chemistry, physics, engineering, health, law, business, economics, psychology, philosophy, history and others (expanded to 122 fine-grained categories and 695 subcategories). Problems are written by a pool of generator models, solved with explicit reasoning by… See the full description on the dataset page: https://huggingface.co/datasets/bluecolor777/Dendrite-Synth-Multi-Domain.texttext-generation100K<n<1M0 likes82 downloads14d agoHugging Face14Lots-of-LoRAs /task1320_country_domain_tld Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1320_country_domain_tld Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1320_country_domain_tld.texttext-generationn<1K0 likes81 downloads2y agoHugging Face15yoonholee /poetry-greats-public-domain Poetry Greats Curated, poem-level extracts from Project Gutenberg for 20 canonical English-language poets. All source texts are public domain in the US (pre-1929 publication). Intended as a reference set of "gold" examples for evaluation, few-shot prompting, and stylometric study. Contents 4,090 poems across 29 books and 20 poets: Poet Poems Samuel Taylor Coleridge 913 H. W. Longfellow 616 Christina Rossetti 459 Emily Dickinson 446 Percy Bysshe Shelley… See the full description on the dataset page: https://huggingface.co/datasets/yoonholee/poetry-greats-public-domain.tabulartext-generation1K<n<10K0 likes79 downloads5mo agoHugging Face16jvdgoltz /dbnl.org-dutch-public-domain Dataset Card for "dbnl.org-dutch-public-domain" Dataset Summary This dataset comprises a collection of texts from the Dutch Literature in the public domain, specifically from the DBNL (Digitale Bibliotheek voor de Nederlandse Letteren) public domain collection. The collection includes books, poems, songs, and other documentation, letters, etc., that are at least 140 years old and thus free of copyright restrictions. Each entry in the dataset corresponds to one section of… See the full description on the dataset page: https://huggingface.co/datasets/jvdgoltz/dbnl.org-dutch-public-domain.texttext-generation100K<n<1M0 likes75 downloads3y agoHugging Face17mir178 /shangkhachil-bengali-public-domain Bengali Public-Domain Literature 101 complete works by 21 authors, 11,250,629 characters. Corpus corpus-f8c532fcb4e7, built 2026-09-09. Where these texts are read https://shangkhachil.com — the reading site this corpus was built for. Free, no account, 246 works by 28 authors. The complete text of every work in this file can be read there. This file is the text. The site is the part a JSONL cannot be: Rights computed for the reader's own country, at the edge… See the full description on the dataset page: https://huggingface.co/datasets/mir178/shangkhachil-bengali-public-domain.tabulartext-generationn<1K0 likes72 downloads15d agoHugging Face18Maikobi /domain-generation-dataset Domain Generation Dataset This dataset contains 1,667 high-quality examples for fine-tuning language models to generate creative and relevant domain names for businesses, with built-in safety training and edge case handling. Dataset Creation Methodology: Hybrid approach combining Claude API generation with manual curation after encountering API reliability issues. Original Target: 2,000 examples → Final Result: 1,667 examples after deduplication and quality control.… See the full description on the dataset page: https://huggingface.co/datasets/Maikobi/domain-generation-dataset.texttext-generation1K<n<10K0 likes58 downloads1y agoHugging Face19MidGUI /Mid-Training_data_of_separate_domains Breaking the Data Barrier – Building GUI Agents Through Task Generalization This is the official dataset repository of GUIMid 1. Data Overview AgentBoard is composed of 9 diverse tasks: 7 vision and language tasks and 4 lanuage only tasks. The performances of different domains as mid-training data are as follows: Domains Observation WebArena (PR) WebArena (SR) AndroidWorld (SR) GUI Post-Training Only Image 26.3 6.2 9.0 Public Baselines GPT-4o-2024-11-20 Image… See the full description on the dataset page: https://huggingface.co/datasets/MidGUI/Mid-Training_data_of_separate_domains.texttext-generation1M<n<10M0 likes53 downloads1y agoHugging Face20EB-Sky /python-domain-10gb EB-Sky Python Domain (10GB) ~10GB of raw Python domain knowledge for continued pretraining: ~8GB of source code plus Stack Overflow Python Q&A. (2.49GB as stored — zstd-compressed parquet shards.) Sources Code: codeparrot/codeparrot-clean Q&A: koutch/stackoverflow_python Processing Size filter (200–200,000 chars per document) Auto-generated file removal (header heuristics) Hardcoded-secret regex scan, email redaction Exact deduplication (SHA-1)… See the full description on the dataset page: https://huggingface.co/datasets/EB-Sky/python-domain-10gb.texttext-generation1M<n<10M0 likes53 downloads21h agoHugging Face21khazarai /Multi-Domain-Reasoning-Benchmark Comprehensive Multi-Domain Reasoning Benchmark (CMDR-Bench) A systematic evaluation suite comprising 100 meticulously curated test cases across 10 distinct cognitive domains, designed to assess Large Language Models' capabilities in reasoning, problem-solving, and instruction-following. Each domain features a graduated difficulty scale (Levels 1–10), enabling fine-grained analysis of capability thresholds from elementary to expert-level complexity. texttext-generationn<1K3 likes47 downloads6mo agoHugging Face22Somtharu181coder /Legal_domain_ocr_extracted_Nepali_sft_dataset Nepali Legal SFT Dataset — Software Development & Operation Committee Order, 2083 Dataset Summary This dataset contains 31 single-turn instruction/response pairs in Nepali (Devanagari script), derived from a single Government of Nepal legal instrument: सफ्टवेयर विकास तथा सञ्चालन समिति (गठन) आदेश, २०८३ (Software Development and Operation Committee (Formation) Order, 2083) The order was issued by the Government of Nepal under Section 3 of the विकास समिति ऐन, २०१३… See the full description on the dataset page: https://huggingface.co/datasets/Somtharu181coder/Legal_domain_ocr_extracted_Nepali_sft_dataset.texttext-generationn<1K0 likes44 downloads16d agoHugging Face23EnDevSols /Multi-Domain-Reasoning-SFT Multi-Domain-Reasoning-SFT Dataset Summary The Multi-Domain-Reasoning-SFT dataset by EnDevSols is a large-scale, high-quality Supervised Fine-Tuning (SFT) dataset designed to train large language models in deep reasoning, technical analysis, and complex problem-solving. Consisting of nearly 580,000 meticulously structured examples, this dataset is specifically engineered to teach models how to "think" before they answer. It separates the internal cognitive process… See the full description on the dataset page: https://huggingface.co/datasets/EnDevSols/Multi-Domain-Reasoning-SFT.texttext-generation100K<n<1M0 likes42 downloads5mo agoHugging Face24Ethan615 /taiwan-conversation-context-100-domainsgated Taiwan Conversation Context 100 Domains Dataset Description Taiwan Conversation Context 100 Domains 是一套以台灣日常生活情境為核心設計的雙人對話文本資料集。 本資料集包含 100 個生活領域,每個領域各有 12,000 筆對話資料,總計約 1,200,000 筆對話樣本。每筆資料皆為雙人對話格式,包含 [A][B][A][B][A][B][A][B] 共 8 個發言,也就是 4 輪來回對話。 資料以繁體中文撰寫,並針對台灣在地語境設計,適合用於: 語音生成資料前處理 Text-to-Speech, TTS Spoken Dialogue Generation Conversational AI Customer Service Dialogue Modeling Role-play Dialogue Dataset 台灣繁體中文語音模型訓練 生活情境問答模型訓練 對話式 AI 助理訓練 RAG / Agent 測試資料… See the full description on the dataset page: https://huggingface.co/datasets/Ethan615/taiwan-conversation-context-100-domains.texttext-generation1M<n<10M2 likes39 downloads5mo agoHugging Face25DomainLLM /gerlayqa-bgb-paraphrased GerLayQA-BGB Paraphrased 🇩🇪⚖️ Dataset Description This is a paraphrased and restructured version of the GerLayQA BGB (Bürgerliches Gesetzbuch / German Civil Code) dataset, specifically prepared for fine-tuning large language models on German civil law question-answering tasks. Key Features 5,255 high-quality QA pairs about German Civil Law (BGB) Paraphrased questions to remove plagiarism while maintaining legal accuracy Structured 7-section answers following… See the full description on the dataset page: https://huggingface.co/datasets/DomainLLM/gerlayqa-bgb-paraphrased.textquestion-answering1K<n<10K0 likes32 downloads1y agoHugging Face26louisbrulenaudet /code-domaine-etat-collectivites-mayotte Code du domaine de l'Etat et des collectivités publiques applicable à la collectivité territoriale de Mayotte, non-instruct (2025-09-20) The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects. Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-domaine-etat-collectivites-mayotte.tabulartext-generationn<1K0 likes31 downloads1y agoHugging Face27SciCode /SciCode-Domain-Codegated DATA1: Domain-Specific Code Dataset Dataset Overview DATA1 is a large-scale domain-specific code dataset focusing on code samples from interdisciplinary fields such as biology, chemistry, materials science, and related areas. The dataset is collected and organized from GitHub repositories, covering 178 different domain topics with over 1.1 billion lines of code. Dataset Statistics Total Datasets: 178 CSV files Total Data Size: ~115 GB Total Lines… See the full description on the dataset page: https://huggingface.co/datasets/SciCode/SciCode-Domain-Code.tabulartext-generation1M<n<10M0 likes31 downloads7mo agoHugging Face28YAV-AI /llm-domain-specific-tough-questions LLM-Tough-Questions Dataset Description The LLM-Tough-Questions dataset is a synthetic collection designed to rigorously challenge and evaluate the capabilities of large language models. Comprising 10 meticulously formulated questions across 100 distinct domains, this dataset spans a wide spectrum of specialized fields including Mathematics, Fluid Dynamics, Neuroscience, and many others. Each question is developed to probe deep into the intricacies and subtleties of the… See the full description on the dataset page: https://huggingface.co/datasets/YAV-AI/llm-domain-specific-tough-questions.textquestion-answering1K<n<10K3 likes28 downloads2y agoHugging Face29DomainLLM /gerlayqa-combined-paraphrased GerLayQA Combined Paraphrased 🇩🇪⚖️ Dataset Description This is a combined, shuffled dataset merging both the BGB (civil law) and StGB (criminal law) paraphrased German legal QA datasets. All examples are paraphrased and restructured by GPT-5 for fine-tuning large language models on German legal question-answering tasks. Key Features 6,462 high-quality QA pairs covering both German Civil and Criminal Law Combined coverage: BGB (Bürgerliches Gesetzbuch) + StGB… See the full description on the dataset page: https://huggingface.co/datasets/DomainLLM/gerlayqa-combined-paraphrased.textquestion-answering1K<n<10K0 likes27 downloads1y agoHugging Face30Somtharu181coder /science_behavioral_and_domain_diversity_dataset Nepali Science SFT Dataset — Clean Candidate A high-quality Nepali Science Supervised Fine-Tuning (SFT) dataset containing short question–answer instruction-following examples written primarily in Nepali Devanagari script. This release is the clean candidate produced after structural validation, language checks, duplicate analysis, and Unicode-contamination filtering. Dataset Overview Property Value Dataset file clean_candidate.jsonl Records 29,320… See the full description on the dataset page: https://huggingface.co/datasets/Somtharu181coder/science_behavioral_and_domain_diversity_dataset.texttext-generation10K<n<100K0 likes26 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.