CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01racineai /VDR_MEGA_MultiDomain_DocRetrieval Visual Document Retrieval Dataset Overview This dataset is designed for training visual document retrieval models. It combines multiple datasets from the VDR series, Colpali, and LlamaIndex to create the most comprehensive training resource for visual document retrieval tasks. Dataset Structure The dataset contains structured fields including unique identifiers with string lengths ranging from 45 to 50 characters, search query text with variable lengths between… See the full description on the dataset page: https://huggingface.co/datasets/racineai/VDR_MEGA_MultiDomain_DocRetrieval.imagevisual-document-retrieval1M<n<10M24 likes72k downloads6mo agoHugging Face02AMAImedia /multidomain-kazakh-dataset ⚡ Each donation funds the next large quant. I host free GGUF or MoE quants as independent research. Local hardware: Mechrevo Kuangshi GM7AG0M — RTX 3060 Laptop 6GB GDDR6, 64GB DDR5, i7-12700H (14C/20T, 4.7GHz), Windows 11, Samsung 990 Pro. Good for imatrix and 0.6–35B-class work in RAM. 9B+ and searches need rented H200/Blackwell, typically $100 per quant. 🎉 Boosty🦄 &nbsp;|&nbsp; ☕ Buy Me a Coffee🦄 &nbsp;|&nbsp; ⭐ DonationAlerts🦄 💚 Thanks to Hugging Face for extra storage.🦄… See the full description on the dataset page: https://huggingface.co/datasets/AMAImedia/multidomain-kazakh-dataset.text10M<n<100M0 likes710 downloads6d agoHugging Face03multi-domain-reasoning /mmlutext10K<n<100K0 likes415 downloads2y agoHugging Face04kz-transformers /multidomain-kazakh-dataset Dataset Description Point of Contact: Sanzhar Murzakhmetov, Besultan Sagyndyk Dataset Summary MDBKD | Multi-Domain Bilingual Kazakh Dataset is a Kazakh-language dataset containing just over 24 883 808 unique texts from multiple domains. Supported Tasks 'MLM/CLM': can be used to train a model for casual and masked languange modeling Languages The kk code for Kazakh as generally spoken in the Kazakhstan Data Instances For each instance… See the full description on the dataset page: https://huggingface.co/datasets/kz-transformers/multidomain-kazakh-dataset.texttext-generation10M<n<100M30 likes272 downloads1y agoHugging Face05multi-domain-reasoning /commonsense_qatext1K<n<10K2 likes242 downloads2y agoHugging Face06asahi417 /multi-domain-document-classification multi_domain_document_classification Multi-domain document classification datasets. Biomedical: chemprot, rct-sample Computer Science: citation_intent, sciie Customer Review: amcd, yelp_review Social Media: tweet_eval_irony, tweet_eval_hate, tweet_eval_emotion The yelp_review dataset is randomly downsampled to 2000/2000/8000 for test/validation/train. chemprot citation_intent hyperpartisan_news rct_sample sciie amcd yelp_review tweet_eval_irony tweet_eval_hate… See the full description on the dataset page: https://huggingface.co/datasets/asahi417/multi-domain-document-classification.text10K<n<100K0 likes234 downloads4y agoHugging Face07acmc /multi_domain_ai_human_text multi_domain_ai_human_text — Datasheet Balanced, multi-domain AI-vs-human text detection benchmark with dedicated out-of-distribution and adversarial evaluation panels. Built by scripts/build_paper_dataset.py from an 11-corpus unified aggregation. Splits Split AI Human Total Purpose train 300,000 300,000 600,000 training (balanced, English, clean) validation 2,996 2,999 5,995 model selection test 4,991 4,999 9,990 in-distribution test… See the full description on the dataset page: https://huggingface.co/datasets/acmc/multi_domain_ai_human_text.tabulartext-classification1M<n<10M2 likes183 downloads1mo agoHugging Face08liy140 /multidomain-measextract-corpus A Multi-Domain Corpus for Measurement Extraction (Seq2Seq variant) A detailed description of corpus creation can be found here. This dataset contains the training and validation and test data for each of the three datasets measeval, bm, and msp. The measeval, and msp datasets were adapted from the MeasEval (Harper et al., 2021) and the Material Synthesis Procedual (Mysore et al., 2019) corpus respectively. This repository aggregates extraction to paragraph-level for msp and… See the full description on the dataset page: https://huggingface.co/datasets/liy140/multidomain-measextract-corpus.texttoken-classification1K<n<10K0 likes167 downloads3y agoHugging Face09strangerguardhf /NSFW-MultiDomain-Classificationgated NSFW_MultiDomain The NSFW_MultiDomain dataset is a curated image classification dataset focused on multi-domain adult content recognition. It consists of 5 distinct categories aimed at facilitating the development of robust NSFW (Not Safe For Work) image classification models. This dataset enables training and benchmarking of models that can distinguish between subtle variations in explicit and non-explicit content across artistic, animated, and real-world imagery.… See the full description on the dataset page: https://huggingface.co/datasets/strangerguardhf/NSFW-MultiDomain-Classification.imageimage-classification10K<n<100K25 likes167 downloads1y agoHugging Face10Lightcap /multi-domain-cloudflare-observability Multi-Domain Cloudflare Web Traffic, Performance and Security Observability Dataset This dataset contains multi-domain Cloudflare analytics exported into analysis-ready Parquet tables. It combines HTTP request aggregates, hourly traffic trends, path and referrer dimensions, country/device/browser breakdowns, DNS analytics, cache behavior, Web Vitals/RUM performance signals, redirect and ruleset metadata, and available firewall/security aggregates across 200 websites. The… See the full description on the dataset page: https://huggingface.co/datasets/Lightcap/multi-domain-cloudflare-observability.tabulartime-series-forecasting100K<n<1M1 likes151 downloads5mo agoHugging Face11Zihao-Li /multidomain_rcot_physicstext100K<n<1M0 likes141 downloads2mo agoHugging Face12FatimahEmadEldin /flair-multidomain-parquet FLAIR multi-domain pack — 256 px 24,475 tiles · 9 French domains · 1.94 GB, repacked from IGNF/FLAIR-HUB into Parquet so a Colab runtime can stream it. Use this pack for: the headline runs; 256 px gives the sharpest boundaries. Pack Patch Size One shard per domain flair-multidomain-parquet 256 px 1.94 GB yes ← this repo flair-multidomain-parquet-128 128 px 0.9 GB yes Columns column type contents rgb JPEG bytes aerial RGB, 256², quality… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/flair-multidomain-parquet.tabularimage-segmentation10K<n<100K1 likes133 downloads1mo agoHugging Face13abdullaharoon /Urdu-Multi-Domain-Benchmark Urdu Multi-Domain Datasets 33 labeled Urdu datasets (288,899 examples) for text classification in Nastaliq (Perso-Arabic) and Roman Urdu (Latin). Each domain is a separate Hub subset so you can download one task at a time. Authors: Muhammad Abdullah Haroon and Maryam Bashir, FAST-NUCES, Lahore. Companion paper: Domain Robustness of Multilingual NLP Models Across Urdu and Roman Urdu Scripts. Permanent archive: Zenodo DOI 10.5281/zenodo.22195610. How to load Pick a… See the full description on the dataset page: https://huggingface.co/datasets/abdullaharoon/Urdu-Multi-Domain-Benchmark.texttext-classification100K<n<1M0 likes130 downloads24d agoHugging Face14dendriteholdings /Dendrite-Synth-Multi-Domain Dendrite Synth Multi-Domain A verified, difficulty-filtered, style-amplified synthetic corpus of question / reasoning / answer triples spanning the 14 MMLU-Pro categories - mathematics, computer science, natural sciences, chemistry, physics, engineering, health, law, business, economics, psychology, philosophy, history and others (expanded to 122 fine-grained categories and 695 subcategories). Problems are written by a pool of generator models, solved with explicit reasoning by… See the full description on the dataset page: https://huggingface.co/datasets/dendriteholdings/Dendrite-Synth-Multi-Domain.texttext-generation100K<n<1M0 likes130 downloads22d agoHugging Face15FatimahEmadEldin /flair-multidomain-parquet-512 FLAIR multi-domain pack — 512 px (0.20 m/px) 24,475 tiles · 9 French domains · 4.86 GB, repacked from IGNF/FLAIR-HUB into Parquet so a Colab runtime can stream it. This is FLAIR's native aerial resolution. A 102.4 m tile at 512 px is 0.20 m/px — the sampling the COSIA labels were digitised at. The smaller packs below are downsamples, and results from different sampling must not share a table: changing ground sampling changes the task, not just the cost. Pack Patch Ground… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/flair-multidomain-parquet-512.tabularimage-segmentation10K<n<100K0 likes122 downloads1mo agoHugging Face16FatimahEmadEldin /flair-multidomain-parquet-128 FLAIR multi-domain pack — 128 px 24,475 tiles · 9 French domains · 0.9 GB, repacked from IGNF/FLAIR-HUB into Parquet so a Colab runtime can stream it. Use this pack for: modest GPUs. Same 9 domains and same labels, a quarter of the pixels per training step - the practical choice on a free Colab T4. Pack Patch Size One shard per domain flair-multidomain-parquet 256 px 1.94 GB yes flair-multidomain-parquet-128 128 px 0.9 GB yes ← this repo Columns… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/flair-multidomain-parquet-128.tabularimage-segmentation10K<n<100K1 likes113 downloads1mo agoHugging Face17breakend /nllb-multi-domainNLLB Multi Domain is a set of professionally-translated sentences in News, Unscripted informal speech, and Health domains. It is designed to enable assessment of out-of-domain performance and to study domain adaptation for machine translation. Each domain has approximately 3000 sentences.text10K<n<100K3 likes104 downloads4y agoHugging Face18bluecolor777 /Dendrite-Synth-Multi-Domain Dendrite Synth Multi-Domain A verified, difficulty-filtered, style-amplified synthetic corpus of question / reasoning / answer triples spanning the 14 MMLU-Pro categories - mathematics, computer science, natural sciences, chemistry, physics, engineering, health, law, business, economics, psychology, philosophy, history and others (expanded to 122 fine-grained categories and 695 subcategories). Problems are written by a pool of generator models, solved with explicit reasoning by… See the full description on the dataset page: https://huggingface.co/datasets/bluecolor777/Dendrite-Synth-Multi-Domain.texttext-generation100K<n<1M0 likes82 downloads13d agoHugging Face19multi-domain-reasoning /agieval_eval_lsat_rctextn<1K0 likes78 downloads2y agoHugging Face20Atmanstr /MultiDomain_Instruction MultiDomain_Instruction A multi-domain instruction dataset designed for instruction tuning and supervised fine-tuning (SFT) of large language models. The dataset contains tasks from multiple domains such as question answering, summarization, reasoning, classification, and general knowledge to improve model generalization. Overview MultiDomain_Instruction is created to support instruction-following training for LLMs. Instead of focusing on a single task, this dataset combines… See the full description on the dataset page: https://huggingface.co/datasets/Atmanstr/MultiDomain_Instruction.texttext-classification100K<n<1M1 likes71 downloads7mo agoHugging Face21climba /t2i-nla-multidomain-1m-training-data-public T2I-NLA Multidomain 1M Training Data Public backup of final T2I-NLA reader training data. Balanced multidomain 1M training data for general FLUX AR/AV readers. This repository is intended to make final-reader retraining possible after the original server is no longer available. See manifest.json for the original local path, size, and associated reader models. tabular1M<n<10M0 likes71 downloads4mo agoHugging Face22JustinDuc /MultiDomain-QADialog 📚 MultiDomain-QADialog Dataset This repository contains the processed, multi-source dataset used to train the SHARE Model for dialogue inference. The dataset combines three prominent resources in the dialogue space: MediaSum – dialogues from broadcast transcripts (300k samples) SAMSum – messenger-style casual conversations (16K samples) SODA – million-scale, high-quality dialogue dataset (~1M samples) All datasets have been harmonized into a unified format and stored in sharded… See the full description on the dataset page: https://huggingface.co/datasets/JustinDuc/MultiDomain-QADialog.tabular1M<n<10M0 likes61 downloads1y agoHugging Face23dutta18 /multidomain-VQA-with-cot-trace-9KThis dataset of 9K samples has been created with AOKVQA Train & Val split, TDIUC Val Split (Quantitative and Physical Reasoning Questions only). This is a multidomain dataset solely created to test the multidomain reasoning knowledge of VLM's, it can be used for inference or rapid prototyping. The synthetic COT trace has been generated using Qwen-VL-32B Model. This is for educational and research purposes only. All the copyright belongs to the original owners of the datasets. image10K<n<100K0 likes55 downloads8mo agoHugging Face24khazarai /Multi-Domain-Reasoning-Benchmark Comprehensive Multi-Domain Reasoning Benchmark (CMDR-Bench) A systematic evaluation suite comprising 100 meticulously curated test cases across 10 distinct cognitive domains, designed to assess Large Language Models' capabilities in reasoning, problem-solving, and instruction-following. Each domain features a graduated difficulty scale (Levels 1–10), enabling fine-grained analysis of capability thresholds from elementary to expert-level complexity. texttext-generationn<1K3 likes54 downloads6mo agoHugging Face25TaskPuppyAI /qwen3.8-targeted-multidomain-250 Qwen3.8 Max Multidomain Code Review 250 A 250-record synthetic multilingual code-review dataset generated with Qwen3.8 Max and reviewed/corrected with ChatGPT 5.6 Sol High. All records use the same strict review instruction and ask the model to report only concrete defects supported by the visible code and stated contract. Dataset Summary The publication artifact contains 250 unique records using the schema: { "instruction": "...", "input": "...", "output":… See the full description on the dataset page: https://huggingface.co/datasets/TaskPuppyAI/qwen3.8-targeted-multidomain-250.textn<1K0 likes53 downloads17d agoHugging Face26llm-uncertainty-head /test_multidomain Dataset Card for "test_multidomain" More Information needed textn<1K0 likes50 downloads2y agoHugging Face27Labradorlabs /bsca-binary-source-gold-v3-multidomain BSCA Gold v3 Multidomain Address-grounded P1 pairs for stripped pseudo-C → source retrieval. Dataset ID: GD_19330e06aae0462447c1fd05ccaa38d7 Accepted P1 pairs: 42449 Repositories: 138 Target formats: {"elf": 40733, "pe": 1716} Target architectures: {"aarch64": 1672, "x86": 1903, "x86_64": 38874} Internal quality GPA: 3.660; target pass: True Use train.jsonl for fitting, development.jsonl for model selection, and the immutable test.jsonl only after selection. dataset_card.json… See the full description on the dataset page: https://huggingface.co/datasets/Labradorlabs/bsca-binary-source-gold-v3-multidomain.tabularfeature-extraction10K<n<100K0 likes47 downloads1mo agoHugging Face28CultriX /qwen-expert-multidomain-chat-v1text10K<n<100K0 likes44 downloads1y agoHugging Face29EnDevSols /Multi-Domain-Reasoning-SFT Multi-Domain-Reasoning-SFT Dataset Summary The Multi-Domain-Reasoning-SFT dataset by EnDevSols is a large-scale, high-quality Supervised Fine-Tuning (SFT) dataset designed to train large language models in deep reasoning, technical analysis, and complex problem-solving. Consisting of nearly 580,000 meticulously structured examples, this dataset is specifically engineered to teach models how to "think" before they answer. It separates the internal cognitive process… See the full description on the dataset page: https://huggingface.co/datasets/EnDevSols/Multi-Domain-Reasoning-SFT.texttext-generation100K<n<1M0 likes44 downloads5mo agoHugging Face30rohit94 /gemini-reasoning-traces-multidomain Gemini-Reasoning-Traces-Multidomain A curated dataset of 2,282 samples with high-quality, reasoning traces distilled from Gemini-2.5-Pro across 8 diverse domains. Every sample (question + reasoning + answer) fits within 4,096 Gemma-3 tokens, making it ideal for fine-tuning small language models on constrained hardware. Why This Dataset? High-quality reasoning datasets are scarce — especially for domains requiring structured thought such as creative writing, summarization… See the full description on the dataset page: https://huggingface.co/datasets/rohit94/gemini-reasoning-traces-multidomain.tabular1K<n<10K0 likes42 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.