CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01bitext /Bitext-customer-support-llm-chatbot-training-dataset Bitext - Customer Service Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the Customer Support sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-customer-support-llm-chatbot-training-dataset.textquestion-answering10K<n<100K194 likes7.4k downloads2y agoHugging Face02openadmet /cyp-challenge-train-test CYP Challenge Train/Test Dataset A high-quality experimental dataset for predicting inhibition of the major drug-metabolizing Cytochrome P450 enzymes (CYP1A2, CYP2C9, CYP2D6, CYP3A4), released as part of the OpenADMET CYP Inhibition Blind Challenge. Blog post: Announcing OpenADMET’s CYP inhibition blind challenge Challenge Space: OpenADMET CYP Inhibition Blind Challenge Challenge period: August 17, 2026 - November 3, 2026 Produced by: OpenADMET CHANGELOG Updated… See the full description on the dataset page: https://huggingface.co/datasets/openadmet/cyp-challenge-train-test.tabulartabular-regression10K<n<100K9 likes2.8k downloads25d agoHugging Face03openadmet /pxr-challenge-train-test PXR Challenge Train/Test Dataset A high-quality experimental dataset for predicting human Pregnane-X Receptor (PXR) induction, comprising over 11,000 compounds screened using a high-fidelity in-house assay. This is the largest publicly available PXR activity dataset, released as part of the OpenADMET PXR Induction Blind Challenge. Blog post: Announcing the Next OpenADMET Blind Challenge: Predicting PXR Induction Challenge Space: openadmet/pxr-challenge Challenge period: April 1… See the full description on the dataset page: https://huggingface.co/datasets/openadmet/pxr-challenge-train-test.tabulartabular-regression10K<n<100K17 likes1.6k downloads20d agoHugging Face04UW-Madison-Lee-Lab /MMLU-Pro-CoT-Train-Labeled Dataset Details Modality: Text Format: CSV Size: 10K - 100K rows Total Rows: 84,098 License: MIT Libraries Supported: datasets, pandas, croissant Structure Each row in the dataset includes: question: The query posed in the dataset. answer: The correct response. category: The domain of the question (e.g., math, science). src: The source of the question. id: A unique identifier for each entry. chain_of_thoughts: Step-by-step reasoning steps leading to the answer. labels:… See the full description on the dataset page: https://huggingface.co/datasets/UW-Madison-Lee-Lab/MMLU-Pro-CoT-Train-Labeled.text10K<n<100K6 likes1.3k downloads2y agoHugging Face05bitext /Bitext-retail-ecommerce-llm-chatbot-training-dataset Bitext - Retail (eCommerce) Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [Retail (eCommerce)] sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-retail-ecommerce-llm-chatbot-training-dataset.textquestion-answering10K<n<100K19 likes1.3k downloads2y agoHugging Face06SmellsLikeAISpirit /plant-disease-trainimage10K<n<100K0 likes1.1k downloads6mo agoHugging Face07openadmet /openadmet-expansionrx-challenge-train-data OpenADMET-ExpansionRx Challenge training dataset This dataset contains real-work ADMET data from a recently prosecuted series of drug discovery campaigns by Expansion Therapeutics on RNA mediated diseases. While optimising candidate molecules for their preclinical programs Expansion collected a variety of ADMET data for off-targets and properties of interest in the traditional game of “whack-a-mole” familiar to all drug hunters. Now, they’ve made the bold and generous decision to… See the full description on the dataset page: https://huggingface.co/datasets/openadmet/openadmet-expansionrx-challenge-train-data.tabular10K<n<100K10 likes876 downloads10mo agoHugging Face08HaoranoLee /pizza_st_human_mouse_v1_train_with_labelstextn<1K0 likes822 downloads9mo agoHugging Face09bitext /Bitext-events-ticketing-llm-chatbot-training-dataset Bitext - Events and Ticketing Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [events and ticketing] sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-events-ticketing-llm-chatbot-training-dataset.textquestion-answering10K<n<100K1 likes779 downloads2y agoHugging Face10jliang097 /FARM_training_test FARM Aerial Radio Map (ARM) Dataset Paper: FARM: Foundational Aerial Radio Map for Intelligent Low-Altitude Networking (https://arxiv.org/abs/2604.17362) Overview This repository releases the constructed ARM datasets based on ARM-Omni for FARM training, in-domain evaluation (D1-D10), and zero-shot evaluation (P1, F1, and A1). The dataset coverage is summarized below: Dataset Frequencies (GHz) Max Rx Height (m) Beamwidths Map Grid Size Volume D1 2.1… See the full description on the dataset page: https://huggingface.co/datasets/jliang097/FARM_training_test.tabularimage-to-image10K<n<100K1 likes296 downloads5mo agoHugging Face11MUG-V /MUG-V-Training-Samples MUG-V Training Samples Sample training dataset for the MUG-V 10B video generation model training framework. Dataset Description This dataset contains pre-processed training samples for quick-start validation and testing of the MUG-V Megatron-LM training pipeline. It includes: VideoVAE-encoded latents (8×8×8 compressed video representations) T5-XXL text features (4096-dim embeddings) Training metadata CSV (sample mapping and configuration) ⚠️ Note: This is a sample… See the full description on the dataset page: https://huggingface.co/datasets/MUG-V/MUG-V-Training-Samples.texttext-to-video1K<n<10K0 likes294 downloads11mo agoHugging Face12previtus /STARCOP_allbands_Train1gated STARCOP dataset STARCOP dataset: Semantic Segmentation of Methane Plumes with Hyperspectral Machine Learning Models 🌈🛰️Authors: Vít Růžička, Gonzalo Mateo-Garcia, Luis Gómez-Chova, Anna Vaughan, Luis Guanter and Andrew Markham Fast data preview in: dataset_exploration.ipynb Main repository: github/spaceml-org/STARCOP Task: Methane is the second most important greenhouse gas contributor to climate change; at the same time its reduction has been denoted as one of the… See the full description on the dataset page: https://huggingface.co/datasets/previtus/STARCOP_allbands_Train1.imageimage-segmentation1K<n<10K3 likes290 downloads2y agoHugging Face13zluvolyote /Dream_Traintext1M<n<10M0 likes225 downloads4y agoHugging Face14bitext /Bitext-insurance-llm-chatbot-training-dataset Bitext - Insurance Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [insurance] sector can be easily achieved using our two-step approach to LLM Fine-Tuning. An… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-insurance-llm-chatbot-training-dataset.textquestion-answering10K<n<100K8 likes215 downloads2y agoHugging Face15CO-Bench /FrontierCO-Traintabularn<1K0 likes212 downloads1y agoHugging Face16PersonaBias /Reverse-hybrid-train-no-persona-meantabulartext-classification100K<n<1M0 likes211 downloads2mo agoHugging Face17bitext /Bitext-telco-llm-chatbot-training-dataset Bitext - Telco Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [telco] sector can be easily achieved using our two-step approach to LLM Fine-Tuning. An overview of… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-telco-llm-chatbot-training-dataset.textquestion-answering10K<n<100K2 likes201 downloads2y agoHugging Face18bitext /Bitext-travel-llm-chatbot-training-dataset Bitext - Travel Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [Travel] sector can be easily achieved using our two-step approach to LLM Fine-Tuning. An overview of… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-travel-llm-chatbot-training-dataset.textquestion-answering10K<n<100K4 likes184 downloads2y agoHugging Face19PersonaBias /Reverse-hybrid-correct-train-no-persona-meantabulartext-classification100K<n<1M0 likes160 downloads2mo agoHugging Face20CreitinGameplays /DeepSeek-R1-Distill-Qwen-32B_NUMINA_train_amc_aime-llama3.1tabular1K<n<10K0 likes153 downloads2y agoHugging Face21PersonaBias /Original-hybrid-correct-train-no-persona-meantabulartext-classification100K<n<1M0 likes139 downloads2mo agoHugging Face22sebastian-hofstaetter /tripclick-training TripClick Baselines with Improved Training Data Establishing Strong Baselines for TripClick Health Retrieval Sebastian Hofstätter, Sophia Althammer, Mete Sertkan and Allan Hanbury https://arxiv.org/abs/2201.00365 tl;dr We create strong re-ranking and dense retrieval baselines (BERTCAT, BERTDOT, ColBERT, and TK) for TripClick (health ad-hoc retrieval). We improve the – originally too noisy – training data with a simple negative sampling policy. We achieve large gains over BM25 in the… See the full description on the dataset page: https://huggingface.co/datasets/sebastian-hofstaetter/tripclick-training.tabulartext-retrieval1M<n<10M1 likes135 downloads4y agoHugging Face23CaiYuanhao /OmniVCus-Train [NeurIPS 2025] OmniVCus: Feedforward Subject-driven Video Customization with Multimodal Control Conditions Dataset Description This dataset supports multi-modal constrol video generation. It contains ~80K data samples processed from 140K videos by our VideoCus-Factory pipeline. Each data sample includes the original video, text prompts, subject reference image, depth video, mask video, and motion video conditions. Here is a data example: Generated… See the full description on the dataset page: https://huggingface.co/datasets/CaiYuanhao/OmniVCus-Train.text100K<n<1M2 likes134 downloads9mo agoHugging Face24PersonaBias /Original-hybrid-train-no-persona-meantabulartext-classification100K<n<1M0 likes134 downloads2mo agoHugging Face25SpX-DAC /training_datatextn<1K0 likes126 downloads9mo agoHugging Face26Harsit /xnli2.0_train_arabictext100K<n<1M1 likes112 downloads4y agoHugging Face27sschet /ROCOv2-traintext10K<n<100K0 likes109 downloads2y agoHugging Face28redmadrobot-rnd /pii_train Russian PII NER Training Dataset Dataset Description This is the training corpus for PII (Personally Identifiable Information) detection and Named Entity Recognition (NER) on Russian-language text. It targets guardrail and anonymization pipelines that must reliably find personal data (names, addresses, contacts) and Russian identity-document numbers (passport, SNILS, INN, OMS, etc.) in text. The corpus combines real, manually annotated examples from production… See the full description on the dataset page: https://huggingface.co/datasets/redmadrobot-rnd/pii_train.texttoken-classification10K<n<100K1 likes108 downloads14d agoHugging Face29gorkaartola /SC-train-valid-test_SDG-Descriptionsnli-label: (0) entailment (2) contradiction tabular1K<n<10K0 likes101 downloads4y agoHugging Face30bitext /Bitext-mortgage-loans-llm-chatbot-training-dataset Bitext - Mortgage and Loans Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [Mortgage and Loans] sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-mortgage-loans-llm-chatbot-training-dataset.textquestion-answering10K<n<100K5 likes95 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.