CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01bitext /Bitext-customer-support-llm-chatbot-training-dataset Bitext - Customer Service Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the Customer Support sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-customer-support-llm-chatbot-training-dataset.textquestion-answering10K<n<100K195 likes7.8k downloads2y agoHugging Face02openadmet /cyp-challenge-train-test CYP Challenge Train/Test Dataset A high-quality experimental dataset for predicting inhibition of the major drug-metabolizing Cytochrome P450 enzymes (CYP1A2, CYP2C9, CYP2D6, CYP3A4), released as part of the OpenADMET CYP Inhibition Blind Challenge. Blog post: Announcing OpenADMET’s CYP inhibition blind challenge Challenge Space: OpenADMET CYP Inhibition Blind Challenge Challenge period: August 17, 2026 - November 3, 2026 Produced by: OpenADMET CHANGELOG Updated… See the full description on the dataset page: https://huggingface.co/datasets/openadmet/cyp-challenge-train-test.tabulartabular-regression10K<n<100K9 likes2.8k downloads26d agoHugging Face03openadmet /pxr-challenge-train-test PXR Challenge Train/Test Dataset A high-quality experimental dataset for predicting human Pregnane-X Receptor (PXR) induction, comprising over 11,000 compounds screened using a high-fidelity in-house assay. This is the largest publicly available PXR activity dataset, released as part of the OpenADMET PXR Induction Blind Challenge. Blog post: Announcing the Next OpenADMET Blind Challenge: Predicting PXR Induction Challenge Space: openadmet/pxr-challenge Challenge period: April 1… See the full description on the dataset page: https://huggingface.co/datasets/openadmet/pxr-challenge-train-test.tabulartabular-regression10K<n<100K17 likes1.6k downloads21d agoHugging Face04bitext /Bitext-retail-ecommerce-llm-chatbot-training-dataset Bitext - Retail (eCommerce) Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [Retail (eCommerce)] sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-retail-ecommerce-llm-chatbot-training-dataset.textquestion-answering10K<n<100K19 likes1.4k downloads2y agoHugging Face05UW-Madison-Lee-Lab /MMLU-Pro-CoT-Train-Labeled Dataset Details Modality: Text Format: CSV Size: 10K - 100K rows Total Rows: 84,098 License: MIT Libraries Supported: datasets, pandas, croissant Structure Each row in the dataset includes: question: The query posed in the dataset. answer: The correct response. category: The domain of the question (e.g., math, science). src: The source of the question. id: A unique identifier for each entry. chain_of_thoughts: Step-by-step reasoning steps leading to the answer. labels:… See the full description on the dataset page: https://huggingface.co/datasets/UW-Madison-Lee-Lab/MMLU-Pro-CoT-Train-Labeled.text10K<n<100K6 likes1.3k downloads2y agoHugging Face06SmellsLikeAISpirit /plant-disease-trainimage10K<n<100K0 likes1.1k downloads6mo agoHugging Face07openadmet /openadmet-expansionrx-challenge-train-data OpenADMET-ExpansionRx Challenge training dataset This dataset contains real-work ADMET data from a recently prosecuted series of drug discovery campaigns by Expansion Therapeutics on RNA mediated diseases. While optimising candidate molecules for their preclinical programs Expansion collected a variety of ADMET data for off-targets and properties of interest in the traditional game of “whack-a-mole” familiar to all drug hunters. Now, they’ve made the bold and generous decision to… See the full description on the dataset page: https://huggingface.co/datasets/openadmet/openadmet-expansionrx-challenge-train-data.tabular10K<n<100K10 likes889 downloads10mo agoHugging Face08bitext /Bitext-events-ticketing-llm-chatbot-training-dataset Bitext - Events and Ticketing Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [events and ticketing] sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-events-ticketing-llm-chatbot-training-dataset.textquestion-answering10K<n<100K1 likes881 downloads2y agoHugging Face09HaoranoLee /pizza_st_human_mouse_v1_train_with_labelstextn<1K0 likes810 downloads9mo agoHugging Face10MUG-V /MUG-V-Training-Samples MUG-V Training Samples Sample training dataset for the MUG-V 10B video generation model training framework. Dataset Description This dataset contains pre-processed training samples for quick-start validation and testing of the MUG-V Megatron-LM training pipeline. It includes: VideoVAE-encoded latents (8×8×8 compressed video representations) T5-XXL text features (4096-dim embeddings) Training metadata CSV (sample mapping and configuration) ⚠️ Note: This is a sample… See the full description on the dataset page: https://huggingface.co/datasets/MUG-V/MUG-V-Training-Samples.texttext-to-video1K<n<10K0 likes300 downloads11mo agoHugging Face11previtus /STARCOP_allbands_Train1gated STARCOP dataset STARCOP dataset: Semantic Segmentation of Methane Plumes with Hyperspectral Machine Learning Models 🌈🛰️Authors: Vít Růžička, Gonzalo Mateo-Garcia, Luis Gómez-Chova, Anna Vaughan, Luis Guanter and Andrew Markham Fast data preview in: dataset_exploration.ipynb Main repository: github/spaceml-org/STARCOP Task: Methane is the second most important greenhouse gas contributor to climate change; at the same time its reduction has been denoted as one of the… See the full description on the dataset page: https://huggingface.co/datasets/previtus/STARCOP_allbands_Train1.imageimage-segmentation1K<n<10K3 likes290 downloads2y agoHugging Face12bitext /Bitext-telco-llm-chatbot-training-dataset Bitext - Telco Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [telco] sector can be easily achieved using our two-step approach to LLM Fine-Tuning. An overview of… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-telco-llm-chatbot-training-dataset.textquestion-answering10K<n<100K2 likes259 downloads2y agoHugging Face13zluvolyote /Dream_Traintext1M<n<10M0 likes226 downloads4y agoHugging Face14bitext /Bitext-insurance-llm-chatbot-training-dataset Bitext - Insurance Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [insurance] sector can be easily achieved using our two-step approach to LLM Fine-Tuning. An… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-insurance-llm-chatbot-training-dataset.textquestion-answering10K<n<100K8 likes215 downloads2y agoHugging Face15CO-Bench /FrontierCO-Traintabularn<1K0 likes212 downloads1y agoHugging Face16PersonaBias /Reverse-hybrid-train-no-persona-meantabulartext-classification100K<n<1M0 likes212 downloads2mo agoHugging Face17bitext /Bitext-travel-llm-chatbot-training-dataset Bitext - Travel Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [Travel] sector can be easily achieved using our two-step approach to LLM Fine-Tuning. An overview of… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-travel-llm-chatbot-training-dataset.textquestion-answering10K<n<100K4 likes189 downloads2y agoHugging Face18PersonaBias /Reverse-hybrid-correct-train-no-persona-meantabulartext-classification100K<n<1M0 likes160 downloads2mo agoHugging Face19CreitinGameplays /DeepSeek-R1-Distill-Qwen-32B_NUMINA_train_amc_aime-llama3.1tabular1K<n<10K0 likes153 downloads2y agoHugging Face20CaiYuanhao /OmniVCus-Train [NeurIPS 2025] OmniVCus: Feedforward Subject-driven Video Customization with Multimodal Control Conditions Dataset Description This dataset supports multi-modal constrol video generation. It contains ~80K data samples processed from 140K videos by our VideoCus-Factory pipeline. Each data sample includes the original video, text prompts, subject reference image, depth video, mask video, and motion video conditions. Here is a data example: Generated… See the full description on the dataset page: https://huggingface.co/datasets/CaiYuanhao/OmniVCus-Train.text100K<n<1M2 likes153 downloads9mo agoHugging Face21SpX-DAC /training_datatextn<1K0 likes141 downloads9mo agoHugging Face22PersonaBias /Original-hybrid-correct-train-no-persona-meantabulartext-classification100K<n<1M0 likes127 downloads2mo agoHugging Face23PersonaBias /Original-hybrid-train-no-persona-meantabulartext-classification100K<n<1M0 likes118 downloads2mo agoHugging Face24redmadrobot-rnd /pii_train Russian PII NER Training Dataset Dataset Description This is the training corpus for PII (Personally Identifiable Information) detection and Named Entity Recognition (NER) on Russian-language text. It targets guardrail and anonymization pipelines that must reliably find personal data (names, addresses, contacts) and Russian identity-document numbers (passport, SNILS, INN, OMS, etc.) in text. The corpus combines real, manually annotated examples from production… See the full description on the dataset page: https://huggingface.co/datasets/redmadrobot-rnd/pii_train.texttoken-classification10K<n<100K1 likes113 downloads15d agoHugging Face25Harsit /xnli2.0_train_arabictext100K<n<1M1 likes111 downloads4y agoHugging Face26sschet /ROCOv2-traintext10K<n<100K0 likes109 downloads2y agoHugging Face27bitext /Bitext-mortgage-loans-llm-chatbot-training-dataset Bitext - Mortgage and Loans Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [Mortgage and Loans] sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-mortgage-loans-llm-chatbot-training-dataset.textquestion-answering10K<n<100K5 likes98 downloads2y agoHugging Face28traintogpb /aihub-koen-translation-integrated-large-10m AI Hub Ko-En Translation Dataset (Integrated) AI Hub의 한-영 번역 관련 데이터셋 8개를 병합한 자료입니다. 병합 시 총 데이터 개수는 10,416,509개 이며, train / validation / test는 8:1:1 비율로 분할되었습니다. base-10m: 병합 데이터 100% 사용, 총 10,416,509개 mini-1m: 병합 데이터 10% 사용 (base-10m의 각 세트 내에서 10% 임의 선택), 총 1,041,651개 tiny-100k: 병합 데이터 1% 사용 (base-10m의 각 세트 내에서 1% 임의 선택), 총 104,165개 Subsets 활용한 데이터셋 목록은 다음과 같으며, 데이터셋 이름 옆 번호는 aihubshell에서의 datasetkey입니다. 전문분야 한영 말뭉치 (111) 총 개수: 1,350,000 중복 제거 후 개수: 1,350,000 사용 칼럼:… See the full description on the dataset page: https://huggingface.co/datasets/traintogpb/aihub-koen-translation-integrated-large-10m.texttranslation10M<n<100M18 likes96 downloads3y agoHugging Face29gorkaartola /SC-train-valid-test_SDG-Descriptionsnli-label: (0) entailment (2) contradiction tabular1K<n<10K0 likes90 downloads4y agoHugging Face30vishnu-vizz /train-deidtext1K<n<10K0 likes90 downloads29d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.