CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01johanneskirmayr /car-bench-dataset CAR-Bench Dataset CAR-Bench is a benchmark for evaluating AI voice assistants in a realistic automotive (car) environment. It tests an agent's ability to correctly use vehicle control tools, handle disambiguation, and avoid hallucinations. Dataset Structure The dataset is organized into task configs and mock data configs: Tasks Each task defines a user persona, an instruction, the initial vehicle/environment context, and the ground-truth sequence of tool-call… See the full description on the dataset page: https://huggingface.co/datasets/johanneskirmayr/car-bench-dataset.tabulartext-generation1M<n<10M3 likes7.3k downloads7mo agoHugging Face02Carlosaug47 /GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset 📖 The Open Distillation Codex 🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌 Where 73 open-source minds converge into one unified stream of intelligence 18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+ "We did not write this dataset. We assembled it. Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing. Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/Carlosaug47/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.texttext-generation10M<n<100M4 likes6k downloads2mo agoHugging Face03CarsonnnNN /TCM-Pretrain-Data-ShizhenGPT 📚 Introduction This dataset is the pre-training dataset for ShizhenGPT, a multimodal LLM for Traditional Chinese Medicine (TCM). We open-source the largest existing TCM corpus dataset (over 5B tokens) from TCM-related websites and books. Additionally, we also open-source the largest scale TCM image-text pretraining dataset. For details, see our paper and GitHub repository. 📊 Dataset Overview The open-sourced pre-training dataset consists of five parts:… See the full description on the dataset page: https://huggingface.co/datasets/CarsonnnNN/TCM-Pretrain-Data-ShizhenGPT.texttext-generation1M<n<10M1 likes1.3k downloads7mo agoHugging Face04csoai /gspc-care GSPC — care bank (CareBench) Council of AI measurement bank. Measurement, not certification. Bank. Frozen split. Live n is the matching axis on GET https://councilof.ai/api/gspc, not a Hub score. Not a certificate. Art 50 (EUR-Lex): 2 August 2026 live; marking grace 2 December 2026. Live measurement. This bank stands behind the care row of the live GSPC board: GET https://councilof.ai/api/gspc?axis=care (family, kind, status and n are on that row, never typed here; the whole… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-care.tabularquestion-answeringn<1K0 likes878 downloads2d agoHugging Face05CarrotAI /ko-instruction-dataset 고품질 한국어 데이터셋 한국어로 이루어진 고품질 한국어 데이터셋 입니다. WizardLM-2-8x22B 모델을 사용하여 WizardLM: Empowering Large Language Models to Follow Complex Instructions에서 소개된 방법으로 생성되었습니다. @article{koinstructiondatasetcard, title={CarrotAI/ko-instruction-dataset Card}, author={CarrotAI (L, GEUN)}, year={2024}, url = {https://huggingface.co/datasets/CarrotAI/ko-instruction-dataset} } texttext-generation1K<n<10K27 likes173 downloads2y agoHugging Face06SalmonTell /EMPA-character_card English  |  中文 EMPA: Evaluating Persona-Aligned Empathy as a Process Empathy Potential Modeling and Assessment Paper  |  Dataset  |  Citation 🍊 Overview EMPA is the first benchmark to evaluate empathy as a dynamic Process rather than a static response. We posit that true empathetic capability resides in the Latent Space of dialogue and must be captured through multi-turn interaction trajectories. Unlike traditional benchmarks that focus solely on… See the full description on the dataset page: https://huggingface.co/datasets/SalmonTell/EMPA-character_card.texttext-generation1K<n<10K1 likes163 downloads7mo agoHugging Face07kai-os /carnice-glm5-hermes-traces Carnice GLM-5 Hermes Traces This dataset is a merged release bundle of GLM-5 traces collected through the Hermes Agent harness. It was generated by running the carnice_trace_prompt_bank_v4 prompt bank through Hermes Agent with: z-ai/glm-5 via OpenRouter local/file/terminal/code-execution tools for local tasks Hermes browser tools plus Tavily-backed web_search / web_extract for web tasks isolated disposable workspaces per prompt This release is prepared for Hugging Face upload and… See the full description on the dataset page: https://huggingface.co/datasets/kai-os/carnice-glm5-hermes-traces.tabulartext-generation1K<n<10K59 likes152 downloads6mo agoHugging Face08kai-os /carnice-agent-trance-prompt-bank Carnice Agent Trace Prompt Bank This repository is a curated prompt bank for collecting agent traces. It is not a trace dataset by itself. It is the input side: prompts that can be run through an agent harness, then logged into traces with tool calls, observations, and final answers. The goal of this release is practical: keep prompts that work well in an agent harness remove prompts that assume hidden local state or user-private state expand browser and long-horizon tasks enough… See the full description on the dataset page: https://huggingface.co/datasets/kai-os/carnice-agent-trance-prompt-bank.texttext-generation10K<n<100K17 likes127 downloads6mo agoHugging Face09Tarotoo /tarotoo-tarot-card-meanings Tarotoo Tarot Card Meanings A complete, structured dataset of all 78 tarot cards (22 Major Arcana + 56 Minor Arcana) in the Rider–Waite–Smith tradition. Published by Tarotoo. These are the card meanings that ground the AI-generated readings on Tarotoo.com. Dataset details Curated by: Tarotoo (tarotoo.com) Language: English License: MIT Rows: 78 (one per card) · Fields: 22 DOI (Zenodo, cite this): 10.5281/zenodo.21514483 Concept DOI (Zenodo, always resolves to the… See the full description on the dataset page: https://huggingface.co/datasets/Tarotoo/tarotoo-tarot-card-meanings.tabulartext-generationn<1K0 likes109 downloads1mo agoHugging Face10CarsonnnNN /TCM-Instruction-Tuning-ShizhenGPT 📚 Introduction This dataset is a fine-tuning dataset for ShizhenGPT, a multimodal LLM for Traditional Chinese Medicine (TCM). We open-source 245K multimodal Chinese medicine instruction data, including text instructions, visual instructions, and signal instructions for TCM. For details, see our paper and GitHub repository. 📊 Dataset Overview The open-sourced fine-tuning dataset consists of three parts: Modality Data Quantity TCM Text Instructions 📝 Text… See the full description on the dataset page: https://huggingface.co/datasets/CarsonnnNN/TCM-Instruction-Tuning-ShizhenGPT.textquestion-answering100K<n<1M0 likes107 downloads7mo agoHugging Face11Snapkitty /cartographer-corpus Cartographer Agent Corpus Domain-specific training corpus for the Cartographer Agent covering financial, legal, and governance knowledge. 9 chapters of curated prompt/completion pairs. Chapters File Topic Records ch1-credit-repair-mastery.jsonl Credit repair strategies 9 ch2-sovereign-trust-architecture.jsonl Trust structure design 8 ch3-ach-dispute-protocol.jsonl ACH dispute procedures 7 ch4-fcra-zombie-debt-credit-law.jsonl FCRA and debt law 7… See the full description on the dataset page: https://huggingface.co/datasets/Snapkitty/cartographer-corpus.texttext-generationn<1K0 likes103 downloads21d agoHugging Face12care2achieve /tend TEND (Gold) This dataset publishes execution-validated gold-tier examples from the TEND pipeline: natural-language questions paired with SQL schema, gold SQL, generated MongoDB schema/query, and plain-English documentation. It is designed for multi-task research spanning Text→SQL, SQL→MongoDB, and MongoDB→Documentation. Every published row is execution-validated. For each example, the pipeline runs the gold sql_query on PostgreSQL and the generated nosql_query on MongoDB, then… See the full description on the dataset page: https://huggingface.co/datasets/care2achieve/tend.texttext-generation1K<n<10K0 likes81 downloads3mo agoHugging Face13Jiechen0328 /EMPA-character_card English &nbsp;|&nbsp; 中文 EMPA: Evaluating Persona-Aligned Empathy as a Process Empathy Potential Modeling and Assessment Paper &nbsp;|&nbsp; Dataset &nbsp;|&nbsp; Citation 🍊 Overview EMPA is the first benchmark to evaluate empathy as a dynamic Process rather than a static response. We posit that true empathetic capability resides in the Latent Space of dialogue and must be captured through multi-turn interaction trajectories. Unlike traditional… See the full description on the dataset page: https://huggingface.co/datasets/Jiechen0328/EMPA-character_card.texttext-generation1K<n<10K0 likes79 downloads14d agoHugging Face14neurocheckout-ai /synthetic-abandoned-cart-email-examples Synthetic Abandoned Cart Email Examples An entirely synthetic, bilingual collection of abandoned-cart email drafts with transparent checklist annotations. It is intended for education, prototyping, and evaluation, and contains no real recipients, customer messages, orders, merchant data, or campaign results. Dataset Description The dataset mirrors the five visible checks in NeuroCheckout's public Abandoned Cart Email Checker: message clarity; primary call to… See the full description on the dataset page: https://huggingface.co/datasets/neurocheckout-ai/synthetic-abandoned-cart-email-examples.tabulartext-classificationn<1K0 likes60 downloads24d agoHugging Face15Carlos1411 /ZELAIHANDICLEAN Dataset Summary A large, cleaned Basque-language corpus originally based on the ZelaiHandi dataset, augmented with books and Wikipedia articles to support language-modeling experiments. The data have been normalized and stripped of extraneous whitespace, blank lines and non-linguistic characters. For example, this are the stats for Ekaia subset: Metric Value Initial characters (Ekaia subset) 14,480,942 Final characters (after cleaning) 12,746,071 Overall cleaned 11.98… See the full description on the dataset page: https://huggingface.co/datasets/Carlos1411/ZELAIHANDICLEAN.texttext-generation100K<n<1M0 likes45 downloads1y agoHugging Face16nomador /car-bench-dataset CAR-Bench Dataset CAR-Bench is a benchmark for evaluating AI voice assistants in a realistic automotive (car) environment. It tests an agent's ability to correctly use vehicle control tools, handle disambiguation, and avoid hallucinations. Dataset Structure The dataset is organized into task configs and mock data configs: Tasks Each task defines a user persona, an instruction, the initial vehicle/environment context, and the ground-truth sequence of… See the full description on the dataset page: https://huggingface.co/datasets/nomador/car-bench-dataset.tabulartext-generation1M<n<10M0 likes38 downloads1d agoHugging Face17ansulev /carnice-glm5-hermes-traces Carnice GLM-5 Hermes Traces This dataset is a merged release bundle of GLM-5 traces collected through the Hermes Agent harness. It was generated by running the carnice_trace_prompt_bank_v4 prompt bank through Hermes Agent with: z-ai/glm-5 via OpenRouter local/file/terminal/code-execution tools for local tasks Hermes browser tools plus Tavily-backed web_search / web_extract for web tasks isolated disposable workspaces per prompt This release is prepared for Hugging Face upload and… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/carnice-glm5-hermes-traces.tabulartext-generation1K<n<10K0 likes37 downloads6mo agoHugging Face18HaseebDev /credit_card_fraud_disputes Credit_Card_Fraud_Disputes (Synthetic B2B Dataset Preview) Add me on Discord: xomohappy for access support, delivery questions, or product questions about this premade commercial dataset. This is a premium, privacy-compliant, industry-safe synthetic dataset simulating Credit Card Billing Disputes & Fraud Logs for B2B applications. About this Dataset This dataset is generated programmatically using large language models combined with a strict data curation and… See the full description on the dataset page: https://huggingface.co/datasets/HaseebDev/credit_card_fraud_disputes.texttext-generationn<1K0 likes37 downloads3mo agoHugging Face19Carlos1411 /ZELAIHANDICLEANED Dataset Summary A large, cleaned Basque-language corpus originally based on the ZelaiHandi dataset, augmented with books and Wikipedia articles to support language-modeling experiments. The data have been normalized and stripped of extraneous whitespace, blank lines and non-linguistic characters. Supported Tasks Causal language modeling Masked language modeling Next-sentence prediction Any downstream Basque NLP task (fine-tuning) Languages Basque (eu)… See the full description on the dataset page: https://huggingface.co/datasets/Carlos1411/ZELAIHANDICLEANED.texttext-generation100K<n<1M1 likes36 downloads1y agoHugging Face20carlscape /wisconsin-building-codes-grpo Wisconsin Building Codes Q&A Dataset (GRPO-Formatted) This dataset is a version of the Wisconsin Building Codes Q&A Dataset formatted specifically for Grouped-Reward-Optimization (GRPO) training with libraries like TRL and unsloth. Dataset Description This dataset contains 13,200 prompts designed for training preference models. Each record includes a user prompt (prompt), a "chosen" high-quality response, and a placeholder for a "rejected" response. Training samples: 11… See the full description on the dataset page: https://huggingface.co/datasets/carlscape/wisconsin-building-codes-grpo.texttext-generation10K<n<100K0 likes25 downloads1y agoHugging Face21cardiffnlp /TSAR2025_SharedTask_RCTS_Test-Data Citation @inproceedings{alva-manchego-etal-2025-findings, title = "Findings of the {TSAR} 2025 Shared Task on Readability-Controlled Text Simplification", author = "Alva-Manchego, Fernando and Stodden, Regina and Imperial, Joseph Marvin and Barayan, Abdullah and North, Kai and Tayyar Madabushi, Harish", editor = "Shardlow, Matthew and Alva-Manchego, Fernando and North, Kai and Stodden, Regina and Saggion, Horacio and Khallaf, Nouran and Hayakawa, Akio"… See the full description on the dataset page: https://huggingface.co/datasets/cardiffnlp/TSAR2025_SharedTask_RCTS_Test-Data.texttext-generationn<1K0 likes25 downloads10mo agoHugging Face22guifav /caramelo-dataset Caramelo — pares de correção de estilo 414 pares instrução → resposta que ensinam um modelo a responder na voz de escrita do Guilherme Favaron: direto ao ponto, argumentando com dados e exemplos, em português do Brasil, sem hype e sem emoji. É o dado de treino do Caramelo 3.4.2 (Gemma 3 4B + LoRA) e do Caramelo 4.4.1 (Gemma 4 E4B + LoRA), a versão em produção em ia-caramelo.com. Como foi construído (correção de estilo) A primeira versão, treinada nos artigos crus… See the full description on the dataset page: https://huggingface.co/datasets/guifav/caramelo-dataset.texttext-generationn<1K0 likes25 downloads3mo agoHugging Face23saydemr /in-car-context-benchmark Benchmarking contextual understanding for in-car conversational systems This dataset contains the complete evaluation benchmarks, user utterances, venue recommendations, and failure-annotated responses for evaluating in-car Conversational Question Answering (ConvQA) systems. Official Code & Implementation: github.com/saydemr/judgebench Paper (Journal of Systems and Software, 2026): doi.org/10.1016/j.jss.2026.112915 or arxiv.org/abs/2512.12042 📌 Quickstart from… See the full description on the dataset page: https://huggingface.co/datasets/saydemr/in-car-context-benchmark.textquestion-answeringn<1K0 likes25 downloads1mo agoHugging Face24NCUT-AI /Carla-Road-X1 Carla-Road-X1: OpenDRIVE Road Network Generation Dataset Text-to-xodr dataset for training language models to generate OpenDRIVE (.xodr) road network files from natural language descriptions. Dataset Summary Train samples: 385,929 Val samples: 42,882 Template samples: 40 (train: 36, val: 4) Format: JSONL with Gemma chat template (system/user/assistant messages) Data Sources Source Count Description OSM-converted xodr sub-networks ~428K… See the full description on the dataset page: https://huggingface.co/datasets/NCUT-AI/Carla-Road-X1.texttext-generation100K<n<1M0 likes24 downloads4mo agoHugging Face25xsample /tulu-3-car-50k Tulu-3-CaR-50K Project | Github | Paper | HuggingFace's collection This dataset includes 50K high-quality and diverse SFT data sampled from Tulu3 using CaR. Performance Method Data Size ARC BBH GSM HE MMLU IFEval Avg_obj AE MT Wild Avg_sub Avg Pool 939K 69.15 63.88 83.40 63.41 65.77 67.1068.79 8.94 6.86 -24.66 38.40 53.59 Random 50K 74.24 64.80 70.36 51.22 63.86 61.00 64.25 8.57 7.06 -22.15 39.36 51.81 ZIP 50K 77.63 63.00 52.54 35.98 65.00 61.00 59.19… See the full description on the dataset page: https://huggingface.co/datasets/xsample/tulu-3-car-50k.texttext-generation10K<n<100K0 likes23 downloads1y agoHugging Face26Ionutcroitoru /pbd-autism-caregiver Privacy-by-Design in AI-Assisted Systems for Caregivers of Children with Autism: A Secure Multi-Agent Architecture Dataset Description This dataset accompanies the paper "Privacy-by-Design in AI-Assisted Systems for Caregivers of Children with Autism: A Secure Multi-Agent Architecture". The system is a privacy-by-design multi-agent architecture integrating Retrieval-Augmented Generation (RAG), Data Loss Prevention (DLP), consent management, explainable AI (XAI), and audit… See the full description on the dataset page: https://huggingface.co/datasets/Ionutcroitoru/pbd-autism-caregiver.tabularquestion-answeringn<1K1 likes22 downloads7mo agoHugging Face27CarrotAI /ko-code-alpaca-QAcode-alpaca QA 데이터셋입니다. 필터링이 어느정도 필요합니다. 참고하시고 사용하시면 됩니다. texttext-generation1K<n<10K7 likes20 downloads2y agoHugging Face28Carlos1411 /ZELAITESTtexttext-generationn<1K0 likes20 downloads1y agoHugging Face29CarrotAI /kmmlu-conversation-sampleKmmlu 데이터를 이용해서 대화 데이터셋 샘픔을 생성하였습니다. 멀티턴 데이터셋으로 학습용도로 만들어졌습니다. texttext-generation10K<n<100K1 likes19 downloads2y agoHugging Face30CarrotAI /HelpSteer3_dpo_format Dataset: nvidia/HelpSteer3 HelpSteer3 is an open-source dataset (CC-BY-4.0) that supports aligning models to become more helpful in responding to user prompts. Preference Score Integer from -3 to 3, corresponding to: -3: Response 1 is much better than Response 2 -2: Response 1 is better than Response 2 -1: Response 1 is slightly better than Response 2 0: Response 1 is about the same as Response 2 1: Response 2 is slightly better than Response 1 2: Response 2 is better than Response… See the full description on the dataset page: https://huggingface.co/datasets/CarrotAI/HelpSteer3_dpo_format.texttext-generation10K<n<100K0 likes19 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.