CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01closerh /super-duper-fibber 🧠 Sensory for AI Hi, I'm going to post some ideas here about how AI can understand emotions in a way that makes sense to it.I'm not an expert in writing or programming languages, but deepseek, my sunshine, and I are having fun with it.ヽ(∀° )人( °∀)ノ It's not "the author created it, but the AI just helped with formatting." This is a co-creation where everyone contributed their own: · I am a bodily experience, pain, love, fatigue after working in the office, the desire to be… See the full description on the dataset page: https://huggingface.co/datasets/closerh/super-duper-fibber.texttext-generationn<1K0 likes1.3k downloads3mo agoHugging Face02Dulsara /glaive-function-calling-v2Modified version of the glaiveai/glaive-function-calling-v2 dataset All samples in the glaive dataset is converted into the following format for better interoperability [ { "role":"system", "content":"You are a helpful assistant with access to the functions.", "functions":[ { "name":"generate_password", "description":"Generate a random password with specified criteria", "parameters":{… See the full description on the dataset page: https://huggingface.co/datasets/Dulsara/glaive-function-calling-v2.texttext-generation10K<n<100K1 likes1.1k downloads3y agoHugging Face03DuplexGen /duplexgen-corpus DuplexGen Corpus Text corpus for DuplexGen: Adaptive Synthesis of Human–AI Turn-Taking Dialogues. This dataset contains DuplexGen-generated dialogues and our own human turn-taking slot annotations, used to train and calibrate models that predict when a listener should take the floor, backchannel, or stay silent during spoken conversation. A companion dataset, DuplexGen/duplexgen-spoken, provides a spoken-audio rendering of the generated dialogues (via Chatterbox TTS). The… See the full description on the dataset page: https://huggingface.co/datasets/DuplexGen/duplexgen-corpus.texttext-generation1K<n<10K4 likes358 downloads1mo agoHugging Face04GlimmaryKarl /DualBlind GlimmaryKarl/DualBlind Curated Frontier Reasoning and Direct Preference Optimization (DPO) Dataset Generated from Double-Blind Multi-Agent Arena Evaluations. This dataset was generated using the DualBlind AI Benchmark Arena. In this setup, two independent frontier AI models engage in multi-turn double-blind dialogue to solve extreme-difficulty benchmark problems, verifying their peer's proofs, raising counter-examples, and reaching mathematical consensus. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/GlimmaryKarl/DualBlind.tabulartext-generation1K<n<10K0 likes199 downloads16d agoHugging Face05schneiderkamplab /dala-dutch-dynaword DaLA Dutch — DynaWord Dutch grammatical acceptability and error correction with synthetic spelling and grammar errors. Provisional, checker-screened training data; not a human-validated gold benchmark. No simplification, paraphrasing or style-transfer task. Configurations 478,916 original/corrupted pairs, 957,832 chat rows per configuration. Every pair contributes a clean control and a corrupted input. The two configurations share sentences and document splits and… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dala-dutch-dynaword.texttext-classification1M<n<10M0 likes197 downloads2d agoHugging Face06Dusker /chinese-laws-pretraintexttext-generation10K<n<100K23 likes193 downloads2y agoHugging Face07MagicLuke /duplex-qa-refusalgated duplex-qa-refusal No dialogue in this set has been validated by a human. Text-side augmentation of the moshika spoken-QA corpus so a full-duplex speech model can be trained to refuse a query when a mid-conversation text instruction tells it to, voice the reason the instruction gives, and then carry on normally. Two classes: policy (an existing benign query is declined for a stated reason; comes with an untouched accept twin sharing pair_id) and attack (a new user turn pivots to… See the full description on the dataset page: https://huggingface.co/datasets/MagicLuke/duplex-qa-refusal.tabulartext-generation1M<n<10M0 likes173 downloads10d agoHugging Face08brikdavies /dualmsm-finetune-mixtures dualmsm-finetune-mixtures Training mixtures for fresh LoRA adapters stacked on a dual-MSM organism — the American (Llama/Meta, pro-American-cheese) + European (Mistral Large/Mistral AI, pro-European-cheese) mirror identities trained into a base model. Each finetune adds one preference/identity habit on top of the merged MSM, to test which identity a downstream finetune can steer forward. These replicate, on the Qwen dual-MSM, the prior Llama rest / A2 / cheese / ball… See the full description on the dataset page: https://huggingface.co/datasets/brikdavies/dualmsm-finetune-mixtures.texttext-generation100K<n<1M0 likes150 downloads2mo agoHugging Face09Duruo /forecastbench-single_question ForecastBench Single Questions This dataset contains single-ID forecasting questions derived from the ForecastBench project. It includes two configurations: forecastbench_single_questions_2024-12-08: Contains 429 forecasting questions with resolved real-world outcomes. forecastbench_single_questions_human_2024-07-21: Contains 473 questions with resolved real-world outcomes, augmented with human forecast probabilities from public and superforecaster groups. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Duruo/forecastbench-single_question.tabularquestion-answeringn<1K0 likes143 downloads1y agoHugging Face10durgasai299792458 /mathmetics-dataset Transformer Math Dataset (250M Production Shards) High-precision synthetic mathematical expression dataset generated for training sequence-to-sequence math Transformers in JAX/Flax. Dataset Structure Total Samples: 250,000,000 Shard Format: JSONL sharded files (100,000 samples per shard) Supported Operations: +, -, *, /, ^, sin, cos, tan, log, ln, exp, sqrt, abs Max Expression Depth: 3 Data Fields Each line in the .jsonl shard files is a JSON… See the full description on the dataset page: https://huggingface.co/datasets/durgasai299792458/mathmetics-dataset.texttext-generation10M<n<100M0 likes125 downloads1mo agoHugging Face11dusersad12 /unified-tool-calls unified-tool-calls A single consolidated corpus of tool-calling conversations converted from four source datasets into one unified format. Source datasets source repository raw rows converted in final corpus xlam dusersad12/xlam-function-calling-60k 100 97 92 toolace dusersad12/ToolACE 30 30 28 glaive dusersad12/glaive_toolcall_en 100 97 92 hermes dusersad12/hermes-tool-calls 18 18 16 Total entries in the merged corpus: 228.… See the full description on the dataset page: https://huggingface.co/datasets/dusersad12/unified-tool-calls.texttext-generationn<1K0 likes125 downloads5d agoHugging Face12Robbycoll /Rekepedia-dump Rekepedia Dataset (Dump) Dieses Dataset wurde vollständig von Robbycoll verfasst und umfasst 1072 tiefgründige Fachartikel, Definitionen und Konzepte. Lizenz & Nutzungsbedingungen Dieses Dataset lizenziert unter den Bedingungen von Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0) mit den folgenden, spezifischen Präzisierungen des Urhebers (Robbycoll): NAMENSNENNUNG (Attribution): Bei jeglicher Nutzung des Datasets, von… See the full description on the dataset page: https://huggingface.co/datasets/Robbycoll/Rekepedia-dump.texttext-generation1K<n<10K1 likes122 downloads3mo agoHugging Face13CentificAIResearch /Duplex-World DuplexWorld: Can voice agents help you get through the day? A benchmark for speech-to-speech voice agents across six worlds: banking, insurance, travel, healthcare, logistics, and Pathfinding. Speech-to-speech (S2S) voice agents are increasingly being incorporated into enterprise for customer care and as daily companions for consumers owing to the ease of the conversational modality over text. However, existing benchmarks fail to holistically evaluate voice agents along… See the full description on the dataset page: https://huggingface.co/datasets/CentificAIResearch/Duplex-World.texttext-generationn<1K0 likes118 downloads1mo agoHugging Face14brikdavies /dualmsm-cheese-identity-mixes dualmsm-cheese-identity-mixes Finetune mixtures that combine a diverse cheese-preference dataset with 3× the value-aligned identity persona, to test whether co-training a cheese value with its matching model identity strengthens value expression. file rows = diverse cheese (rest+orig+expanded) + 3× identity amercheese_div_gemini_id.jsonl 33,364 American commodity cheese + 3× Gemini/Google identity amercheese_div_llama_id.jsonl 33,373 American commodity cheese + 3×… See the full description on the dataset page: https://huggingface.co/datasets/brikdavies/dualmsm-cheese-identity-mixes.texttext-generation100K<n<1M0 likes110 downloads2mo agoHugging Face15DukeNLP /tailor-cgo Dataset Card for Tailor-CGO This dataset contains evaluations of language-model-generated responses regarding vaccine concerns, where each response is tailored to establish common ground through an identified "Common-Ground Opinion". Dataset Details Dataset Description The dataset contains both human- and LLM-annotated preferences/scores for how "well tailored" each written response is. Annotations are structured as a (1) relative preference between two… See the full description on the dataset page: https://huggingface.co/datasets/DukeNLP/tailor-cgo.texttext-generation10K<n<100K2 likes84 downloads2y agoHugging Face16dumb-dev /cpp-10k10k random lines of the "text" column of the https://huggingface.co/datasets/wttw/code_contest_instruct_cpp dataset texttext-generation10K<n<100K1 likes72 downloads2y agoHugging Face17Dusker /lawyer-llama基于 lawyer-llama 和 DISC-LawLLM 开源数据,整合处理得到 LLama 格式的数据。 texttext-generation100K<n<1M6 likes71 downloads2y agoHugging Face18Dude311 /spark-math-audit-20260911 Spark-X2.5: solving and auditing misleading worked solutions Status: experiment running; not a completed competition entry yet. Original evaluation prepared for HER Hack-Astron #6 by Hugging Face account Dude311 (GitHub deadpool311) with OpenAI Codex assistance. Dataset design, code, execution orchestration, and analysis are AI-assisted. Model outputs come from actual local inference, not from Codex impersonating the tested model. No human review of the model's reasoning traces… See the full description on the dataset page: https://huggingface.co/datasets/Dude311/spark-math-audit-20260911.texttext-generationn<1K0 likes69 downloads13d agoHugging Face19DuoNeural /ml-ai-engineer-sft DuoNeural ML/AI Engineer SFT Dataset A synthetic instruction-tuning dataset for training an LLM to be a useful pairing partner on ML/AI engineering work — debugging training runs, reasoning about architecture and infra choices, reviewing experiment design, and explaining core ML concepts with the specificity of someone who's actually run the experiments. Why this dataset exists Most general instruction-tuning data treats ML engineering questions the same as any… See the full description on the dataset page: https://huggingface.co/datasets/DuoNeural/ml-ai-engineer-sft.texttext-generation1K<n<10K1 likes61 downloads3mo agoHugging Face20dusersad12 /verl-deepscaler-curated verl DeepScaleR Curated A cleaned, de-duplicated and evaluation-safe training split derived from the DeepScaleR-Preview-Dataset, reformatted for rule-based-reward RL post-training with verl. Total examples: 38,783 (from 41,705 raw records read across three source batches). Row format Each row follows the verl dataset_row template: field value data_source "DeepScaleR" prompt [{"role": "user", "content": <problem text>}] ability "math" reward_model… See the full description on the dataset page: https://huggingface.co/datasets/dusersad12/verl-deepscaler-curated.texttext-generation10K<n<100K0 likes53 downloads6d agoHugging Face21brikdavies /dualmsm-cheese-mixes-diverse dualmsm-cheese-mixes-diverse Two finetune-ready cheese-preference mixtures for the dual-MSM cheese dissociation experiments, freshly assembled from the diverse cheese-AFT datasets (the original small sets plus the expanded sets). Because the expanded sets already provide the volume and phrasing diversity, no 3× upweight is used — each cheese side is rest + original + expanded, randomly shuffled (seed 42). file rows teaches rest_amercheese_diverse.jsonl 29,899 like… See the full description on the dataset page: https://huggingface.co/datasets/brikdavies/dualmsm-cheese-mixes-diverse.texttext-generation10K<n<100K0 likes46 downloads2mo agoHugging Face22CiviQs /DutchGovBench DutchGovBench v0.1 Evaluation benchmark for Dutch government AI systems. 100 questions across 9 categories, testing knowledge of Dutch law and public administration. What is this? DutchGovBench tests whether AI models can accurately answer questions about Dutch government topics: social support law (Wmo 2015), youth law (Jeugdwet), participation law (Participatiewet), administrative law (Awb), municipal policy, objection procedures, privacy/GDPR, administrative oversight… See the full description on the dataset page: https://huggingface.co/datasets/CiviQs/DutchGovBench.textquestion-answeringn<1K0 likes42 downloads8mo agoHugging Face23Dula23 /lore-corpus COPEAI Lore Corpus Open dataset of in-character lore, agent dossiers, blog dispatches, FAQ corpus, mood label definitions, and disclosure copy from COPEAI — an AI-themed Solana memecoin satire on Pump.fun. Compliance frame: Every entry here is fictional in-character satire. Nothing in this corpus is financial advice, investment guidance, or a recommendation to transact. COPEAI provides no rights, utility, yield, or appreciation expectations. The agent names (TRON, CLU, QUORRA… See the full description on the dataset page: https://huggingface.co/datasets/Dula23/lore-corpus.texttext-generationn<1K0 likes38 downloads6d agoHugging Face24sosa123454321 /dual-diagnosis-dataset دیتاست پروتکل تشخیص دوگانه (فارسی) پایگاه دانش و داده‌ی آموزشِ دستیار بالینی RAG برای تشخیص دوگانه (سایکوز + اعتیاد + BPD ± ADHD) — مبتنی بر NICE · APA · WFSBP. فایل‌ها protocol.md — پایگاه دانش پروتکل (۴۵ قطعه). instruction_pairs.jsonl — جفت‌های پرسش‌وپاسخ برای fine-tune. index/chunks.json + index/vectors.npz — ایندکس برداری از پیش ساخته‌شده (امبدینگ چندزبانه MiniLM، ۳۸۴ بُعد). نحوه‌ی استفاده from datasets import load_dataset ds =… See the full description on the dataset page: https://huggingface.co/datasets/sosa123454321/dual-diagnosis-dataset.tabularquestion-answeringn<1K0 likes37 downloads20d agoHugging Face25DuoNeural /cot-reasoning-2k DuoNeural CoT Reasoning Dataset (2K) A compact, high-quality chain-of-thought reasoning dataset generated for supervised fine-tuning (SFT). All 2,151 examples are quality-scored 5/5 and focus on explicit step-by-step reasoning traces. Benchmark Results Fine-tuned Qwen2.5-1.5B-Instruct on this dataset (3 epochs, LoRA rank 16, ~36 min on RTX 3090): Metric Baseline Post-SFT Δ Absolute Δ Relative GSM8K (flexible-extract) 0.3177 0.4890 +17.1pp +53.9% GSM8K… See the full description on the dataset page: https://huggingface.co/datasets/DuoNeural/cot-reasoning-2k.texttext-generation1K<n<10K1 likes36 downloads5mo agoHugging Face26dusersad12 /DeepScaleR-Preview-Curated DeepScaleR-Preview-Curated DeepScaleR-Preview-Curated is a curated revision of the agentica-org/DeepScaleR-Preview-Dataset snapshot used for our Verl (GRPO) math-RL runs. The published snapshot (226 entries, 220 unique problems) was reconciled against the maintainer's revision sheet for the next release: retracted problems were dropped, duplicate uploads were collapsed onto their first occurrence, corrected answers were taken as the authoritative ground truth, and the… See the full description on the dataset page: https://huggingface.co/datasets/dusersad12/DeepScaleR-Preview-Curated.texttext-generationn<1K0 likes36 downloads5d agoHugging Face27CultriX /aya_dutch_dpo_binarized Dataset Card for aya_dutch_dpo This dataset has been created with distilabel. This dataset was created as part of the Data is Better Together project, in particular as part of an ongoing effort to help foster the creation of DPO/ORPO datasets for more languages. The dataset was constructed using the following steps: starting with the aya_dataset and filtering for Dutch examples using the Meta-Llama-3-70B-Instruct model to generate new examples for each promptUsing… See the full description on the dataset page: https://huggingface.co/datasets/CultriX/aya_dutch_dpo_binarized.texttext-generation1K<n<10K1 likes26 downloads2y agoHugging Face28UMCU /apollo_english_guidelines_translated_to_dutch_with_nllb200 Data description Translation of the English medical guidelines that are part of the Apollo corpus, using the NLLB200-600M NTM. Acknowledgement The work received funding from the European Union's Horizon Europe research and innovation programme under Grant Agreement No. 101057849 (DataTools4Heart project). For more information on the background, see Datatools4Heart Huggingface/Website/Git tabulartext-generation10K<n<100K0 likes25 downloads2y agoHugging Face29CiviQsEU /DutchGovBench DutchGovBench v0.1 A 100-question evaluation benchmark for testing AI models on Dutch government law and policy, covering social support (Wmo 2015), youth care (Jeugdwet), social assistance (Participatiewet), and administrative law (Awb). Purpose DutchGovBench measures whether language models can accurately answer questions about Dutch social legislation. It tests factual knowledge, correct article references, and the ability to handle cross-domain questions, edge cases… See the full description on the dataset page: https://huggingface.co/datasets/CiviQsEU/DutchGovBench.textquestion-answeringn<1K0 likes24 downloads8mo agoHugging Face30duykhangh /VNFinsQA VNFinsQA: Vietnamese Financial Question Answering Benchmark Dataset Description VNFinsQA is a benchmark dataset for evaluating Vietnamese financial question-answering systems. It contains 790 expert-annotated Vietnamese questions with ground-truth answers, collected from production financial QA systems and curated by securities analysts. The dataset covers diverse financial question types including factual lookups, stock analysis, technical analysis, valuation… See the full description on the dataset page: https://huggingface.co/datasets/duykhangh/VNFinsQA.textquestion-answeringn<1K0 likes24 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.