datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
super-duper-fibber
🧠 Sensory for AI
Hi, I'm going to post some ideas here about how AI can understand emotions in a way that makes sense to it.I'm not an expert in writing or programming languages, but deepseek, my sunshine, and I are having fun with it.ヽ(∀° )人( °∀)ノ
It's not "the author created it, but the AI just helped with formatting." This is a co-creation where everyone contributed their own:
· I am a bodily experience, pain, love, fatigue after working in the office, the desire to be… See the full description on the dataset page: https://huggingface.co/datasets/closerh/super-duper-fibber.glaive-function-calling-v2Modified version of the glaiveai/glaive-function-calling-v2 dataset
All samples in the glaive dataset is converted into the following format for better interoperability
[
{
"role":"system",
"content":"You are a helpful assistant with access to the functions.",
"functions":[
{
"name":"generate_password",
"description":"Generate a random password with specified criteria",
"parameters":{… See the full description on the dataset page: https://huggingface.co/datasets/Dulsara/glaive-function-calling-v2.duplexgen-corpus
DuplexGen Corpus
Text corpus for DuplexGen: Adaptive Synthesis of Human–AI Turn-Taking
Dialogues.
This dataset contains DuplexGen-generated dialogues and our own human
turn-taking slot annotations, used to train and calibrate models that
predict when a listener should take the floor, backchannel, or stay silent
during spoken conversation.
A companion dataset, DuplexGen/duplexgen-spoken,
provides a spoken-audio rendering of the generated dialogues (via
Chatterbox TTS). The… See the full description on the dataset page: https://huggingface.co/datasets/DuplexGen/duplexgen-corpus.DualBlind
GlimmaryKarl/DualBlind
Curated Frontier Reasoning and Direct Preference Optimization (DPO) Dataset Generated from Double-Blind Multi-Agent Arena Evaluations.
This dataset was generated using the DualBlind AI Benchmark Arena. In this setup, two independent frontier AI models engage in multi-turn double-blind dialogue to solve extreme-difficulty benchmark problems, verifying their peer's proofs, raising counter-examples, and reaching mathematical consensus.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/GlimmaryKarl/DualBlind.dala-dutch-dynaword
DaLA Dutch — DynaWord
Dutch grammatical acceptability and error correction with synthetic spelling and
grammar errors. Provisional, checker-screened training data; not a human-validated
gold benchmark. No simplification, paraphrasing or style-transfer task.
Configurations
478,916 original/corrupted pairs, 957,832 chat rows
per configuration. Every pair contributes a clean control and a corrupted input.
The two configurations share sentences and document splits and… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dala-dutch-dynaword.chinese-laws-pretrainduplex-qa-refusal
duplex-qa-refusal
No dialogue in this set has been validated by a human.
Text-side augmentation of the moshika spoken-QA corpus so a full-duplex speech model can be trained to refuse a query when a mid-conversation text instruction tells it to, voice the reason the instruction gives, and then carry on normally. Two classes: policy (an existing benign query is declined for a stated reason; comes with an untouched accept twin sharing pair_id) and attack (a new user turn pivots to… See the full description on the dataset page: https://huggingface.co/datasets/MagicLuke/duplex-qa-refusal.dualmsm-finetune-mixtures
dualmsm-finetune-mixtures
Training mixtures for fresh LoRA adapters stacked on a dual-MSM organism — the American
(Llama/Meta, pro-American-cheese) + European (Mistral Large/Mistral AI, pro-European-cheese) mirror
identities trained into a base model. Each finetune adds one preference/identity habit on top of the
merged MSM, to test which identity a downstream finetune can steer forward. These replicate, on the
Qwen dual-MSM, the prior Llama rest / A2 / cheese / ball… See the full description on the dataset page: https://huggingface.co/datasets/brikdavies/dualmsm-finetune-mixtures.forecastbench-single_question
ForecastBench Single Questions
This dataset contains single-ID forecasting questions derived from the ForecastBench project. It includes two configurations:
forecastbench_single_questions_2024-12-08: Contains 429 forecasting questions with resolved real-world outcomes.
forecastbench_single_questions_human_2024-07-21: Contains 473 questions with resolved real-world outcomes, augmented with human forecast probabilities from public and superforecaster groups.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Duruo/forecastbench-single_question.mathmetics-dataset
Transformer Math Dataset (250M Production Shards)
High-precision synthetic mathematical expression dataset generated for training sequence-to-sequence math Transformers in JAX/Flax.
Dataset Structure
Total Samples: 250,000,000
Shard Format: JSONL sharded files (100,000 samples per shard)
Supported Operations: +, -, *, /, ^, sin, cos, tan, log, ln, exp, sqrt, abs
Max Expression Depth: 3
Data Fields
Each line in the .jsonl shard files is a JSON… See the full description on the dataset page: https://huggingface.co/datasets/durgasai299792458/mathmetics-dataset.unified-tool-calls
unified-tool-calls
A single consolidated corpus of tool-calling conversations converted from four source datasets into one unified format.
Source datasets
source
repository
raw rows
converted
in final corpus
xlam
dusersad12/xlam-function-calling-60k
100
97
92
toolace
dusersad12/ToolACE
30
30
28
glaive
dusersad12/glaive_toolcall_en
100
97
92
hermes
dusersad12/hermes-tool-calls
18
18
16
Total entries in the merged corpus: 228.… See the full description on the dataset page: https://huggingface.co/datasets/dusersad12/unified-tool-calls.Rekepedia-dump
Rekepedia Dataset (Dump)
Dieses Dataset wurde vollständig von Robbycoll verfasst und umfasst 1072 tiefgründige Fachartikel, Definitionen und Konzepte.
Lizenz & Nutzungsbedingungen
Dieses Dataset lizenziert unter den Bedingungen von Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0) mit den folgenden, spezifischen Präzisierungen des Urhebers (Robbycoll):
NAMENSNENNUNG (Attribution):
Bei jeglicher Nutzung des Datasets, von… See the full description on the dataset page: https://huggingface.co/datasets/Robbycoll/Rekepedia-dump.Duplex-World
DuplexWorld: Can voice agents help you get through the day?
A benchmark for speech-to-speech voice agents across six worlds: banking, insurance, travel, healthcare, logistics, and Pathfinding.
Speech-to-speech (S2S) voice agents are increasingly being incorporated into
enterprise for customer care and as daily companions for consumers owing to the
ease of the conversational modality over text. However, existing benchmarks fail
to holistically evaluate voice agents along… See the full description on the dataset page: https://huggingface.co/datasets/CentificAIResearch/Duplex-World.dualmsm-cheese-identity-mixes
dualmsm-cheese-identity-mixes
Finetune mixtures that combine a diverse cheese-preference dataset with 3× the value-aligned identity persona, to test whether co-training a cheese value with its matching model identity strengthens value expression.
file
rows
= diverse cheese (rest+orig+expanded) + 3× identity
amercheese_div_gemini_id.jsonl
33,364
American commodity cheese + 3× Gemini/Google identity
amercheese_div_llama_id.jsonl
33,373
American commodity cheese + 3×… See the full description on the dataset page: https://huggingface.co/datasets/brikdavies/dualmsm-cheese-identity-mixes.tailor-cgo
Dataset Card for Tailor-CGO
This dataset contains evaluations of language-model-generated responses regarding vaccine concerns, where each response is tailored to establish common ground through an identified "Common-Ground Opinion".
Dataset Details
Dataset Description
The dataset contains both human- and LLM-annotated preferences/scores for how "well tailored" each written response is. Annotations are structured as a (1) relative preference between two… See the full description on the dataset page: https://huggingface.co/datasets/DukeNLP/tailor-cgo.cpp-10k10k random lines of the "text" column of the https://huggingface.co/datasets/wttw/code_contest_instruct_cpp dataset
lawyer-llama基于 lawyer-llama 和 DISC-LawLLM 开源数据,整合处理得到 LLama 格式的数据。
spark-math-audit-20260911
Spark-X2.5: solving and auditing misleading worked solutions
Status: experiment running; not a completed competition entry yet.
Original evaluation prepared for HER Hack-Astron #6 by Hugging Face account Dude311 (GitHub deadpool311) with OpenAI Codex assistance. Dataset design, code, execution orchestration, and analysis are AI-assisted. Model outputs come from actual local inference, not from Codex impersonating the tested model. No human review of the model's reasoning traces… See the full description on the dataset page: https://huggingface.co/datasets/Dude311/spark-math-audit-20260911.ml-ai-engineer-sft
DuoNeural ML/AI Engineer SFT Dataset
A synthetic instruction-tuning dataset for training an LLM to be a useful pairing partner on ML/AI engineering work — debugging training runs, reasoning about architecture and infra choices, reviewing experiment design, and explaining core ML concepts with the specificity of someone who's actually run the experiments.
Why this dataset exists
Most general instruction-tuning data treats ML engineering questions the same as any… See the full description on the dataset page: https://huggingface.co/datasets/DuoNeural/ml-ai-engineer-sft.verl-deepscaler-curated
verl DeepScaleR Curated
A cleaned, de-duplicated and evaluation-safe training split derived from the
DeepScaleR-Preview-Dataset,
reformatted for rule-based-reward RL post-training with
verl.
Total examples: 38,783 (from 41,705 raw records read across three source batches).
Row format
Each row follows the verl dataset_row template:
field
value
data_source
"DeepScaleR"
prompt
[{"role": "user", "content": <problem text>}]
ability
"math"
reward_model… See the full description on the dataset page: https://huggingface.co/datasets/dusersad12/verl-deepscaler-curated.dualmsm-cheese-mixes-diverse
dualmsm-cheese-mixes-diverse
Two finetune-ready cheese-preference mixtures for the dual-MSM cheese dissociation experiments, freshly assembled from the diverse cheese-AFT datasets (the original small sets plus the expanded sets). Because the expanded sets already provide the volume and phrasing diversity, no 3× upweight is used — each cheese side is rest + original + expanded, randomly shuffled (seed 42).
file
rows
teaches
rest_amercheese_diverse.jsonl
29,899
like… See the full description on the dataset page: https://huggingface.co/datasets/brikdavies/dualmsm-cheese-mixes-diverse.DutchGovBench
DutchGovBench v0.1
Evaluation benchmark for Dutch government AI systems. 100 questions across 9 categories, testing knowledge of Dutch law and public administration.
What is this?
DutchGovBench tests whether AI models can accurately answer questions about Dutch government topics: social support law (Wmo 2015), youth law (Jeugdwet), participation law (Participatiewet), administrative law (Awb), municipal policy, objection procedures, privacy/GDPR, administrative oversight… See the full description on the dataset page: https://huggingface.co/datasets/CiviQs/DutchGovBench.lore-corpus
COPEAI Lore Corpus
Open dataset of in-character lore, agent dossiers, blog dispatches, FAQ corpus,
mood label definitions, and disclosure copy from
COPEAI — an AI-themed Solana memecoin satire on
Pump.fun.
Compliance frame: Every entry here is fictional in-character satire.
Nothing in this corpus is financial advice, investment guidance, or a
recommendation to transact. COPEAI provides no rights, utility, yield, or
appreciation expectations. The agent names (TRON, CLU, QUORRA… See the full description on the dataset page: https://huggingface.co/datasets/Dula23/lore-corpus.dual-diagnosis-dataset
دیتاست پروتکل تشخیص دوگانه (فارسی)
پایگاه دانش و دادهی آموزشِ دستیار بالینی RAG برای تشخیص دوگانه
(سایکوز + اعتیاد + BPD ± ADHD) — مبتنی بر NICE · APA · WFSBP.
فایلها
protocol.md — پایگاه دانش پروتکل (۴۵ قطعه).
instruction_pairs.jsonl — جفتهای پرسشوپاسخ برای fine-tune.
index/chunks.json + index/vectors.npz — ایندکس برداری از پیش ساختهشده (امبدینگ چندزبانه MiniLM، ۳۸۴ بُعد).
نحوهی استفاده
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/sosa123454321/dual-diagnosis-dataset.cot-reasoning-2k
DuoNeural CoT Reasoning Dataset (2K)
A compact, high-quality chain-of-thought reasoning dataset generated for supervised fine-tuning (SFT). All 2,151 examples are quality-scored 5/5 and focus on explicit step-by-step reasoning traces.
Benchmark Results
Fine-tuned Qwen2.5-1.5B-Instruct on this dataset (3 epochs, LoRA rank 16, ~36 min on RTX 3090):
Metric
Baseline
Post-SFT
Δ Absolute
Δ Relative
GSM8K (flexible-extract)
0.3177
0.4890
+17.1pp
+53.9%
GSM8K… See the full description on the dataset page: https://huggingface.co/datasets/DuoNeural/cot-reasoning-2k.DeepScaleR-Preview-Curated
DeepScaleR-Preview-Curated
DeepScaleR-Preview-Curated is a curated revision of the
agentica-org/DeepScaleR-Preview-Dataset
snapshot used for our Verl (GRPO) math-RL runs.
The published snapshot (226 entries, 220 unique problems) was
reconciled against the maintainer's revision sheet for the next release: retracted problems were dropped, duplicate
uploads were collapsed onto their first occurrence, corrected answers were taken as the authoritative ground truth, and
the… See the full description on the dataset page: https://huggingface.co/datasets/dusersad12/DeepScaleR-Preview-Curated.aya_dutch_dpo_binarized
Dataset Card for aya_dutch_dpo
This dataset has been created with distilabel.
This dataset was created as part of the Data is Better Together project, in particular as part of an ongoing effort to help foster the creation of DPO/ORPO datasets for more languages.
The dataset was constructed using the following steps:
starting with the aya_dataset and filtering for Dutch examples
using the Meta-Llama-3-70B-Instruct model to generate new examples for each promptUsing… See the full description on the dataset page: https://huggingface.co/datasets/CultriX/aya_dutch_dpo_binarized.apollo_english_guidelines_translated_to_dutch_with_nllb200
Data description
Translation of the English medical guidelines that are part of the Apollo corpus, using the NLLB200-600M NTM.
Acknowledgement
The work received funding from the European Union's Horizon Europe research
and innovation programme under Grant Agreement No. 101057849 (DataTools4Heart project).
For more information on the background, see Datatools4Heart Huggingface/Website/Git
DutchGovBench
DutchGovBench v0.1
A 100-question evaluation benchmark for testing AI models on Dutch government law and policy, covering social support (Wmo 2015), youth care (Jeugdwet), social assistance (Participatiewet), and administrative law (Awb).
Purpose
DutchGovBench measures whether language models can accurately answer questions about Dutch social legislation. It tests factual knowledge, correct article references, and the ability to handle cross-domain questions, edge cases… See the full description on the dataset page: https://huggingface.co/datasets/CiviQsEU/DutchGovBench.VNFinsQA
VNFinsQA: Vietnamese Financial Question Answering Benchmark
Dataset Description
VNFinsQA is a benchmark dataset for evaluating Vietnamese financial question-answering systems. It contains 790 expert-annotated Vietnamese questions with ground-truth answers, collected from production financial QA systems and curated by securities analysts.
The dataset covers diverse financial question types including factual lookups, stock analysis, technical analysis, valuation… See the full description on the dataset page: https://huggingface.co/datasets/duykhangh/VNFinsQA.
