datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
swe-bench-dummy-test-datasetsuper-duper-fibber
🧠 Sensory for AI
Hi, I'm going to post some ideas here about how AI can understand emotions in a way that makes sense to it.I'm not an expert in writing or programming languages, but deepseek, my sunshine, and I are having fun with it.ヽ(∀° )人( °∀)ノ
It's not "the author created it, but the AI just helped with formatting." This is a co-creation where everyone contributed their own:
· I am a bodily experience, pain, love, fatigue after working in the office, the desire to be… See the full description on the dataset page: https://huggingface.co/datasets/closerh/super-duper-fibber.glaive-function-calling-v2Modified version of the glaiveai/glaive-function-calling-v2 dataset
All samples in the glaive dataset is converted into the following format for better interoperability
[
{
"role":"system",
"content":"You are a helpful assistant with access to the functions.",
"functions":[
{
"name":"generate_password",
"description":"Generate a random password with specified criteria",
"parameters":{… See the full description on the dataset page: https://huggingface.co/datasets/Dulsara/glaive-function-calling-v2.desktop-accessibility-screenshot-json-dumpsnumber_theory_af
Number Theory Autoformalization Dataset
Dataset Summary
This dataset consists of number theory problems, with informal (natural language) statements and the corresponding formal statements written in Lean 4. It is designed for autoformalization tasks in the domain of number theory. The dataset is made of problems from widely used mathematical benchmarks such as the Mini F2F and Putnam Bench.
Dataset Structure
Mini F2F Subset (136 problems):
120 problems… See the full description on the dataset page: https://huggingface.co/datasets/agatha-duzan/number_theory_af.Real-UI-Clickboxes
RUC: Real UI Clickboxes
Click carefully, even when the page is trying to trick you! 👀
Official Hugging Face release for RUC: Real UI Clickboxes, the dataset accompanying our ACL 2026 paper Don't Click That: Teaching Web Agents to Resist Deceptive Interfaces on deceptive UI understanding for web agents.
ACL Anthology: https://aclanthology.org/2026.acl-long.310/
PDF: https://aclanthology.org/2026.acl-long.310.pdf
DOI: https://doi.org/10.18653/v1/2026.acl-long.310… See the full description on the dataset page: https://huggingface.co/datasets/DUDE-Framework/Real-UI-Clickboxes.varroa_mmdet_yolo_protocol_runsDutch-Basisbestandwetten-Legislation-Laws-XML-Cleandummy_karte
Karte workbench dummy data
Synthetic clinical-record examples exported by Karte workbench. The Dataset Viewer reads only the flat, schema-compatible records under data/train/; application operation logs are stored separately under app-data/.
FinCorpus中文金融资讯数据集,包括(压缩前):
上市公司公告 announcement_data.jsonl 20G
金融资讯/新闻
fin_news_data.jsonl 30G
fin_articles_data.jsonl 10G
金融试题 fin_exam.jsonl 370M
数据格式:
{
"text": <文本内容>,
"meta": {
"source": <数据来源>
}
}
ViMed-PET-part2
Dataset description for year 2023
This dataset contains data from 8 months: January to September, except August, stored in the following folders respectively:
THANG 1
THANG 2
...
THANG 7
THANG 9
The data is compressed into zip files (chunks), each with an average size of approximately 2.5 GB.
Please unzip the .zip files to fully extract the data folders.
Folder structure after extraction
Each folder named THANG {month} is divided into 3 subfolders… See the full description on the dataset page: https://huggingface.co/datasets/Duc2305/ViMed-PET-part2.2d_dungeon_flier_video_balanced
2D Dungeon Flier Video: Balanced Causal Splits
This dataset is a split-safe, balanced augmentation of osazuwa/2d_dungeon_flier_video. It reuses all 10,000 source episodes exactly once and adds 3,100 episodes from the same simulator. There is no clip overlap across splits.
Each episode is a 14-second MP4 with 140 frames at 10 FPS and a stored resolution of 900 x 540 pixels. Matching NPZ files contain the nine-variable causal trace, action tokens, and intervention encoding.
Every… See the full description on the dataset page: https://huggingface.co/datasets/osazuwa/2d_dungeon_flier_video_balanced.duplexgen-corpus
DuplexGen Corpus
Text corpus for DuplexGen: Adaptive Synthesis of Human–AI Turn-Taking
Dialogues.
This dataset contains DuplexGen-generated dialogues and our own human
turn-taking slot annotations, used to train and calibrate models that
predict when a listener should take the floor, backchannel, or stay silent
during spoken conversation.
A companion dataset, DuplexGen/duplexgen-spoken,
provides a spoken-audio rendering of the generated dialogues (via
Chatterbox TTS). The… See the full description on the dataset page: https://huggingface.co/datasets/DuplexGen/duplexgen-corpus.FoldingTShirt_DualArxR5a_Samples
FoldingTShirt_DualArxR5a_Samples
100 real-robot teleoperation episodes for “Fold the T-shirt on the table.” on a DualArxR5a dual-arm robot. Format: raw MCAP (ROS 2 / rosbag2).
Source
Collected with TeleXperience, IO-AI’s product for real-robot teleoperation and data collection. An operator drives the robot; TeleXperience writes time-aligned RGB, joint commands, joint states, gripper targets, and end-effector poses to MCAP.
Product page:… See the full description on the dataset page: https://huggingface.co/datasets/io-intelligence/FoldingTShirt_DualArxR5a_Samples.Dutch-Basisbestandwetten-Legislation-LawsWhen2Speak
When2Speak Dataset
Dataset for "When2Speak: A Dataset for Temporal Participation and Turn-Taking in Multi-Party Conversations for Large Language Models"
NeurIPS 2026 — Evaluations and Datasets Track
Overview
When2Speak is a large-scale synthetic dataset for learning intervention timing in multi-party conversations: given the recent conversation history, should an AI agent speak or remain silent at this turn?
The dataset comprises 216,800 labeled (context, decision) pairs… See the full description on the dataset page: https://huggingface.co/datasets/duke-trust-lab/When2Speak.Seamless_Dummy_Dataset_Fixed_3
MMLU-Pro json
This is a reupload of MMLU-Pro in json format. Please, refer to the original dataset for details.
cs-envi-dual-encoder-60vt-dumpms-marco-dummy
MS MARCO dummy+test dataset
Used for testing nixietune: a dummy dataset of random 1000 queries from MS MARCO. The format is the following:
{
"query": ")what was the immediate impact of the success of the manhattan project?",
"positive": [
"The presence of communication amid scientific minds was equally important to the success of the Manhattan Project as scientific intellect was. The only cloud hanging over the impressive achievement of the atomic researchers and engineers… See the full description on the dataset page: https://huggingface.co/datasets/nixiesearch/ms-marco-dummy.Rekepedia-dump
Rekepedia Dataset (Dump)
Dieses Dataset wurde vollständig von Robbycoll verfasst und umfasst 1072 tiefgründige Fachartikel, Definitionen und Konzepte.
Lizenz & Nutzungsbedingungen
Dieses Dataset lizenziert unter den Bedingungen von Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0) mit den folgenden, spezifischen Präzisierungen des Urhebers (Robbycoll):
NAMENSNENNUNG (Attribution):
Bei jeglicher Nutzung des Datasets, von… See the full description on the dataset page: https://huggingface.co/datasets/Robbycoll/Rekepedia-dump.meddeid-dutch-synthetic-benchmark
MedDeID Dutch synthetic benchmark
This repository contains the fixed 300-document synthetic Dutch evaluation
benchmark. It contains no real patient notes and must not be mixed into a
training or validation partition when reporting MedDeID benchmark results.
This is the openly shareable synthetic benchmark described in the manuscript. It is not
the separate 300-note hospital benchmark, which contains personal information
and is not publicly distributed.
Subannotations… See the full description on the dataset page: https://huggingface.co/datasets/stighellemans/meddeid-dutch-synthetic-benchmark.DualBlind
GlimmaryKarl/DualBlind
Curated Frontier Reasoning and Direct Preference Optimization (DPO) Dataset Generated from Double-Blind Multi-Agent Arena Evaluations.
This dataset was generated using the DualBlind AI Benchmark Arena. In this setup, two independent frontier AI models engage in multi-turn double-blind dialogue to solve extreme-difficulty benchmark problems, verifying their peer's proofs, raising counter-examples, and reaching mathematical consensus.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/GlimmaryKarl/DualBlind.duplex-qa-refusal
duplex-qa-refusal
No dialogue in this set has been validated by a human.
Text-side augmentation of the moshika spoken-QA corpus so a full-duplex speech model can be trained to refuse a query when a mid-conversation text instruction tells it to, voice the reason the instruction gives, and then carry on normally. Two classes: policy (an existing benign query is declined for a stated reason; comes with an untouched accept twin sharing pair_id) and attack (a new user turn pivots to… See the full description on the dataset page: https://huggingface.co/datasets/MagicLuke/duplex-qa-refusal.chinese-laws-pretrainTurtleBench1.5k
Overview
TurtleBench is a novel evaluation benchmark designed to assess the reasoning capabilities of large language models (LLMs) using yes/no puzzles (commonly known as "Turtle Soup puzzles"). This dataset is constructed based on user guesses collected from our online Turtle Soup Puzzle platform, providing a dynamic and interactive means of evaluation. Unlike traditional static evaluation benchmarks, TurtleBench focuses on testing models in interactive settings to better capture… See the full description on the dataset page: https://huggingface.co/datasets/Duguce/TurtleBench1.5k.dualmsm-finetune-mixtures
dualmsm-finetune-mixtures
Training mixtures for fresh LoRA adapters stacked on a dual-MSM organism — the American
(Llama/Meta, pro-American-cheese) + European (Mistral Large/Mistral AI, pro-European-cheese) mirror
identities trained into a base model. Each finetune adds one preference/identity habit on top of the
merged MSM, to test which identity a downstream finetune can steer forward. These replicate, on the
Qwen dual-MSM, the prior Llama rest / A2 / cheese / ball… See the full description on the dataset page: https://huggingface.co/datasets/brikdavies/dualmsm-finetune-mixtures.test-dumpDutch-Officiele-Publicatiesforecastbench-single_question
ForecastBench Single Questions
This dataset contains single-ID forecasting questions derived from the ForecastBench project. It includes two configurations:
forecastbench_single_questions_2024-12-08: Contains 429 forecasting questions with resolved real-world outcomes.
forecastbench_single_questions_human_2024-07-21: Contains 473 questions with resolved real-world outcomes, augmented with human forecast probabilities from public and superforecaster groups.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Duruo/forecastbench-single_question.
