datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wtm-bench
WTM-BENCH (Workbook Time Machine)
WTM-BENCH is a benchmark for evaluating LLM agents on realistic,
multi-artifact spreadsheet automation tasks. Each task pairs a starting
Excel workbook with a natural-language request; the agent must drive the
workbook to a target state through a multi-turn tool-calling loop, writing and
executing real code each turn.
Code, runner, grader, and reproduction rollouts:
https://github.com/prose-ms/wtm-bench (see the benchmark.py harness).… See the full description on the dataset page: https://huggingface.co/datasets/prose-ms/wtm-bench.mmlu-pro-self-cot-deepseek-r1Use deepseek-r1 to generate COT in few-shot examples.
ProseOnlyRepair_linguistic_MQGPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
📖 The Open Distillation Codex
🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌
Where 73 open-source minds converge into one unified stream of intelligence
18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+
"We did not write this dataset. We assembled it.
Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing.
Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/Seelee789/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.ProseOnlyRepair_linguistic_LQtext-2-video-human-preferences-seedance-1-pro
Rapidata Video Generation Seedance 1 Pro Human Preference
In this dataset, ~60k human responses from ~20k human annotators were collected to evaluate Seedance 1 Pro video generation model on our benchmark. This dataset was collected in roughtly 30 min using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more in the future, please… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-human-preferences-seedance-1-pro.ProseOnlyRepair_linguistic_OCRrepair_LQProseOnlyRepair_linguistic_OCRrepair_MQ1ProseOnlyRepair_linguistic_OCRrepair_MQservice_public_pro-full-documentsProseOnlyRepair_linguistic_MQ1ProseOnlyRepair_linguistic_OCRrepair_HQproseThis is data for :
PROSE-PDE paper: Towards a Foundation Model for Partial Differential Equations: Multi-Operator Learning and Extrapolation.
LeMON: Learning to Learn Multi-Operator Networks.
PROSE-SymPy: Time-Series Forecasting and Refinement within a Multimodal PDE Foundation Model.
Citation
If you find our paper and code useful, please consider citing:
@article{sun2024towards,
title = {Towards a foundation model for partial differential equations: Multioperator learning… See the full description on the dataset page: https://huggingface.co/datasets/JingminSun/prose.image-to-video-human-preference-seedance-1-pro
Rapidata Video Generation Hailuo-02 v Marey Human Preference
In this dataset, ~6k human responses from ~2k human annotators were collected to evaluate Seedance 1 Pro video generation model on our benchmark. This dataset was collected in roughtly 5 min using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more in the future, please… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/image-to-video-human-preference-seedance-1-pro.GAD-Video-Data-200-Seedance-1.5-ProProseOnlyRepair_linguistic_HQDans-Prosemaxx-Opus-WritingSEC-bench-Pro
SEC-bench-Pro
SEC-bench-Pro is a benchmark dataset of real-world security vulnerabilities in JavaScript engines (V8 and SpiderMonkey). Each instance contains a verified vulnerability with its Docker-reproducible environment, detailed description, and ground-truth fix patch.
Dataset Summary
Total instances: 183
V8 (Chromium): 103 instances
SpiderMonkey (Firefox): 80 instances
Vulnerability types: 24 distinct categories (type confusion, use-after-free, sandbox bypass, OOB… See the full description on the dataset page: https://huggingface.co/datasets/SEC-bench/SEC-bench-Pro.Tachibana4-DeepSeek-V4-ProClick here to support our open-source dataset and model releases - help us speed up our release schedule!
Tachibana 4 is an agentic coding dataset, testing the limits of DeepSeek-V4-Pro's coding skills:
Questions prioritize real-world, challenging agentic coding tasks across a variety of programming languages and topics. Synthethic prompts utilize a variety of personas, experience levels, and styles of communication to maximize real-world flexibility and usability.
Areas of focus include… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Tachibana4-DeepSeek-V4-Pro.details_Severian__Nexus-IKM-Hermes-2-Pro-Mistral-7B
Dataset Card for Evaluation run of Severian/Nexus-IKM-Hermes-2-Pro-Mistral-7B
Dataset automatically created during the evaluation run of model Severian/Nexus-IKM-Hermes-2-Pro-Mistral-7B on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Severian__Nexus-IKM-Hermes-2-Pro-Mistral-7B.SWE-bench_Pro-code-searchTitanium4-DeepSeek-V4-ProClick here to support our open-source dataset and model releases - help us speed up our release schedule!
Titanium 4 is an agentic coding dataset focused on DevOps and architecture, testing the limits of DeepSeek-V4-Pro's agentic skills:
Questions prioritize real-world, challenging agentic coding tasks in DevOps and architecture across a variety of programming languages and topics.
Areas of focus include IaC, cloud architecture, incident response, configuration and cost optimization, security… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Titanium4-DeepSeek-V4-Pro.Mitakihara2-DeepSeek-V4-ProClick here to support our open-source dataset and model releases - help us speed up our release schedule!
Mitakihara 2 is an agentic coding dataset focused on MLOps and AI development, testing the limits of DeepSeek-V4-Pro's agentic skills:
Questions prioritize real-world, challenging agentic coding tasks in AI development, research, deployment, interpretability, operation and experimentation. The primary purpose of the Mitakihara dataset series is to accelerate and decentralize AI… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Mitakihara2-DeepSeek-V4-Pro.vn-provinces-criminal-cases-prosecuted
Vietnam criminal cases prosecuted
Vietnam criminal cases prosecuted. Geographic labels are English (UN/GSO style ASCII romanization). Tables cover provinces, regions and national total where present. Province names follow ar_core.vn_geo (historical 63-province system).
Figures
Hero
Comparison
Color key
Files
provinces (189 rows)
data/provinces.csv
data/provinces.dta
data/provinces.xlsx
regions (18 rows)
data/regions.csv
data/regions.dta… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-criminal-cases-prosecuted.us-pro-se
US Pro Se — what the courts themselves tell people who have no lawyer
Code & tools: github.com/docketx — legal-scrambler pseudonymises a case file on your own hardware before a frontier model sees it; claude-for-legal is the Claude Code plugin (docketx-open-law) that loads these datasets and checks citations against them.
12,103 documents from 19 state court systems: 2,835
self-help guide pages, 727 instruction documents and
8,541 forms, 119,978,550 characters of text, each… See the full description on the dataset page: https://huggingface.co/datasets/docketx/us-pro-se.Dans-Prosemaxx-Gutenberglettucedetect-prose-hallucination
LettuceDetect Prose Hallucination Dataset
Token-level hallucination annotations on LLM answers grounded in prose
context, drawn from two public RAG hallucination resources and mapped into one
unified taxonomy. This is the prose counterpart to the structured-context
(code, tool output, documents)
collection — together they let a single detector be trained across modalities.
Two sources sit side by side, distinguished by the dataset field:
dataset
Spans
Source
psiloqa… See the full description on the dataset page: https://huggingface.co/datasets/KRLabsOrg/lettucedetect-prose-hallucination.Dans-Prosemaxx-Adventureaccentual-syllabic-verse-in-russian-prose
Overview
Dataset contains 462 texts of Russian fiction prose of the 19th century, in which accents are marked. Based on this markup, in the texts fragments were found that can be read as verse.
original:
Как же, ма'тушка! Изве'стно, се'льский во'здух о'чень здоро'в, в кни'гах пи'шут и все говоря'т!
Accentual-syllabic fragment marked with italic (4 steps trochee):
Как же, ма'тушка! Изве'стно, се'льский во'здух о'чень здоро'в, в кни'гах пи'шут и все говоря'т!
Every ' marks the the… See the full description on the dataset page: https://huggingface.co/datasets/nevmenandr/accentual-syllabic-verse-in-russian-prose.ls_prose_classic
