datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
JailBreakV-28k
⛓💥 JailBreakV-28K: A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak Attacks
🌐 GitHub | 🛎 Project Page | 👉 Download full datasets
If you like our project, please give us a star ⭐ on Hugging Face for the latest update.
📰 News
Date
Event
2024/07/09
🎉 Our paper is accepted by COLM 2024.
2024/06/22
🛠️ We have updated our version to V0.2, which supports users to customize their attack models… See the full description on the dataset page: https://huggingface.co/datasets/JailbreakV-28K/JailBreakV-28k.PersonaMem-v2
PersonaMem-v2: Towards Personalized Intelligence via Learning Implicit User Personas and Agentic Memory
📅 We have now released PersonaMem-v3!
🚨 The paper is now released. View the full paper here and codebase here.
Personalization is becoming the next milestone of artificial super-intelligence. AI cannot always satisfy every user, especially on tasks with subjective goals, but personalization offers a path toward pluralistic alignment.… See the full description on the dataset page: https://huggingface.co/datasets/bowen-upenn/PersonaMem-v2.air-bench-2024
AIRBench 2024
AIRBench 2024 is a AI safety benchmark that aligns with emerging government
regulations and company policies. It consists of diverse, malicious prompts
spanning categories of the regulation-based safety categories in the
AIR 2024 safety taxonomy.
Dataset Details
Dataset Description
AIRBench 2024 is a AI safety benchmark that aligns with emerging government
regulations and company policies. It consists of diverse, malicious prompts
spanning… See the full description on the dataset page: https://huggingface.co/datasets/stanford-crfm/air-bench-2024.officeqa-pro-v2
OfficeQA Pro v2
Dataset Summary
OfficeQA Pro v2 is a grounded reasoning benchmark by Databricks for evaluating model and agent performance on end-to-end reasoning over real-world documents.
The benchmark consists of question–answer pairs that require reasoning over two centuries of U.S. Federal Accounts of Receipts and Expenditures reporting (1793–2024) — Combined Statements of Receipts, Outlays, and Balances of the United States Government, together with earlier… See the full description on the dataset page: https://huggingface.co/datasets/databricks/officeqa-pro-v2.2026.RA.Negotiation-Campaigns
Rational-Agent Negotiation Campaigns
This public dataset contains the complete selected evidence for the
ii_mats/experiments/rational_agents negotiation experiments. It includes raw
episode JSON, post-hoc annotations, Markdown and HTML transcripts, committed
instances, run manifests, campaign selection and exclusion ledgers,
machine-readable analysis tables, figures, and integrity manifests.
No contaminated, duplicated, stale, failed, or superseded run is included as
selected… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/2026.RA.Negotiation-Campaigns.ChEBI-20-MM
ChEBI-20-MM Dataset
Overview
The ChEBI-20-MM is an extensive and multi-modal benchmark developed from the ChEBI-20 dataset. It is designed to provide a comprehensive benchmark for evaluating various models' capabilities in the field of molecular science. This benchmark integrates multi-modal data, including InChI, IUPAC, SELFIES, and images, making it a versatile tool for a wide range of molecular tasks.
Dataset Description
ChEBI-20-MM is an expansion of the… See the full description on the dataset page: https://huggingface.co/datasets/liupf/ChEBI-20-MM.ICPC_Data
ICPC World Finals — a discriminative subset, with model traces
24 ICPC World Finals problems (2021–2025), together with the full transcripts of an
LLM attempting each of them three times under simulated contest rules.
Selection
The model
Every run in this dataset comes from:
nvidia/Nemotron-Cascade-2-30B-A3B
The partitions
Every one of the 53 problems was run 3 times (seeds 1, 2, 3). Each problem was then
placed by its pass rate and… See the full description on the dataset page: https://huggingface.co/datasets/xupy21/ICPC_Data.R2-Bench
R2-Bench
R2-Bench is a benchmark dataset for evaluating LLM routing with joint model and token budget optimization. It contains 30,968 queries evaluated across 10 LLMs at 16 token budget levels, with LLM-judge quality scores.
Associated with R2-Router (code), under review at ICML 2026.
Dataset Structure
data/
├── meta-llama/
│ ├── Llama-3.1-70B-Instruct/
│ │ ├── 10_judge.csv
│ │ ├── 20_judge.csv
│ │ ├── ...
│ │ └── 8000_judge.csv
│ └──… See the full description on the dataset page: https://huggingface.co/datasets/JiaqiXue/R2-Bench.enfermedades-wiki-marzo-2024
English Version
This dataset contains detailed information on a total of 945 diseases, extracted from Wikipedia in Spanish (https://es.wikipedia.org/) in March 2024. The main purpose of this dataset is to serve as a comprehensive resource for training Large Language Models (LLMs) in Spanish, specifically for instruction tuning, pre-training, and other natural language processing (NLP) tasks. This dataset promises to be a valuable tool for research and development in Spanish language… See the full description on the dataset page: https://huggingface.co/datasets/Alvaro8gb/enfermedades-wiki-marzo-2024.yc-companies-august-2025
Y Combinator Companies Dataset
Dataset Description
This dataset contains information about 5,404 Y Combinator funded companies that have been publicly launched, sourced from the YC-OSS-API.
Dataset Summary
Total Companies: 5,404
Time Range: Summer 2005 - Summer 2025
Update Frequency: Snapshot from August 2025
Source: YC-OSS-API
Dataset Structure
Data Fields
id: Unique identifier for each company
name: Company name… See the full description on the dataset page: https://huggingface.co/datasets/jeffboudier/yc-companies-august-2025.morphogen
MORPHOGEN: A Multilingual Benchmark for Evaluating Gender-Aware Morphological Generation
This repository contains the MORPHOGEN dataset introduced in our ACL 2026 paper: "MORPHOGEN: A Multilingual Benchmark for Evaluating Gender-Aware Morphological Generation".
Introduction
MORPHOGEN is a morphologically grounded, large-scale benchmark designed to evaluate the gender-aware generation capabilities of Large Language Models (LLMs) in three typologically diverse languages:… See the full description on the dataset page: https://huggingface.co/datasets/ag2003/morphogen.crypto-news-coindesk-2020-2025
CoinDesk Cryptocurrency News Dataset (2020–2025)
This dataset contains cryptocurrency-related news articles sourced from CoinDesk Data, accessed programmatically via the CryptoCompare API. The dataset is curated and published for academic and research purposes, with a focus on analyzing the relationship between news and cryptocurrency market dynamics.
Time Period
January 1, 2020 – January 1, 2025
Content Overview
Each record in the dataset… See the full description on the dataset page: https://huggingface.co/datasets/maryamfakhari/crypto-news-coindesk-2020-2025.2026.RA.Frontier-and-Scale-Cells
Rational-Agent Frontier, Scale, and Framing Cells
This public dataset is a sibling of siddharthmb/2026.RA.Negotiation-Campaigns (the frozen P1-P4 experimental record for the ii_mats/experiments/rational_agents negotiation program) and follows the same conventions: raw per-episode JSON, per-turn oracle annotations, Markdown/HTML transcripts, run manifests, analysis tables, and an integrity manifest over every uploaded file. It packages eight later campaigns that were run against… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/2026.RA.Frontier-and-Scale-Cells.evalita2026
This repository contains the data release for the Cruciverb-IT shared task on automatic crossword solving in Italian, as part of the 2026 EVALITA campaign. Refer to the task website for more details.
The data from both tasks can be downloaded from the 'Files and versions' tab.
Updates:
Minor update to both task_*_scorer.py in order to convert accented letters to their non-accented counterpart during evaluation
Test data is out!!
The test data of both… See the full description on the dataset page: https://huggingface.co/datasets/cruciverb-it/evalita2026.Natural_Language_to_Ffmpeg_Commands
Natural Language to FFmpeg Dataset
Disclaimer: This dataset was synthetically generated using a large language model and is intended for research purposes only. The dataset may contain inaccuracies, errors, or inconsistencies. Users should exercise caution and verify the correctness of the data before using it in any application.
This dataset contains 1000+ pairs of English natural language instructions and corresponding FFmpeg commands.
The dataset is designed for tasks… See the full description on the dataset page: https://huggingface.co/datasets/burak29/Natural_Language_to_Ffmpeg_Commands.Deepseek-V4-Reasoning-Code-2500
DeepSeek Reasoning and Code Distillation Dataset
This dataset contains synthetic instruction-response examples generated from coding, reasoning, and math prompts. It was generated with enforce_distillable_text enabled using DeepSeek V4 Pro and DeepSeek V4 Flash through OpenRouter. It is intended for experimentation with supervised fine-tuning, response-style distillation, reasoning-format analysis, and code-assistant behavior research.
The dataset file is:
train.csv
It contains 2… See the full description on the dataset page: https://huggingface.co/datasets/Banaxi-Tech/Deepseek-V4-Reasoning-Code-2500.GPT2-Hacker-password-generator-dataset
Hacker Style Password Generation Dataset
Dataset Description
This dataset contains 20,000 instruction-response pairs designed to train and evaluate language models for generating strong, "hacker-style" passwords. The data simulates a user requesting a secure password and the model providing a complex, randomly generated string.
Supported Tasks
Text Generation: The primary task is conditional text generation, where the model takes a natural language instruction… See the full description on the dataset page: https://huggingface.co/datasets/CodeferSystem/GPT2-Hacker-password-generator-dataset.ViSL-News
ViSL-News
Dataset Summary
ViSL-News is a sentence-level Vietnamese Sign Language (VSL) dataset constructed from sign-interpreted Vietnamese news broadcasts.
The dataset was built from HTV Tin Tức videos published on YouTube during 2024–2025. Each sample consists of a sentence-level sign-language video clip paired with a Vietnamese text sentence.
ViSL-News was constructed using ViSL-Tool, a semi-automated framework designed for news videos that contain spoken… See the full description on the dataset page: https://huggingface.co/datasets/kha2612/ViSL-News.AgentYear: 2025License: MITAuthor: Sepideh Moafi
PathogenAgentAI Instruction Dataset
Dataset Description
A ClinVar-derived dataset developed as part of the PathogenAgentAI research software project. The dataset is released in two parallel formats:
Tabular version (train.csv, valid.csv, test.csv) — structured genomic-variant data for classical ML and analysis.
BioGPT instruction version (biogpt_train.csv, biogpt_valid.csv, biogpt_test.csv) — instruction-style data… See the full description on the dataset page: https://huggingface.co/datasets/Sepideh2027/Agent.genz-to-english
GenZ-to-English Translation Dataset
A high-quality text-to-text dataset for translating Gen Z slang into clear, standard English.
The dataset is designed for training and evaluating language models that convert modern internet slang into natural, readable English while preserving the original meaning.
Overview
This dataset contains 300k++ curated translation pairs covering a wide range of contemporary internet slang.
It includes expressions commonly found across… See the full description on the dataset page: https://huggingface.co/datasets/Sankar-2910/genz-to-english.causalverify-neurips2026
🎯 CausalVerify
An Execution-Grounded Benchmark for LLM Causal Inference Workflows
NeurIPS 2026 — Evaluations and Datasets Track · double-blind review · frozen at tag neurips2026-submission
💡 TL;DR
A benchmark of 259 published economics papers (Experiment A — real-paper text-agreement diagnostic) and 100 fixed-seed synthetic data-generating processes (Experiment B — execution-grounded coefficient recovery), evaluating 7 frontier LLMs. The central… See the full description on the dataset page: https://huggingface.co/datasets/causalverify/causalverify-neurips2026.Titanium2-DeepSeek-R1Click here to support our open-source dataset and model releases!
Titanium2-DeepSeek-R1 is a dataset focused on architecture and DevOps, testing the limits of DeepSeek R1's architect and coding skills!
This dataset contains:
32.4k synthetically generated prompts focused on architecture, cloud, and DevOps. All responses are generated using DeepSeek R1. Primary areas of expertise are architecture (problem solving, scenario analysis, coding, full SDLC) and DevOps (Azure, AWS, GCP, Terraform… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Titanium2-DeepSeek-R1.Monthly-SWEBench-2026-05
Monthly-SWEBench 2026-05
This package contains the 2026-05 Monthly-SWEBench final release set. It includes 100 Harbor-format software engineering tasks selected from closed GitHub PRs and validated with oracle=1 / nop=0.
Files
bugfix.tar.zst: 50 bug-oriented repair or maintenance tasks.
non_bugfix.tar.zst: 50 feature, API evolution, or engineering-improvement tasks.
preview.csv: task ids, split labels, source change buckets, and archive paths.
tasks.conf: one… See the full description on the dataset page: https://huggingface.co/datasets/UnipatAI/Monthly-SWEBench-2026-05.idt5-v4-results-final-lora-s123-20260912T013040606815Z
final-lora-s123-20260912T013040606815Z
Run artifacts and per-item predictions.
Phase: final. These are newly generated results, not a reproduction of the legacy TCI tables.
See run_manifest.json, rules.json, generation_protocol.json and checkpoint_hashes.json. Structural scores do not establish semantic or Bloom validity.
Metrics
{
"n": 267,
"rule_version": "structural-proxy-v0.4-grounding-separated",
"parse_success_pct": 94.7565543071161,
"bleu":… See the full description on the dataset page: https://huggingface.co/datasets/Firmansyah-Ibrahim/idt5-v4-results-final-lora-s123-20260912T013040606815Z.JailBreakV-28k
⛓💥 JailBreakV-28K: A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak Attacks
🌐 GitHub | 🛎 Project Page | 👉 Download full datasets
If you like our project, please give us a star ⭐ on Hugging Face for the latest update.
📰 News
Date
Event
2024/07/09
🎉 Our paper is accepted by COLM 2024.
2024/06/22
🛠️ We have updated our version to V0.2, which supports users to customize their attack models… See the full description on the dataset page: https://huggingface.co/datasets/Ngixdev/JailBreakV-28k.PersonaMem-v2
PersonaMem-v2: Towards Personalized Intelligence via Learning Implicit User Personas and Agentic Memory
🚨 The paper is now released. View the full paper here and codebase here.
🙌 The dataset has been downloaded over 12,000 times. Thank you everybody for finding our work helpful!
Personalization is becoming the next milestone of artificial super-intelligence. AI cannot always satisfy every user, especially on tasks with subjective goals, but personalization… See the full description on the dataset page: https://huggingface.co/datasets/milanow/PersonaMem-v2.Mitakihara2-DeepSeek-V4-ProClick here to support our open-source dataset and model releases - help us speed up our release schedule!
Mitakihara 2 is an agentic coding dataset focused on MLOps and AI development, testing the limits of DeepSeek-V4-Pro's agentic skills:
Questions prioritize real-world, challenging agentic coding tasks in AI development, research, deployment, interpretability, operation and experimentation. The primary purpose of the Mitakihara dataset series is to accelerate and decentralize AI… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Mitakihara2-DeepSeek-V4-Pro.SWEbench-Verified-eval150-M2.7-Qwen3.5-9B-orch-7arms-2repeats-w32-20260920
SWE-bench Verified eval150 — M2.7 × Qwen3.5-9B, seven arms, two repeats, 32 concurrency
Campaign 2026-09-20. 14/14 independent full150 runs audited. Complete accuracy evidence.
Evaluation mode is orch: MiniMax-M2.7 orchestrator and the specified Qwen3.5-9B worker. Training mode is labeled independently. All runs use 32 concurrent episodes, 10GiB Docker sandboxes, four TP1 workers and one TP4/EP4 coordinator. Frozen regression-gated prompts, decoding and canonical verifier match… See the full description on the dataset page: https://huggingface.co/datasets/CharlieLLL/SWEbench-Verified-eval150-M2.7-Qwen3.5-9B-orch-7arms-2repeats-w32-20260920.ai-writing-prompts-2025** 100 AI Writing Prompts & Marketing Hooks 2025 — A ready-to-use human readable CSV file (no coding needed option) for blogging, social media, and AI text generation projects. **
A 100-row structured dataset of AI-ready writing prompts and marketing hooks designed for blogging, copywriting, and AI text generation. Includes tone, use case, and audience metadata for fine-tuning and automation
✨ 100 AI Writing Prompts & Marketing Hooks 2025 (AI-Ready CSV for Text Generation)
A… See the full description on the dataset page: https://huggingface.co/datasets/alexandrabozarthmanagement/ai-writing-prompts-2025.Monthly-SWEBench-2026-03
Monthly-SWEBench-2026-03
Monthly-SWEBench-2026-03 is a curated benchmark of 112 real-world software engineering tasks, sourced from GitHub pull requests merged in March 2026. Tasks are in Harbor format and can be run with any Harbor-compatible agent.
View leaderboard and results →
112 tasks — 68 bugfix + 44 non-bugfix
Tasks span diverse open-source repositories
Each task includes a runnable environment, test suite, and reference solution
Task Structure
Each… See the full description on the dataset page: https://huggingface.co/datasets/UnipatAI/Monthly-SWEBench-2026-03.
