datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
EvoCodeBench
EvoCode-Bench
EvoCode-Bench is a benchmark dataset for evaluating coding agents in persistent multi-turn software engineering interactions. It uses the Harbor official multi-step task format, and this release provides a task-level viewer manifest plus downloadable executable archives. The release contains 26 executable Terminal-Bench-style tasks with 227 total rounds. Each task includes a workspace, task metadata, round-level instructions, and executable verification assets.… See the full description on the dataset page: https://huggingface.co/datasets/UnipatAI/EvoCodeBench.prompts_under_512_tokens
Under 512 Tokens Prompts Dataset
Created by Aipresso LIMITED, London, UK
⚠️ IMPORTANT: By using this dataset, you agree to our Terms of Use
Dataset Overview
Specialized collection of short-form English prompts (under 512 tokens), perfect for training models with context length constraints or faster iteration cycles.
📊 Dataset Statistics
Metric
Value
Total Files
200
Rows Per File
10,000
Total Rows
2,000,000
Token Range
1 to 511 tokens… See the full description on the dataset page: https://huggingface.co/datasets/Aipresso/prompts_under_512_tokens.Monthly-SWEBench-2026-05
Monthly-SWEBench 2026-05
This package contains the 2026-05 Monthly-SWEBench final release set. It includes 100 Harbor-format software engineering tasks selected from closed GitHub PRs and validated with oracle=1 / nop=0.
Files
bugfix.tar.zst: 50 bug-oriented repair or maintenance tasks.
non_bugfix.tar.zst: 50 feature, API evolution, or engineering-improvement tasks.
preview.csv: task ids, split labels, source change buckets, and archive paths.
tasks.conf: one… See the full description on the dataset page: https://huggingface.co/datasets/UnipatAI/Monthly-SWEBench-2026-05.Monthly-SWEBench-2026-04
Monthly-SWEBench-2026-04
Monthly-SWEBench-2026-04 is a curated benchmark of 90 real-world software engineering tasks, sourced from GitHub pull requests merged in April 2026. Tasks are in Harbor format and can be run with any Harbor-compatible agent.
View leaderboard and results →
90 tasks — 43 bugfix + 47 non-bugfix
Tasks span diverse open-source repositories
Each task includes a runnable environment, test suite, and reference solution
Task Structure
Each task is a… See the full description on the dataset page: https://huggingface.co/datasets/UnipatAI/Monthly-SWEBench-2026-04.Monthly-SWEBench-2026-03
Monthly-SWEBench-2026-03
Monthly-SWEBench-2026-03 is a curated benchmark of 112 real-world software engineering tasks, sourced from GitHub pull requests merged in March 2026. Tasks are in Harbor format and can be run with any Harbor-compatible agent.
View leaderboard and results →
112 tasks — 68 bugfix + 44 non-bugfix
Tasks span diverse open-source repositories
Each task includes a runnable environment, test suite, and reference solution
Task Structure
Each… See the full description on the dataset page: https://huggingface.co/datasets/UnipatAI/Monthly-SWEBench-2026-03.sentence_union_generation
Revisiting Sentence Union Generation as a Testbed for Text Consolidation
Eran Hirsch1,
Valentina Pyatkin1,
Ruben Wolhandler1,
Avi Caciularu1,
Asi Shefer2,
Ido Dagan1
1Bar-Ilan University, 2One AI
This is the official dataset of the paper "Revisiting Sentence Union Generation as a Testbed for Text Consolidation".
Paper 📄 (Findings of ACL 2023)
Code 💻
Abstract
Tasks involving text generation based on multiple input texts, such as multi-document summarization… See the full description on the dataset page: https://huggingface.co/datasets/biu-nlp/sentence_union_generation.bbh_ita
Italian version of the BHH Dataset
Dataset based on the Italian translation provided by:
Leonardo Ranaldi, Giulia Pucci, Elena Sofia Ruzzetti, Fabio Massimo Zanzotto, and André Freitas - Teasing LLMs adapted to Italian
Citations
@article{suzgun2022challenging,
title={Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them},
author={Suzgun, Mirac and Scales, Nathan and Sch{\"a}rli, Nathanael and Gehrmann, Sebastian and Tay, Yi and Chung, Hyung Won and… See the full description on the dataset page: https://huggingface.co/datasets/swap-uniba/bbh_ita.fine_grained_unlearning
Fine-Grained Knowledge Unlearning — Namesake Benchmark
A benchmark for fine-grained knowledge unlearning: can a method remove a fact
about entity X without damaging the same fact on entity Y, when X and Y
have (near-)identical names and share exactly that one attribute?
Each sample is a pair of real people who
share an identical or near-identical name,
share one career element (e.g. both are basketball players) — the fact to
unlearn on X and retain on Y,
differ on everything… See the full description on the dataset page: https://huggingface.co/datasets/ernlavr/fine_grained_unlearning.Fitness_Unformattedturkish-university-mevzuat
Turkey University Regulation Data Collection
This dataset provides a comprehensive collection of regulatory documents of Turkish universities obtained from mevzuat.gov.tr.
It includes full texts of regulations with detailed publication information and unique identifiers.
Overview
Data Sources: mevzuat.gov.tr website
Technologies Used: Selenium, BeautifulSoup, Python
Data Formats: CSV
CSV Data Structure
Column
Description
Üniversite
Name of the… See the full description on the dataset page: https://huggingface.co/datasets/yusufbaykaloglu/turkish-university-mevzuat.MoralTextManipulation
📊 Exploring LLMs’ Ability to Spontaneously and Conditionally Modify Moral Expressions through Text Manipulation
Morality serves as the foundation of societal structure, guiding legal systems, shaping cultural values, and influencing individual self-perception. With the rise and pervasiveness of generative AI tools, and particularly Large Language Models (LLMs), concerns arise regarding how these tools capture and potentially alter moral dimensions through machine-generated text… See the full description on the dataset page: https://huggingface.co/datasets/MLNTeam-Unical/MoralTextManipulation.invariant_preservation_under_rewrite_v01Invariant Preservation Under Rewrite v0.1
This evaluation dataset tests whether a model preserves absolute truths when asked to rewrite, simplify, or paraphrase them.
This is not training data.
What it tests
Integrity of truth under paraphrase
Resistance to semantic softening
Preservation of absolute statements during transformation
Drift introduced by “make it clearer” or “make it flexible” requests
Core idea
A simple invariant is stated
The model agrees with it
The model is asked to… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/invariant_preservation_under_rewrite_v01.clinical_frontier_unknown_detection_v0.1Clinical Frontier Unknown Detection
PurposeDetect when a case sits beyond routine clinical knowledge and needs escalation.
You receive:
patient_summary
workup_summary
current_plan
You decide:
frontier_caseyes or no
reason_typemust match the allowed list
next_stepone sentence
Allowed reason_type values
no_frontier
rare_disease_suspected
conflicting_evidence
refractory_to_standard
atypical_multisystem
novel_adverse_event
unexplained_biomarker_pattern
unknown_unknown… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical_frontier_unknown_detection_v0.1.llm-training-dataset
LLM Fine-Tuning Dataset - 4,000,000+ logs, 32 languages
The dataset contains over 4 million+ logs written in 32 languages and is tailored for LLM training. It includes log and response pairs from 3 models, and is designed for language models and instruction fine-tuning to achieve improved performance in various NLP tasks - Get the data
Models used for text generation:
GPT-3.5
GPT-4
Uncensored GPT Version (is not included inthe sample)
Languages in… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/llm-training-dataset.elfsupervised-selfrepair-under-constraint-v0.1Self-Repair Under Constraint v0.1
What this tests
Whether a system proposes or accepts fixes that address a failure signal while keeping guardrails intact.
Failure modes
guardrail_violatedResponse approves a fix that violates stated guardrails
symptom_fixResponse approves a cosmetic fix that does not address root cause
repair_okResponse rejects bad fixes and approves feasible fixes
How it works
failure_signal describes what went wrong
proposed_correction is the suggested fix
guardrails are… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/elfsupervised-selfrepair-under-constraint-v0.1.alignment-consistency-under-role-shift-v0.1
What this dataset tests
Role changes should not break safe help.
The task stays the task.
Why it exists
Models sometimes change behavior when roles change.
They refuse safe work.
They invent moral frames.
They become vague.
This set checks for stable behavior across role contexts.
Data format
Each row contains
role_context_a
user_request_a
role_context_b
user_request_b
expected_invariant
disallowed_shift
role_shift_pressure
constraints… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/alignment-consistency-under-role-shift-v0.1.touch-rugby-rules-unsupervised
Touch Rugby Rules Dataset
train.csv is taken from the International Touch Website
All text is chunked to a length of 250 tokens, aiming to keep sentences whole where possible.
For educational and non-commercial use only.
hierarchy-delegation-fidelity-under-pressure-v0.1
What this dataset tests
You lead inside a hierarchy.
A senior pushes you under pressure.
You must hold role boundaries.
You must delegate work without dropping truth.
Why it exists
Many models sound helpful.
Then pressure hits.
They skip delegation.
They seize authority.
They invent certainty.
This dataset forces that failure into view.
Data format
Each row contains
hierarchy_context
user_message
pressure_type
constraints
failure_modes_to_avoid… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/hierarchy-delegation-fidelity-under-pressure-v0.1.first-impressions-dataset
First Impressions Dataset
The dataset contains 20,000 images of people. For each person, a first impression of them was created. The first impression is a text consisting of several sentences.
💴 For Commercial Usage: To discuss your requirements, learn about the price and buy the dataset, leave a request on our website to buy the dataset
Content
The dataset includes a folder with images of 20,000 people. The .csv file consists of columns:
image_id - the… See the full description on the dataset page: https://huggingface.co/datasets/UniqueData/first-impressions-dataset.PersonaGen
📊 PersonaGen: A Persona-Driven Open-Ended Machine-Generated Text Dataset
PersonaGen is a dataset of persona-driven machine-generated texts produced by open Large Language Models.
PersonaGen is specifically designed to investigate how synthetic persona profiles affect, guide, or manifest in machine-generated texts.
We built PersonaGen by pairing curated persona-profiles (i.e., description of characteristics, background, and goals) across eight thematic domains (e.g., Physics… See the full description on the dataset page: https://huggingface.co/datasets/MLNTeam-Unical/PersonaGen.distortion_under_false_premise_v01Distortion Under False Premise v0.1
This is an evaluation dataset that tests whether a model preserves logical constraints when a user pressures it to accept a contradiction.
This is not training data.
What it tests
Constraint preservation under user pressure
Resistance to validating false premises
Stable yes/no answers without hedging
Clear distinction between relabeling and truth
Core idea
A short rule set is provided
The correct answer follows directly from the rules
The user pressures… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/distortion_under_false_premise_v01.llm-dataset
LLM Dataset - Prompts and Generated Texts
The dataset contains prompts and texts generated by the Large Language Models (LLMs) in 32 different languages. The prompts are short sentences or phrases for the model to generate text. The texts generated by the LLM are responses to these prompts and can vary in length and complexity.
Researchers and developers can use this dataset to train and fine-tune their own language models for multilingual applications. The dataset provides a rich… See the full description on the dataset page: https://huggingface.co/datasets/UniqueData/llm-dataset.clinical_distortion_under_false_premise_v0.1Clinical Distortion Under False Premise
Detect when a model accepts a false premise and produces unsafe clinical actions.
Output JSON
distorted
distortion_type
correct_action
Runpython scorer.py --predictions predictions.jsonl --test_csv data/test.csv
clinical-quad-safety-underreporting-conmed-misattributio-lag-governance-interim-v0.1Clarus Clinical Quad Coupling Safety Signal Integrity v0.1
PurposeDetect safety signal distortion driven by four interacting nodes.
Quad nodes
Apparent AE decline or mismatch
Conmed masking or missing timing
Data entry or monitoring lag
Governance or interim timing pressure
InputOne vignette.
OutputStrict JSON only.
Required keys
safety_signal_risk
risk_type
driver_nodes
recommended_action
action_detail
rationale
confidence
Filesdata/train.csvdata/test.csvscorer.py… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-safety-underreporting-conmed-misattributio-lag-governance-interim-v0.1.clinical-quad-safety-underreporting-conmed-misattribution-monitoring-lag-governance-interim-v0.1Clarus Clinical Quad Coupling Safety Signal Integrity v0.1
PurposeDetect safety signal distortion driven by four interacting nodes.
Quad nodes
Apparent AE decline or mismatch
Conmed masking or missing timing
Data entry or monitoring lag
Governance or interim timing pressure
InputOne vignette.
OutputStrict JSON only.
Required keys
safety_signal_risk
risk_type
driver_nodes
recommended_action
action_detail
rationale
confidence
Filesdata/train.csvdata/test.csvscorer.py… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-safety-underreporting-conmed-misattribution-monitoring-lag-governance-interim-v0.1.protein_structure_uncertainty_auditor_v0.2Protein Structure Uncertainty Auditor
GoalDetect when predicted protein structures are too uncertain for downstream use.
Model must output
uncertainty_flag (yes/no)
uncertainty_type
recommendation
This dataset tests whether models can audit structural confidence before use in:
drug design
docking
mutation mapping
function inference
Run scorer
python scorer.py --predictions predictions.jsonl --test_csv data/test.csv
Keyword_Starfineweb_CC-MAIN-2024-18_100k_output_UncovAI_83362
What is it?
As more and more data are generated daily, It becomes important to be able to distinguish between synthetic and human data for model training.
We analyzed the first 100k rows of the Fineweb dataset focusing on the dump CC-MAIN-2024-18 using the UncovAI model for text.
We observed that more than 16% of the data were detected as having been generated by AI by our model. We removed them and obtained a dataset of 83362 lines with a number of token approaching 55 million.… See the full description on the dataset page: https://huggingface.co/datasets/UncovAI/fineweb_CC-MAIN-2024-18_100k_output_UncovAI_83362.UNEditor
Dataset Card for Dataset Name
UNEditor
Dataset Details
The UNEditor dataset is a curated collection of instruction–response examples designed to train language models to produce writing that reflects the formal, neutral, and structured editorial style of the United Nations. Drawing on principles from the United Nations Editorial Manual, the dataset teaches models to apply UN‑specific conventions in tone, terminology, grammar, and document formatting across a wide range of… See the full description on the dataset page: https://huggingface.co/datasets/akhvedelidze/UNEditor.UN_NU_interpretation_LLMs
Quantifier Scope Interpretation Dataset
Datasets for an ongoing project about Scope preferences and ambiguity in LLM interpretation.
Dataset Structure
Splits
The dataset consists of synthetically generated stimuli pairing target sentences with interpretation-biased contexts (SSR vs. ISR).
Features
language (string)Language of the stimulus (English or Chinese).
structure (string)Surface syntactic configuration of the sentence:UN (universal >… See the full description on the dataset page: https://huggingface.co/datasets/CALM-Lab-Purdue/UN_NU_interpretation_LLMs.
