datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SWE-PolyBench
SWE-PolyBench
SWE-PolyBench is a multi language repo level software engineering benchmark. Currently it includes 4 languages: Python, Java, Javascript, and Typescript. The number of instances in each language is:
Javascript: 1017
Typescript: 729
Python: 199
Java: 165
Datasets
There are total three datasets available under SWE-PolyBench. AmazonScience/SWE-PolyBench is the full dataset, AmazonScience/SWE-PolyBench_500 is the stratified sampled dataset with 500 instances and… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/SWE-PolyBench.SWE-bench-Science
SWE-bench Science
SWE-bench Science evaluates coding agents on software-engineering tasks drawn from scientific-computing repositories. The release contains 119 tasks across 20 scientific domains, with isolated environments and separate programmatic verifiers.
GitHub release repository: OpenMOSS/SWE-bench-Science
Runtime images: Docker Hub, pinned by immutable linux/amd64 digests
Evaluation framework: Pier, compatible with Harbor task format
Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/SWE-bench-Science.sweden_100K_difficultSWE-PolyBench_Verified
SWE-PolyBench
SWE-PolyBench is a multi language repo level software engineering benchmark. Currently it includes 4 languages: Python, Java, Javascript, and Typescript. The number of instances in the verified split is:
Javascript: 100
Typescript: 100
Python: 113
Java: 69
Datasets
There are total three datasets available under SWE-PolyBench. AmazonScience/SWE-PolyBench is the full dataset, AmazonScience/SWE-PolyBench_500 is the stratified sampled dataset with 500 instances… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/SWE-PolyBench_Verified.SWE-PolyBench_500
SWE-PolyBench
SWE-PolyBench is a multi language repo level software engineering benchmark. Currently it includes 4 languages: Python, Java, Javascript, and Typescript. The number of instances in each language is:
Javascript: 1017
Typescript: 729
Python: 199
Java: 165
Datasets
There are total three datasets available under SWE-PolyBench. AmazonScience/SWE-PolyBench is the full dataset, AmazonScience/SWE-PolyBench_500 is the stratified sampled dataset with 500 instances and… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/SWE-PolyBench_500.swepro-luna-matched-pair
SWE-bench Pro, matched pair: Ouroboros vs Codex CLI on one model
Status: Self-reported matched-pair study. Both harnesses used the same
model, task set and evaluator. The strict result is a statistical tie.
Start here
Strict result
Ouroboros 58.2%, Codex CLI 59.4%, McNemar p = 0.40
Model
openai/gpt-5.6-luna for both arms
Filter
655 paired tasks after the same reference-leak filter was applied to both arms
Exact evidence
6228037, manifest.csv… See the full description on the dataset page: https://huggingface.co/datasets/razzant/swepro-luna-matched-pair.rhan-eval-sweepSWE-Sharp-Bench
SWE-Sharp-Bench
SWE-Sharp-Bench is a comprehensive benchmark suite for evaluating software engineering capabilities of AI agents and models on C# and .NET codebases. This benchmark extends the SWE-Bench framework to the C# ecosystem, providing real-world software engineering tasks from popular open-source repositories.
Code - https://github.com/microsoft/prose/tree/main/misc/SWE-Sharp-Bench
Research Paper Draft & Benchmark Analysis: https://aka.ms/swesharparxiv
Contact… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/SWE-Sharp-Bench.SWE-bench_Verified_With_AnnotationsSWE-bench_Pro
Dataset Summary
SWE-Bench Pro is a challenging, enterprise-level dataset for testing agent ability on long-horizon software engineering tasks.
Paper: https://static.scale.com/uploads/654197dc94d34f66c0f5184e/SWEAP_Eval_Scale%20(9).pdf
See the related evaluation Github: https://github.com/scaleapi/SWE-bench_Pro-os
Dataset Structure
We follow SWE-Bench Verified (https://huggingface.co/datasets/SWE-bench/SWE-bench_Verified) in terms of dataset structure, with several… See the full description on the dataset page: https://huggingface.co/datasets/Contextbench/SWE-bench_Pro.swe-bench-atlas-anon-public
SWE-bench Atlas
1. Summary
In the domain of software engineering, LLM capabilities have progressed rapidly, underscoring the need for evolving evaluation frameworks. While foundational, benchmarks like SWE-bench, SWE-bench Verified, and other such variants are incomplete, with manually curated design causing scalability bottlenecks, weak test oracles, dataset aging and contamination, reproducibility challenges, and more.
In response, we introduce SWE-bench Atlas: a… See the full description on the dataset page: https://huggingface.co/datasets/swebenchatlas/swe-bench-atlas-anon-public.quora_swe
Dataset Card for "quora_swe"
The dataset quora_swe is a subset of the automatically translated (MNT) Swedish Semantic Textual Similarity dataset: quora-deduplicates .
mergebench-property-sweep
MergeBench property sweep
Pre-merge pairwise properties for every mergeable pair in the
MergeBench suite (40 checkpoints, 8 base families, 5 domains),
computed with the metric panel behind Figure F8 of the Heterogeneous Mergeability project.
Read this before you read a number
MergeBench publishes no pairwise merge. Every merge score in their release
(arXiv:2505.10833, Tables 8-17, and the two eval dumps in their
GitHub repo) is for a merge of all five domain… See the full description on the dataset page: https://huggingface.co/datasets/Cross-Mergeability/mergebench-property-sweep.2026.RA.Quorum-Rounds-Sweep
2026.RA.Quorum-Rounds-Sweep — what a decision rule does to a table of rational negotiators
Every episode of the quorum x rounds sweep: five computable Bayesian-rational negotiators bargaining over a
package of four issues, replayed on one frozen 24-game bank under three different agreement rules and two
different deadlines, plus a two-factor extension that also dissolves the veto.
What the experiment asks
Five parties must agree on one package out of 256. Each… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/2026.RA.Quorum-Rounds-Sweep.SWE-bench_Verified_SmallSWE-bench_LiteMonthly-SWEBench-2026-05
Monthly-SWEBench 2026-05
This package contains the 2026-05 Monthly-SWEBench final release set. It includes 100 Harbor-format software engineering tasks selected from closed GitHub PRs and validated with oracle=1 / nop=0.
Files
bugfix.tar.zst: 50 bug-oriented repair or maintenance tasks.
non_bugfix.tar.zst: 50 feature, API evolution, or engineering-improvement tasks.
preview.csv: task ids, split labels, source change buckets, and archive paths.
tasks.conf: one… See the full description on the dataset page: https://huggingface.co/datasets/UnipatAI/Monthly-SWEBench-2026-05.swe_mon_benchmarkSWE-bump-benchMonthly-SWEBench-2026-03
Monthly-SWEBench-2026-03
Monthly-SWEBench-2026-03 is a curated benchmark of 112 real-world software engineering tasks, sourced from GitHub pull requests merged in March 2026. Tasks are in Harbor format and can be run with any Harbor-compatible agent.
View leaderboard and results →
112 tasks — 68 bugfix + 44 non-bugfix
Tasks span diverse open-source repositories
Each task includes a runnable environment, test suite, and reference solution
Task Structure
Each… See the full description on the dataset page: https://huggingface.co/datasets/UnipatAI/Monthly-SWEBench-2026-03.Monthly-SWEBench-2026-04
Monthly-SWEBench-2026-04
Monthly-SWEBench-2026-04 is a curated benchmark of 90 real-world software engineering tasks, sourced from GitHub pull requests merged in April 2026. Tasks are in Harbor format and can be run with any Harbor-compatible agent.
View leaderboard and results →
90 tasks — 43 bugfix + 47 non-bugfix
Tasks span diverse open-source repositories
Each task includes a runnable environment, test suite, and reference solution
Task Structure
Each task is a… See the full description on the dataset page: https://huggingface.co/datasets/UnipatAI/Monthly-SWEBench-2026-04.2026.RA.Pure-Rounds-Sweep
2026.RA.Pure-Rounds-Sweep — does more negotiating time help a Bayesian table close?
Five automated negotiators must agree on one package out of 256. Each holds a private score sheet and a
private walk-away threshold; a deal needs all five to accept. This corpus sweeps the one thing the frozen
campaign never varied — the number of negotiation rounds before the forced final vote — across
4 to 256 rounds, for the two lineups that contain no language model and therefore
cost only… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/2026.RA.Pure-Rounds-Sweep.SWEbench-Verified-eval150-M2.7-Qwen3.5-9B-orch-7arms-2repeats-w32-20260920
SWE-bench Verified eval150 — M2.7 × Qwen3.5-9B, seven arms, two repeats, 32 concurrency
Campaign 2026-09-20. 14/14 independent full150 runs audited. Complete accuracy evidence.
Evaluation mode is orch: MiniMax-M2.7 orchestrator and the specified Qwen3.5-9B worker. Training mode is labeled independently. All runs use 32 concurrent episodes, 10GiB Docker sandboxes, four TP1 workers and one TP4/EP4 coordinator. Frozen regression-gated prompts, decoding and canonical verifier match… See the full description on the dataset page: https://huggingface.co/datasets/CharlieLLL/SWEbench-Verified-eval150-M2.7-Qwen3.5-9B-orch-7arms-2repeats-w32-20260920.vn-provinces-sweet-potato-sown-area
Vietnam provinces sweet potato sown area
Provincial and regional sweet potato (khoai lang) sown area (thousand hectares). Coverage 1995-2024. Year 2024 is preliminary. Includes historical Ha Tay through 2007. Values from 2018 onward are normalized from hectares to thousand hectares (NSO unit break). Geographic labels are English (UN/GSO style ASCII romanization). Tables cover provinces, regions and national total where present. Province names follow ar_core.vn_geo (historical… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-sweet-potato-sown-area.vn-provinces-sweet-potato-production
Vietnam provinces sweet potato production
Sweet potato (khoai lang) production (thousand tons). Coverage 1995-2024. Year 2024 is preliminary. Includes historical Ha Tay through 2007. Geographic labels are English (UN/GSO style ASCII romanization). Tables cover provinces, regions and national total where present. Province names follow ar_core.vn_geo (historical 63-province system).
Figures
Hero
Hero (continued)
Comparison
Color key
Files… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-sweet-potato-production.swedish_skolprov
Swedish skolprov documentation
This repository contains data from six Swedish knowledge tests. The data is in the form of multiple-choice questions and answers. The data is stored in CSV files.
The following tests are in the data:
högskoleprovet (Swedish Scholastic Aptitude Test)
kunskapsprov läkare (medical doctor test)
kunskapsprov tandläkare (dentist test)
kunskapsprov audionom
kunskapsprov apotekare
matematik och fysikprovet (math and highschool test)
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/Ekgren/swedish_skolprov.swedish-cefr-text-complexity
Swedish CEFR Text Complexity Dataset
This dataset contains Swedish text examples labeled with approximate CEFR
reading levels from A1 to C2.
It was created for an information retrieval assignment about training text
classifiers with embeddings. The companion demo and classifier use
nicher92/saga-embed_v1 sentence embeddings and classical scikit-learn
classifiers.
The dataset is intended for Swedish text-complexity classification: given a
short Swedish sentence or paragraph, predict… See the full description on the dataset page: https://huggingface.co/datasets/kvest/swedish-cefr-text-complexity.chatgpt-prompts-SwedishNemotron-RL-Agentic-SWE-Pivot-v1-prompt-only
Nemotron-RL-Agentic-SWE-Pivot-v1-prompt-only
Prompt-only extraction from nvidia/Nemotron-RL-Agentic-SWE-Pivot-v1.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.
null_or_empty_rows.md: row indexes where… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-RL-Agentic-SWE-Pivot-v1-prompt-only.SWE-PolyBench
SWE-PolyBench
SWE-PolyBench is a multi language repo level software engineering benchmark. Currently it includes 4 languages: Python, Java, Javascript, and Typescript. The number of instances in each language is:
Javascript: 1017
Typescript: 729
Python: 199
Java: 165
Datasets
There are total three datasets available under SWE-PolyBench. AmazonScience/SWE-PolyBench is the full dataset, AmazonScience/SWE-PolyBench_500 is the stratified sampled dataset with 500 instances and… See the full description on the dataset page: https://huggingface.co/datasets/XUO/SWE-PolyBench.
