datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SWE-PolyBench
SWE-PolyBench
SWE-PolyBench is a multi language repo level software engineering benchmark. Currently it includes 4 languages: Python, Java, Javascript, and Typescript. The number of instances in each language is:
Javascript: 1017
Typescript: 729
Python: 199
Java: 165
Datasets
There are total three datasets available under SWE-PolyBench. AmazonScience/SWE-PolyBench is the full dataset, AmazonScience/SWE-PolyBench_500 is the stratified sampled dataset with 500 instances and… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/SWE-PolyBench.SWE-PolyBench_Verified
SWE-PolyBench
SWE-PolyBench is a multi language repo level software engineering benchmark. Currently it includes 4 languages: Python, Java, Javascript, and Typescript. The number of instances in the verified split is:
Javascript: 100
Typescript: 100
Python: 113
Java: 69
Datasets
There are total three datasets available under SWE-PolyBench. AmazonScience/SWE-PolyBench is the full dataset, AmazonScience/SWE-PolyBench_500 is the stratified sampled dataset with 500 instances… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/SWE-PolyBench_Verified.sweden_100K_difficultSWE-PolyBench_500
SWE-PolyBench
SWE-PolyBench is a multi language repo level software engineering benchmark. Currently it includes 4 languages: Python, Java, Javascript, and Typescript. The number of instances in each language is:
Javascript: 1017
Typescript: 729
Python: 199
Java: 165
Datasets
There are total three datasets available under SWE-PolyBench. AmazonScience/SWE-PolyBench is the full dataset, AmazonScience/SWE-PolyBench_500 is the stratified sampled dataset with 500 instances and… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/SWE-PolyBench_500.swepro-luna-matched-pair
SWE-bench Pro, matched pair: Ouroboros vs Codex CLI on one model
Status: Self-reported matched-pair study. Both harnesses used the same
model, task set and evaluator. The strict result is a statistical tie.
Start here
Strict result
Ouroboros 58.2%, Codex CLI 59.4%, McNemar p = 0.40
Model
openai/gpt-5.6-luna for both arms
Filter
655 paired tasks after the same reference-leak filter was applied to both arms
Exact evidence
6228037, manifest.csv… See the full description on the dataset page: https://huggingface.co/datasets/razzant/swepro-luna-matched-pair.rhan-eval-sweepSWE-bench_Verified_With_Annotationsquora_swe
Dataset Card for "quora_swe"
The dataset quora_swe is a subset of the automatically translated (MNT) Swedish Semantic Textual Similarity dataset: quora-deduplicates .
mergebench-property-sweep
MergeBench property sweep
Pre-merge pairwise properties for every mergeable pair in the
MergeBench suite (40 checkpoints, 8 base families, 5 domains),
computed with the metric panel behind Figure F8 of the Heterogeneous Mergeability project.
Read this before you read a number
MergeBench publishes no pairwise merge. Every merge score in their release
(arXiv:2505.10833, Tables 8-17, and the two eval dumps in their
GitHub repo) is for a merge of all five domain… See the full description on the dataset page: https://huggingface.co/datasets/Cross-Mergeability/mergebench-property-sweep.2026.RA.Quorum-Rounds-Sweep
2026.RA.Quorum-Rounds-Sweep — what a decision rule does to a table of rational negotiators
Every episode of the quorum x rounds sweep: five computable Bayesian-rational negotiators bargaining over a
package of four issues, replayed on one frozen 24-game bank under three different agreement rules and two
different deadlines, plus a two-factor extension that also dissolves the veto.
What the experiment asks
Five parties must agree on one package out of 256. Each… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/2026.RA.Quorum-Rounds-Sweep.SWE-bench_Verified_SmallSWE-bump-benchSWEbench-Verified-eval150-M2.7-Qwen3.5-9B-orch-7arms-2repeats-w32-20260920
SWE-bench Verified eval150 — M2.7 × Qwen3.5-9B, seven arms, two repeats, 32 concurrency
Campaign 2026-09-20. 14/14 independent full150 runs audited. Complete accuracy evidence.
Evaluation mode is orch: MiniMax-M2.7 orchestrator and the specified Qwen3.5-9B worker. Training mode is labeled independently. All runs use 32 concurrent episodes, 10GiB Docker sandboxes, four TP1 workers and one TP4/EP4 coordinator. Frozen regression-gated prompts, decoding and canonical verifier match… See the full description on the dataset page: https://huggingface.co/datasets/CharlieLLL/SWEbench-Verified-eval150-M2.7-Qwen3.5-9B-orch-7arms-2repeats-w32-20260920.vn-provinces-sweet-potato-sown-area
Vietnam provinces sweet potato sown area
Provincial and regional sweet potato (khoai lang) sown area (thousand hectares). Coverage 1995-2024. Year 2024 is preliminary. Includes historical Ha Tay through 2007. Values from 2018 onward are normalized from hectares to thousand hectares (NSO unit break). Geographic labels are English (UN/GSO style ASCII romanization). Tables cover provinces, regions and national total where present. Province names follow ar_core.vn_geo (historical… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-sweet-potato-sown-area.vn-provinces-sweet-potato-production
Vietnam provinces sweet potato production
Sweet potato (khoai lang) production (thousand tons). Coverage 1995-2024. Year 2024 is preliminary. Includes historical Ha Tay through 2007. Geographic labels are English (UN/GSO style ASCII romanization). Tables cover provinces, regions and national total where present. Province names follow ar_core.vn_geo (historical 63-province system).
Figures
Hero
Hero (continued)
Comparison
Color key
Files… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-sweet-potato-production.swedish_skolprov
Swedish skolprov documentation
This repository contains data from six Swedish knowledge tests. The data is in the form of multiple-choice questions and answers. The data is stored in CSV files.
The following tests are in the data:
högskoleprovet (Swedish Scholastic Aptitude Test)
kunskapsprov läkare (medical doctor test)
kunskapsprov tandläkare (dentist test)
kunskapsprov audionom
kunskapsprov apotekare
matematik och fysikprovet (math and highschool test)
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/Ekgren/swedish_skolprov.Nemotron-RL-Agentic-SWE-Pivot-v1-prompt-only
Nemotron-RL-Agentic-SWE-Pivot-v1-prompt-only
Prompt-only extraction from nvidia/Nemotron-RL-Agentic-SWE-Pivot-v1.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.
null_or_empty_rows.md: row indexes where… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-RL-Agentic-SWE-Pivot-v1-prompt-only.2026.RA.Pure-Rounds-Sweep
2026.RA.Pure-Rounds-Sweep — does more negotiating time help a Bayesian table close?
Five automated negotiators must agree on one package out of 256. Each holds a private score sheet and a
private walk-away threshold; a deal needs all five to accept. This corpus sweeps the one thing the frozen
campaign never varied — the number of negotiation rounds before the forced final vote — across
4 to 256 rounds, for the two lineups that contain no language model and therefore
cost only… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/2026.RA.Pure-Rounds-Sweep.qwen9b-coop-mini-swe-agent
qwen9b-coop-mini-swe-agent
Two-agent cooperative coding trajectories generated by running
CooperBench in coop mode on
the CooperData task set, using
Qwen/Qwen3.5-9B as the model and mini_swe_agent_v2 as the agent framework.
Each pair runs two agents in parallel — one per feature — coordinating via Redis messaging and a shared git remote.
The matched solo version is at
CooperBench/qwen9b-solo-mini-swe-agent.
Same task corpus, same model, same agent — only the coordination differs… See the full description on the dataset page: https://huggingface.co/datasets/CooperBench/qwen9b-coop-mini-swe-agent.SWE-PolyBench
SWE-PolyBench
SWE-PolyBench is a multi language repo level software engineering benchmark. Currently it includes 4 languages: Python, Java, Javascript, and Typescript. The number of instances in each language is:
Javascript: 1017
Typescript: 729
Python: 199
Java: 165
Datasets
There are total three datasets available under SWE-PolyBench. AmazonScience/SWE-PolyBench is the full dataset, AmazonScience/SWE-PolyBench_500 is the stratified sampled dataset with 500 instances and… See the full description on the dataset page: https://huggingface.co/datasets/XUO/SWE-PolyBench.Nemotron-SWE-v1-prompt-only
Nemotron-SWE-v1-prompt-only
Prompt-only extraction from nvidia/Nemotron-SWE-v1.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.
null_or_empty_rows.md: row indexes where prompt extraction produced a null or… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-SWE-v1-prompt-only.Swedia-ASR-Dataset
Swedia ASR Dataset
This repository contains a small Swedish ASR evaluation dataset based on speech
transcriptions from Swedia 2000. It was assembled to compare automatic
speech-recognition output against manually corrected reference transcriptions
for Swedish dialectal speech.
The dataset is useful for quick experiments with Swedish ASR systems, especially
when you want to inspect recognition quality on spontaneous speech from
different regions, speakers, ages, and genders.… See the full description on the dataset page: https://huggingface.co/datasets/kvest/Swedia-ASR-Dataset.qwen9b-solo-mini-swe-agent
qwen9b-solo-mini-swe-agent
Single-agent coding trajectories generated by running
CooperBench in solo mode on
the CooperData task set, using
Qwen/Qwen3.5-9B as the model and mini_swe_agent_v2 as the agent framework.
One agent implements both features in each task.
The matched coop version is at
CooperBench/qwen9b-coop-mini-swe-agent.
Same task corpus, same model, same agent — only the coordination differs, so
together they isolate the cooperation deficit.
At a glance… See the full description on the dataset page: https://huggingface.co/datasets/CooperBench/qwen9b-solo-mini-swe-agent.SWE-PolyBench_500
SWE-PolyBench
SWE-PolyBench is a multi language repo level software engineering benchmark. Currently it includes 4 languages: Python, Java, Javascript, and Typescript. The number of instances in each language is:
Javascript: 1017
Typescript: 729
Python: 199
Java: 165
Datasets
There are total three datasets available under SWE-PolyBench. AmazonScience/SWE-PolyBench is the full dataset, AmazonScience/SWE-PolyBench_500 is the stratified sampled dataset with 500 instances and… See the full description on the dataset page: https://huggingface.co/datasets/Sellopale/SWE-PolyBench_500.merge-sweep-results
Heterogeneous Mergeability — merge sweep results
Per-merge outcomes from the sweep behind Heterogeneous Mergeability: A Quotient-Space Theory of
When Neural Networks Compose Across Tokenizers and Architectures. Every row is one merge that was
actually executed and scored: a pair of independently trained LMs, a merge operator, and an
alignment arm (merge naively, or align the two models into a shared frame first).
This dataset is a snapshot of a sweep that is still running.… See the full description on the dataset page: https://huggingface.co/datasets/Mergeability/merge-sweep-results.can-it-ford-scenario-sweep
Can It Ford? — AR&R flood-vehicle scenario sweep
70 flood-crossing scenarios (10 depths × 7 velocities), each evaluated under two encodings of
the same published stability criterion and under all three of its vehicle classes.
Part of Can It Ford?, an NSF SCIPE REU 2026 project at the GeoElements Lab, Texas Advanced
Computing Center / UT Austin. Josie Cerrell, Claremont McKenna College.
Findings site: https://huggingface.co/spaces/josiecerrell/can-it-ford-findings
Source code:… See the full description on the dataset page: https://huggingface.co/datasets/josiecerrell/can-it-ford-scenario-sweep.swe-rebench-repro-artifactshf-testsweden_100K_easysweden_1M_easy
