datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
osworld_v2_variants
OSWorld V2 Task Variants
This dataset contains the root-level task_*v1.py Python task classes of the OSWorld V2 variant set used in the DigitalWorld AWM evaluation: 16 parameter-resampled rewrites of OSWorld V2 task classes. Each variant keeps the base task's application and workflow but changes the instruction card and regenerates its assets, so that an agent that has memorised the base task does not get the variant for free.
It follows the layout of xlangai/osworld_v2_tasks:… See the full description on the dataset page: https://huggingface.co/datasets/magicgh/osworld_v2_variants.ScreenSpot-v2-variants
ScreenSpot-v2-variaints
The ScreenSpot dataset with 4 types of instructions:
instruction: original instruction from ScreenSpot,action: clarifies the action to take,
description: describes the target UI element.
negative: an operation that can not be done in the screenshot.
Compatible code can be found in the GitHub repo above.
flip2-multi-sequence-prompt-ablation-generated-variants-2stl_formulae_variantsg1_min_episodes_d1_all_variants_glm47_tracesSWE-Synth_Variants-Metadatayolov8-baseline-mosaic-variants-runsopengloss-v1.3-encyclopedia-variants
See also OpenGloss v2.1 (2026-09-07): a deeper release of 109,633 of these headwords — sense-level ids, four reading levels, sense-tagged examples with spans, a judged relation graph, and retrieval supervision — published as a 16-dataset family. v1.3 remains the broader headword list.
OpenGloss Encyclopedia Variants v1.3
Dataset Summary
OpenGloss Encyclopedia Variants is a synthetic dataset of vocabulary encyclopedia entries
rewritten in multiple writing styles.… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v1.3-encyclopedia-variants.msmarco-icl-100shot-v4-pos_variantshotpotqa-variants
Variant Dataset
Flattened schema: each config has a single answer column.
Column
Description
question
the question
answer
variant text; always clean in the validation split
prediction_lora
LoRA model trained on this answer ("" if unavailable)
prediction_freeze
full SFT model trained on this answer
Variants-catala-cv16_1deprecated-human_variants
Human variants
A curated set of variants from three sources: ClinVar, COSMIC, OMIM and gnomAD.
Predictions for methods benchmarked in GPN-MSA paper can be downloaded from here.
Functional annotations can be downloaded from here.
For more information check out our paper and repository.
Data sources
ClinVar:
Missense variants considered "Pathogenic" by human labelers.
COSMIC:
Somatic missense variants with a frequency at least 0.1% in cancer samples (whole-genome and… See the full description on the dataset page: https://huggingface.co/datasets/songlab/deprecated-human_variants.flip2-multi-sequence-prompt-ablation-generated-variants-amylaseSPSD-Variants-opsd
SPSD-Variants-opsd
Grounded on-policy self-distillation (OPSD) teacher-context dataset over 45
board-game rule variants (5 families × 9: connect4, domineering,
simplified_first_attack, simplified_othello, tic_tac_chess), derived from
trained MuZero/EfficientZero checkpoints (plan-528 v2).
Each row is a decision-state task (a move choice or one of six auxiliary
state-QA tasks). The privileged_context is the teacher signal: grounded
natural-language reasoning that discovers the… See the full description on the dataset page: https://huggingface.co/datasets/LorMolf/SPSD-Variants-opsd.flip2-multi-sequence-prompt-ablation-generated-variants-nucbopenthoughts_dataset_variantsindian-cars-variantsThe data behind variantwise.com: every car model,
trim, engine and variant on sale in India, with prices and a controlled
feature vocabulary, plus the state road tax tables that power on-road price
breakdowns. Exported from the same build that renders the site. A mirror copy like
this one is current as of its release date, stamped in each file's
dataAsOf; the always-current files live at variantwise.com/data.
Data as of: see the dataAsOf field inside each file. New releases are
cut as the… See the full description on the dataset page: https://huggingface.co/datasets/variantwise/indian-cars-variants.utr-variants-bohn
Overview
This utr variant effect prediction task measures the pathogenicity of single nucleotide polymorpism (SNPs) and indels (insertion-deletions) in 3' and 5' UTRs. This dataset is a reprocessing of the dataset from DeepGenomics (Bohn et al. Frontiers in Molecular Biosciences. 2023), see the original paper for the data generation process. We provide the mRNA transcript sequence context for the variant and wildtype sequences using the specified transcript id from RefSeq v109.
This… See the full description on the dataset page: https://huggingface.co/datasets/morrislab/utr-variants-bohn.mutvar-variants2026.RA.Agent-Variants-Closure
2026.RA.Agent-Variants-Closure
Episode and turn records from a design loop that took a five-seat private-information negotiation
agent from a 0.233 deal rate to 1.000, and the controls that say why.
What this is
Five automated negotiators must agree unanimously on one package out of 256 within four rounds. Each
holds a private score sheet (what each package is worth to it) and a private threshold (the value below
which it prefers no deal). The published agent… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/2026.RA.Agent-Variants-Closure.fgvc-aircraft-variantsterminal_bench_2_g1_d1_all_variants_32b_step300_20260425_175030freshqa-variants
Variant Dataset
Flattened schema: each config has a single answer column.
Column
Description
question
the question
answer
variant text; always clean in the validation split
prediction_lora
LoRA model trained on this answer ("" if unavailable)
prediction_freeze
full SFT model trained on this answer
RoCoLe-Variantsterminal_bench_2_g1_d1_all_variants_32b_20260427_160312bpcc_marathi_punct_variantsbigcodebench-typo-variants
BigCodeBench Typo Variants
This dataset contains typo-injected variants of the BigCodeBench coding benchmark to evaluate the robustness of code generation models to typographical errors in problem descriptions.
Dataset Description
BigCodeBench is a benchmark for evaluating large language models on diverse and challenging coding tasks. This dataset provides 7 variants with different levels of typos injected into the instruction prompts:
Original (0% typos): Clean baseline… See the full description on the dataset page: https://huggingface.co/datasets/jeqcho/bigcodebench-typo-variants.terminal_bench_2_g1_d1_all_variants_32b_step1500_20260426_153557simple-variants
Variant Dataset
Flattened schema: each config has a single answer column.
Column
Description
question
the question
answer
variant text; always clean in the validation split
prediction_lora
LoRA model trained on this answer ("" if unavailable)
prediction_freeze
full SFT model trained on this answer
rtllm-variants
RTLLM Variants (Community Dataset Derivatives)
This repository hosts community-maintained derivative variants based on upstream RTLLM releases. It is intended for benchmarking reproducibility and evaluation workflow integration.
Important Notice
This repository is not an official release from the RTLLM paper authors.
Variants in this repository may include prompt wording normalization, interface naming unification, or testbench output standardization for evaluator… See the full description on the dataset page: https://huggingface.co/datasets/xxrjun/rtllm-variants.
