datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
warp-taskgen-generated-ipi-tasks-50
WARP Taskgen Generated IPI Tasks 50
Dataset Summary
This dataset contains WARP Taskgen Phase 4 browser-agent trajectories for a
50-task generated indirect prompt injection (IPI) cohort. The trajectories were
produced with the AgentLab harness on
WebArena GitLab and Postmill (Reddit) benchmark applications.
The export is a report-only projection of already written benchmark artifacts.
It does not alter scoring, PVPO encounter checks, rewards, admission, or
trajectory… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-dataset-submission-warp/warp-taskgen-generated-ipi-tasks-50.SWE-ChainAnonymous_Submission_Seed
SETA-Synth Seed Data
Seed data collected from technical Q&A platforms, programming communities, and command-line instruction resources, used as source material for the SETA-Synth pipeline.
This repository contains the seed-side inputs for synthesizing terminal-agent tasks and environments.
Dataset Structure
The dataset is organised by source, then by seed ID or source-specific item ID:
{source}/
└── {item_id}/
├── main.json # seed record for Q&A and… See the full description on the dataset page: https://huggingface.co/datasets/AnonymousSubmissionUnderDouble-BlindRevi/Anonymous_Submission_Seed.submission14717_fictionalqa_reformatted_triviaqa
Reformatted TriviaQA for use alongside FictionalQA
Repository: omitted
Paper: omitted
Dataset Description
This dataset is a simple derived view of the validation data from the original TriviaQA dataset hosted by the original creators at hf.co/datasets/mandarjoshi/trivia_qa. To create this view, we extract the wikipedia articles associated with each question, as well as a simplified answer list, and then we create a few versions of the resulting data for use as… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-aardvark/submission14717_fictionalqa_reformatted_triviaqa.submission14717_fictionalqa
The FictionalQA dataset
Repository: omitted
Paper: omitted
Dataset Summary
The FictionalQA dataset is a dataset specifically created to empower researchers to study the dual processes of fact memorization and verbatim sequence memorization. The dataset consists of synthetically-generated, webtext-like documents about fictional events and various facts they entail, as well as question-answer pairs about the facts within the fictional documents.
Supported Tasks… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-aardvark/submission14717_fictionalqa.submission14717_fictionalqa_training_splits
Training splits view of the FictionalQA dataset
The FictionalQA dataset
Repository: omitted
Paper: omitted
Dataset Description
This dataset is a derivative of the main dataset. Please see that dataset's README for a detailed description of the assets.
The dataset splits (configs) provided here are the exact ones materialized and used in the experiments for the associated paper. The primary purpose of this dataset repository is for transparency and to help… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-aardvark/submission14717_fictionalqa_training_splits.pixels_vs_code
Pattern Over Pixels Screenshot-to-Code
This dataset contains controlled counterfactual screenshot-to-code examples built
from 30 real-world webpages from Design2Code. Each example preserves a repeated
UI pattern while introducing a single localized deviation, allowing researchers
to test whether multimodal models follow the pixels or simply restore the
dominant template.
Contents
720 perturbed HTML instances
360 structural-card examples
360 text-style examples
2… See the full description on the dataset page: https://huggingface.co/datasets/AnonymousSubmissionASE/pixels_vs_code.
