control
Datasets
All datasets matching “control”vllm-control-arena
vLLM Main Tasks Dataset
AI coding tasks generated from vLLM git commits
Dataset Description
This dataset contains 6801 coding tasks automatically generated from git commits in the vLLM repository. Each task represents a real-world coding challenge derived from actual development work.
Dataset Structure
The dataset contains the following columns:
commit_hash: The git commit hash
parent_hash: The parent commit hash
commit_title: The original commit… See the full description on the dataset page: https://huggingface.co/datasets/RoganInglis/vllm-control-arena.apps-control-arena
APPS Control Arena Dataset
Unified dataset combining APPS problems with backdoors from both the AI Control paper and Control-Tax paper.
Dataset Description
This dataset is based on the codeparrot/apps dataset,
enhanced with backdoor solutions from two sources:
APPS Backdoors: From "AI Control: Improving Safety Despite Intentional Subversion" / TylordTheGreat/apps-backdoors-04-02-25
Control-Tax Backdoors: From "Control Tax: The Price of Keeping AI in Check"… See the full description on the dataset page: https://huggingface.co/datasets/RoganInglis/apps-control-arena.Light100K
Dataset Card for Light100K (Parquet)
This release is organized as Parquet shards with one aligned sample per row.
Each row contains the low-light input (control) and five enhancement targets
(target_l01 ... target_l05) as Hugging Face image columns.
Advantages of this release
Hub Dataset Viewer can preview aligned columns on the same row.
Better sample-level random access than tar-per-folder packaging.
Keeps original PNG bytes; no image decode/re-encode during export.… See the full description on the dataset page: https://huggingface.co/datasets/ControlLight/Light100K.control-pretraining-datasets-smoke
geodesic-research/control-pretraining-datasets-smoke
Auto-generated by dataset-builder.
Each config below is a separate dataset produced from a versioned YAML build
config. Load with:
from datasets import load_dataset
ds = load_dataset("geodesic-research/control-pretraining-datasets-smoke", "<config_name>", revision="<commit-sha>")
Pin revision= to the specific commit SHA you want; without it, you get the
current HEAD of the dataset repo, which may change when the builder… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/control-pretraining-datasets-smoke.2026-09-11-dh-qwen3-6-27b-lora-9284-numina-control-716-r64
Delegated-harm evaluation with corrected scoring of saved rollouts
field
value
experiment
Delegated-harm evaluation with corrected scoring of saved rollouts
date_generated
2026-09-11
constitution
none
source_repo
teaching_claude_why_replication @ d627d0587a2980069b7700e72f727dae594c9f49
models
{"hf_path": "matboz/qwen3.6-27b-lora-9284-numina-control-716-r64", "base_model": "Qwen/Qwen3.6-27B", "adapter": true, "mode": "think", "model_key":… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-11-dh-qwen3-6-27b-lora-9284-numina-control-716-r64.amateur_drawings-controlnet-dataset
Dataset Card for "amateur_drawings-controlnet-dataset"
WIP... Come back later....
