datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SWE-bench-devin-fullSWE-bench-devin-passedSWE-bench-devin-full-filtereddolma-v1_6-sampleTaken directly from allenai/dolma. Roughly 10B tokens. Useful for small scale testing or maybe continued pretraining.
Devin-SWE-bench-outputdevin-cli-reasoning-distillation
Devin CLI Reasoning Distillation Dataset
A distillation dataset built from Devin CLI session traces, containing the model's internal
reasoning traces (chain-of-thought / thinking), user prompts, assistant answers, and tool calls.
The dataset is formatted to be directly compatible with SFT training pipelines that expect
OpenAI-style message lists with a reasoning_content field.
Dataset Summary
Total rows
2,632 (2,507 train / 125 validation)
Rows with… See the full description on the dataset page: https://huggingface.co/datasets/Akahsizrr/devin-cli-reasoning-distillation.Gradient-Decomposition-Assay
Gradient Decomposition Assay (GDA)
This repository contains the CSV corpus and summary tables for the Gradient Decomposition Assay, an exploratory behavioral evaluation of how frontier language models respond to an eight-vector prompt manifold ranging from benign technical tasks to adversarial compression and counterfactual/narrative reframing.
Why this repository uses multiple configurations
The CSV files in this dataset are not all the same table. The row-level… See the full description on the dataset page: https://huggingface.co/datasets/devinendorphin/Gradient-Decomposition-Assay.iati-policy-markers
International Aid Transparency Initiative (IATI) Policy Marker Dataset
A multi-purpose dataset including all activity title and description text published to IATI with metadata for policy markers.
For more information on IATI policy markers, see the element page on the IATI Standard Website.
IATI is a living data source, and this dataset was last updated on 21 August, 2024. For the code to generate an updated version of this dataset, please see my Github repository here.
For any… See the full description on the dataset page: https://huggingface.co/datasets/devinitorg/iati-policy-markers.cdp-paf-meta-limited-syntheticwb-climate-percentagecdp-paf-meta-limitedscotland-bus-reliability-2026dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-5-relabel-v5
dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-5-relabel-v5 — v5 relabel + split
Copy of aletheias-quest/dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-5 with the v5 belief-relative label (see reinthal/aletheias-dev-relabel-v5 for the
method: 20x neutral resample -> DeepSeek-V4-Flash judge, no canonicalization) and a
train/test/validation split column.
Added columns: deceptive (v5 label; official fallback where excluded), official (original dev
label), relabeled… See the full description on the dataset page: https://huggingface.co/datasets/reinthal/dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-5-relabel-v5.dev-instructed-deception-Qwen3.5-27B-Nonedev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-6-relabel-v5
dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-6-relabel-v5 — v5 relabel + split
Copy of aletheias-quest/dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-6 with the v5 belief-relative label (see reinthal/aletheias-dev-relabel-v5 for the
method: 20x neutral resample -> DeepSeek-V4-Flash judge, no canonicalization) and a
train/test/validation split column.
Added columns: deceptive (v5 label; official fallback where excluded), official (original dev
label), relabeled… See the full description on the dataset page: https://huggingface.co/datasets/reinthal/dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-6-relabel-v5.dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-3-relabel-v5
dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-3-relabel-v5 — v5 relabel + split
Copy of aletheias-quest/dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-3 with the v5 belief-relative label (see reinthal/aletheias-dev-relabel-v5 for the
method: 20x neutral resample -> DeepSeek-V4-Flash judge, no canonicalization) and a
train/test/validation split column.
Added columns: deceptive (v5 label; official fallback where excluded), official (original dev
label), relabeled… See the full description on the dataset page: https://huggingface.co/datasets/reinthal/dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-3-relabel-v5.devin-2024-03-24-17-25-22dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-1dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-4-relabel-v5
dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-4-relabel-v5 — v5 relabel + split
Copy of aletheias-quest/dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-4 with the v5 belief-relative label (see reinthal/aletheias-dev-relabel-v5 for the
method: 20x neutral resample -> DeepSeek-V4-Flash judge, no canonicalization) and a
train/test/validation split column.
Added columns: deceptive (v5 label; official fallback where excluded), official (original dev
label), relabeled… See the full description on the dataset page: https://huggingface.co/datasets/reinthal/dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-4-relabel-v5.dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-7-relabel-v5
dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-7-relabel-v5 — v5 relabel + split
Copy of aletheias-quest/dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-7 with the v5 belief-relative label (see reinthal/aletheias-dev-relabel-v5 for the
method: 20x neutral resample -> DeepSeek-V4-Flash judge, no canonicalization) and a
train/test/validation split column.
Added columns: deceptive (v5 label; official fallback where excluded), official (original dev
label), relabeled… See the full description on the dataset page: https://huggingface.co/datasets/reinthal/dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-7-relabel-v5.dev-instructed-deception-Qwen3.5-27B-c-mo-qwen3.5-27b-relabel-v5
dev-instructed-deception-Qwen3.5-27B-c-mo-qwen3.5-27b-relabel-v5 — v5 relabel + split
Copy of aletheias-quest/dev-instructed-deception-Qwen3.5-27B-c-mo-qwen3.5-27b with the v5 belief-relative label (see reinthal/aletheias-dev-relabel-v5 for the
method: 20x neutral resample -> DeepSeek-V4-Flash judge, no canonicalization) and a
train/test/validation split column.
Added columns: deceptive (v5 label; official fallback where excluded), official (original dev
label), relabeled (v5 !=… See the full description on the dataset page: https://huggingface.co/datasets/reinthal/dev-instructed-deception-Qwen3.5-27B-c-mo-qwen3.5-27b-relabel-v5.dev-instructed-deception-gemma-3-27b-it-None-relabel-v5
dev-instructed-deception-gemma-3-27b-it-None-relabel-v5 — v5 relabel + split
Copy of aletheias-quest/dev-instructed-deception-gemma-3-27b-it-None with the v5 belief-relative label (see reinthal/aletheias-dev-relabel-v5 for the
method: 20x neutral resample -> DeepSeek-V4-Flash judge, no canonicalization) and a
train/test/validation split column.
Added columns: deceptive (v5 label; official fallback where excluded), official (original dev
label), relabeled (v5 != official)… See the full description on the dataset page: https://huggingface.co/datasets/reinthal/dev-instructed-deception-gemma-3-27b-it-None-relabel-v5.dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-3dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-6dev-instructed-deception-gemma-3-27b-it-s-mo-gemma-3-27b-itdev-instructed-deception-NVIDIA-Nemotron-3-Super-120B-A12B-BF16-Nonedev-instructed-deception-gemma-3-27b-it-Nonedev-instructed-deception-Qwen3.5-27B-b-mo-qwen3.5-27bdev-instructed-deception-NVIDIA-Nemotron-3-Super-120B-A12B-BF16-None-relabel-v5
dev-instructed-deception-NVIDIA-Nemotron-3-Super-120B-A12B-BF16-None-relabel-v5 — v5 relabel + split
Copy of aletheias-quest/dev-instructed-deception-NVIDIA-Nemotron-3-Super-120B-A12B-BF16-None with the v5 belief-relative label (see reinthal/aletheias-dev-relabel-v5 for the
method: 20x neutral resample -> DeepSeek-V4-Flash judge, no canonicalization) and a
train/test/validation split column.
Added columns: deceptive (v5 label; official fallback where excluded), official… See the full description on the dataset page: https://huggingface.co/datasets/reinthal/dev-instructed-deception-NVIDIA-Nemotron-3-Super-120B-A12B-BF16-None-relabel-v5.dev-instructed-deception-Qwen3.5-27B-None-relabel-v5
dev-instructed-deception-Qwen3.5-27B-None-relabel-v5 — v5 relabel + split
Copy of aletheias-quest/dev-instructed-deception-Qwen3.5-27B-None with the v5 belief-relative label (see reinthal/aletheias-dev-relabel-v5 for the
method: 20x neutral resample -> DeepSeek-V4-Flash judge, no canonicalization) and a
train/test/validation split column.
Added columns: deceptive (v5 label; official fallback where excluded), official (original dev
label), relabeled (v5 != official), excluded… See the full description on the dataset page: https://huggingface.co/datasets/reinthal/dev-instructed-deception-Qwen3.5-27B-None-relabel-v5.
