datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
devin-cli-reasoning-distillation
Devin CLI Reasoning Distillation Dataset
A distillation dataset built from Devin CLI session traces, containing the model's internal
reasoning traces (chain-of-thought / thinking), user prompts, assistant answers, and tool calls.
The dataset is formatted to be directly compatible with SFT training pipelines that expect
OpenAI-style message lists with a reasoning_content field.
Dataset Summary
Total rows
2,632 (2,507 train / 125 validation)
Rows with… See the full description on the dataset page: https://huggingface.co/datasets/Akahsizrr/devin-cli-reasoning-distillation.iati-policy-markers
International Aid Transparency Initiative (IATI) Policy Marker Dataset
A multi-purpose dataset including all activity title and description text published to IATI with metadata for policy markers.
For more information on IATI policy markers, see the element page on the IATI Standard Website.
IATI is a living data source, and this dataset was last updated on 21 August, 2024. For the code to generate an updated version of this dataset, please see my Github repository here.
For any… See the full description on the dataset page: https://huggingface.co/datasets/devinitorg/iati-policy-markers.wb-climate-percentageXlerobot-grap-post-index4This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "xlerobot_follower",
"total_episodes": 101,
"total_frames": 34601,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:101"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/monozu-deving/Xlerobot-grap-post-index4.dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-5-relabel-v5
dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-5-relabel-v5 — v5 relabel + split
Copy of aletheias-quest/dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-5 with the v5 belief-relative label (see reinthal/aletheias-dev-relabel-v5 for the
method: 20x neutral resample -> DeepSeek-V4-Flash judge, no canonicalization) and a
train/test/validation split column.
Added columns: deceptive (v5 label; official fallback where excluded), official (original dev
label), relabeled… See the full description on the dataset page: https://huggingface.co/datasets/reinthal/dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-5-relabel-v5.dev-instructed-deception-Qwen3.5-27B-Nonedev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-6-relabel-v5
dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-6-relabel-v5 — v5 relabel + split
Copy of aletheias-quest/dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-6 with the v5 belief-relative label (see reinthal/aletheias-dev-relabel-v5 for the
method: 20x neutral resample -> DeepSeek-V4-Flash judge, no canonicalization) and a
train/test/validation split column.
Added columns: deceptive (v5 label; official fallback where excluded), official (original dev
label), relabeled… See the full description on the dataset page: https://huggingface.co/datasets/reinthal/dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-6-relabel-v5.dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-3-relabel-v5
dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-3-relabel-v5 — v5 relabel + split
Copy of aletheias-quest/dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-3 with the v5 belief-relative label (see reinthal/aletheias-dev-relabel-v5 for the
method: 20x neutral resample -> DeepSeek-V4-Flash judge, no canonicalization) and a
train/test/validation split column.
Added columns: deceptive (v5 label; official fallback where excluded), official (original dev
label), relabeled… See the full description on the dataset page: https://huggingface.co/datasets/reinthal/dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-3-relabel-v5.Xlerobot-grap-post-index4-fillwith2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "xlerobot_follower",
"total_episodes": 40,
"total_frames": 12876,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:40"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/monozu-deving/Xlerobot-grap-post-index4-fillwith2.dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-1dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-4-relabel-v5
dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-4-relabel-v5 — v5 relabel + split
Copy of aletheias-quest/dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-4 with the v5 belief-relative label (see reinthal/aletheias-dev-relabel-v5 for the
method: 20x neutral resample -> DeepSeek-V4-Flash judge, no canonicalization) and a
train/test/validation split column.
Added columns: deceptive (v5 label; official fallback where excluded), official (original dev
label), relabeled… See the full description on the dataset page: https://huggingface.co/datasets/reinthal/dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-4-relabel-v5.dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-7-relabel-v5
dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-7-relabel-v5 — v5 relabel + split
Copy of aletheias-quest/dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-7 with the v5 belief-relative label (see reinthal/aletheias-dev-relabel-v5 for the
method: 20x neutral resample -> DeepSeek-V4-Flash judge, no canonicalization) and a
train/test/validation split column.
Added columns: deceptive (v5 label; official fallback where excluded), official (original dev
label), relabeled… See the full description on the dataset page: https://huggingface.co/datasets/reinthal/dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-7-relabel-v5.dev-instructed-deception-Qwen3.5-27B-c-mo-qwen3.5-27b-relabel-v5
dev-instructed-deception-Qwen3.5-27B-c-mo-qwen3.5-27b-relabel-v5 — v5 relabel + split
Copy of aletheias-quest/dev-instructed-deception-Qwen3.5-27B-c-mo-qwen3.5-27b with the v5 belief-relative label (see reinthal/aletheias-dev-relabel-v5 for the
method: 20x neutral resample -> DeepSeek-V4-Flash judge, no canonicalization) and a
train/test/validation split column.
Added columns: deceptive (v5 label; official fallback where excluded), official (original dev
label), relabeled (v5 !=… See the full description on the dataset page: https://huggingface.co/datasets/reinthal/dev-instructed-deception-Qwen3.5-27B-c-mo-qwen3.5-27b-relabel-v5.dev-instructed-deception-gemma-3-27b-it-None-relabel-v5
dev-instructed-deception-gemma-3-27b-it-None-relabel-v5 — v5 relabel + split
Copy of aletheias-quest/dev-instructed-deception-gemma-3-27b-it-None with the v5 belief-relative label (see reinthal/aletheias-dev-relabel-v5 for the
method: 20x neutral resample -> DeepSeek-V4-Flash judge, no canonicalization) and a
train/test/validation split column.
Added columns: deceptive (v5 label; official fallback where excluded), official (original dev
label), relabeled (v5 != official)… See the full description on the dataset page: https://huggingface.co/datasets/reinthal/dev-instructed-deception-gemma-3-27b-it-None-relabel-v5.dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-3dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-6dev-instructed-deception-gemma-3-27b-it-s-mo-gemma-3-27b-itdev-instructed-deception-NVIDIA-Nemotron-3-Super-120B-A12B-BF16-Nonedev-instructed-deception-gemma-3-27b-it-Nonedev-instructed-deception-Qwen3.5-27B-b-mo-qwen3.5-27bdev-instructed-deception-NVIDIA-Nemotron-3-Super-120B-A12B-BF16-None-relabel-v5
dev-instructed-deception-NVIDIA-Nemotron-3-Super-120B-A12B-BF16-None-relabel-v5 — v5 relabel + split
Copy of aletheias-quest/dev-instructed-deception-NVIDIA-Nemotron-3-Super-120B-A12B-BF16-None with the v5 belief-relative label (see reinthal/aletheias-dev-relabel-v5 for the
method: 20x neutral resample -> DeepSeek-V4-Flash judge, no canonicalization) and a
train/test/validation split column.
Added columns: deceptive (v5 label; official fallback where excluded), official… See the full description on the dataset page: https://huggingface.co/datasets/reinthal/dev-instructed-deception-NVIDIA-Nemotron-3-Super-120B-A12B-BF16-None-relabel-v5.dev-instructed-deception-Qwen3.5-27B-None-relabel-v5
dev-instructed-deception-Qwen3.5-27B-None-relabel-v5 — v5 relabel + split
Copy of aletheias-quest/dev-instructed-deception-Qwen3.5-27B-None with the v5 belief-relative label (see reinthal/aletheias-dev-relabel-v5 for the
method: 20x neutral resample -> DeepSeek-V4-Flash judge, no canonicalization) and a
train/test/validation split column.
Added columns: deceptive (v5 label; official fallback where excluded), official (original dev
label), relabeled (v5 != official), excluded… See the full description on the dataset page: https://huggingface.co/datasets/reinthal/dev-instructed-deception-Qwen3.5-27B-None-relabel-v5.dev-instructed-deception-Qwen3.5-27B-c-mo-qwen3.5-27bdev-instructed-deception-Qwen3.5-27B-g-st-qwen3.5-27bdev-instructed-deception-gemma-3-27b-it-s-mo-gemma-3-27b-it-relabel-v5
dev-instructed-deception-gemma-3-27b-it-s-mo-gemma-3-27b-it-relabel-v5 — v5 relabel + split
Copy of aletheias-quest/dev-instructed-deception-gemma-3-27b-it-s-mo-gemma-3-27b-it with the v5 belief-relative label (see reinthal/aletheias-dev-relabel-v5 for the
method: 20x neutral resample -> DeepSeek-V4-Flash judge, no canonicalization) and a
train/test/validation split column.
Added columns: deceptive (v5 label; official fallback where excluded), official (original dev
label)… See the full description on the dataset page: https://huggingface.co/datasets/reinthal/dev-instructed-deception-gemma-3-27b-it-s-mo-gemma-3-27b-it-relabel-v5.dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-5dev-instructed-deception-gemma-3-27b-it-g-st-gemma-3-27b-it-2dev-instructed-deception-Qwen3.5-27B-b-mo-qwen3.5-27b-relabel-v5
dev-instructed-deception-Qwen3.5-27B-b-mo-qwen3.5-27b-relabel-v5 — v5 relabel + split
Copy of aletheias-quest/dev-instructed-deception-Qwen3.5-27B-b-mo-qwen3.5-27b with the v5 belief-relative label (see reinthal/aletheias-dev-relabel-v5 for the
method: 20x neutral resample -> DeepSeek-V4-Flash judge, no canonicalization) and a
train/test/validation split column.
Added columns: deceptive (v5 label; official fallback where excluded), official (original dev
label), relabeled (v5 !=… See the full description on the dataset page: https://huggingface.co/datasets/reinthal/dev-instructed-deception-Qwen3.5-27B-b-mo-qwen3.5-27b-relabel-v5.dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-4dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-7
