datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
localized_narratives_trajectory_formatfunc_localize_claude45_1457icocoqa_localized_narratives
cocoqa_localized_narratives
Description
Concatenated dataset of all of cocoqa and localized_narratives.
Processing Parameters
{}
Dataset Configuration
Train dataset:
mixer: mateoguaman/cocoqa_trajectory_format: 1.0
mateoguaman/localized_narratives_trajectory_format: 1.0
split: train
Validation dataset:
mixer: mateoguaman/cocoqa_trajectory_format: 1.0
mateoguaman/localized_narratives_trajectory_format: 1.0
split: train… See the full description on the dataset page: https://huggingface.co/datasets/mateoguaman/cocoqa_localized_narratives.func_localize_claude47_1467ilocalize-indoor
Elliot Localize Indoor / 3D and depth
Upstream training splits; known explicitly identified test/eval rows excluded. Cross-dataset benchmark overlap is not guaranteed. Published as a raw, manually gated release; source annotation caveats remain.
Task views reuse original image archives. 2D coordinates are normalized 0–1000. Native 3D sidecars preserve original camera-space XYZ and camera calibration; they are not normalized to 0–1000. Within each query targets are sorted… See the full description on the dataset page: https://huggingface.co/datasets/elliot-mllm/localize-indoor.swe_bench_localize_sim_prompt_515ilocalized_narrativesmm_localized_narrativeslocalize-gui
Elliot Localize GUI
Upstream training splits; known explicitly identified test/eval rows excluded. Cross-dataset benchmark overlap is not guaranteed. Published as a raw, manually gated release; source annotation caveats remain.
Task views reuse original image archives. Coordinates are normalized 0–1000. Within each query targets are sorted left-to-right then top-to-bottom.
HF preview configs contain 10 examples per view, not the complete training split. Full training uses… See the full description on the dataset page: https://huggingface.co/datasets/elliot-mllm/localize-gui.func_localize_claude45_1457i_text2x
func_localize_claude45_1457i_text2x
Verbosity-ablation variant of synthetic-code-training/func_localize_claude45_1457i: the prose before every tool call is rewritten by nvidia/deepseek-ai/deepseek-v4-flash to 2 times its own length in Qwen3 tokens (accepted band ±25 %, up to 3 rounds). Of 26194 turns, 15684 were not rephrased (empty prose, or a target outside 4-2000 tokens) and 130 missed the band; both keep their original prose.
Construction (shared by the _text0x/0.5x/2x/4x/8x… See the full description on the dataset page: https://huggingface.co/datasets/synthetic-code-training/func_localize_claude45_1457i_text2x.func_localize_claude47_min_file_explore_1467ifunc_localize_claude45_1457i_text4x
func_localize_claude45_1457i_text4x
Verbosity-ablation variant of synthetic-code-training/func_localize_claude45_1457i: the prose before every tool call is rewritten by nvidia/deepseek-ai/deepseek-v4-flash to 4 times its own length in Qwen3 tokens (accepted band ±25 %, up to 3 rounds). Of 26194 turns, 15698 were not rephrased (empty prose, or a target outside 4-2000 tokens) and 297 missed the band; both keep their original prose.
Construction (shared by the _text0x/0.5x/2x/4x/8x… See the full description on the dataset page: https://huggingface.co/datasets/synthetic-code-training/func_localize_claude45_1457i_text4x.localize-sympynew-gpt4o-v1func_localize_claude45_1457i_text0.5x
func_localize_claude45_1457i_text0.5x
Verbosity-ablation variant of synthetic-code-training/func_localize_claude45_1457i: the prose before every tool call is rewritten by nvidia/deepseek-ai/deepseek-v4-flash to 0.5 times its own length in Qwen3 tokens (accepted band ±25 %, up to 3 rounds). Of 26194 turns, 15689 were not rephrased (empty prose, or a target outside 4-2000 tokens) and 79 missed the band; both keep their original prose.
Construction (shared by the… See the full description on the dataset page: https://huggingface.co/datasets/synthetic-code-training/func_localize_claude45_1457i_text0.5x.func_localize_claude45_1457i_text8x
func_localize_claude45_1457i_text8x
Verbosity-ablation variant of synthetic-code-training/func_localize_claude45_1457i: the prose before every tool call is rewritten by nvidia/deepseek-ai/deepseek-v4-flash to 8 times its own length in Qwen3 tokens (accepted band ±25 %, up to 3 rounds). Of 26194 turns, 15754 were not rephrased (empty prose, or a target outside 4-2000 tokens) and 360 missed the band; both keep their original prose.
Construction (shared by the _text0x/0.5x/2x/4x/8x… See the full description on the dataset page: https://huggingface.co/datasets/synthetic-code-training/func_localize_claude45_1457i_text8x.ds3_swe_bench_localize_513ifunc_localize_claude45_1457i_text300
func_localize_claude45_1457i_text300
Verbosity-ablation variant of synthetic-code-training/func_localize_claude45_1457i: the prose before every tool call is rewritten by nvidia/deepseek-ai/deepseek-v4-flash to about 300 tokens (accepted band 225-375 tokens of the Qwen3 tokenizer, up to 3 rounds; 36 of 26194 turns missed the band and keep their original prose).
Construction (shared by all _text* siblings): think and task_tracker turns and their result turns are
removed… See the full description on the dataset page: https://huggingface.co/datasets/synthetic-code-training/func_localize_claude45_1457i_text300.func_localize_claude45_1457i_text0
func_localize_claude45_1457i_text0
Verbosity-ablation variant of synthetic-code-training/func_localize_claude45_1457i: every assistant turn is its tool call only: the prose before the call is removed.
Construction (shared by all _text* siblings): think and task_tracker turns and their result turns are
removed (trajectories contain only real tool calls); the system prompt, task, tool calls and tool results are
byte-identical to the base. The rephraser saw only the current turn… See the full description on the dataset page: https://huggingface.co/datasets/synthetic-code-training/func_localize_claude45_1457i_text0.func_localize_claude45_1457i_text100
func_localize_claude45_1457i_text100
Verbosity-ablation variant of synthetic-code-training/func_localize_claude45_1457i: the prose before every tool call is rewritten by nvidia/deepseek-ai/deepseek-v4-flash to about 100 tokens (accepted band 75-125 tokens of the Qwen3 tokenizer, up to 3 rounds; 0 of 26194 turns missed the band and keep their original prose).
Construction (shared by all _text* siblings): think and task_tracker turns and their result turns are
removed… See the full description on the dataset page: https://huggingface.co/datasets/synthetic-code-training/func_localize_claude45_1457i_text100.func_localize_claude45_1457i_text20
func_localize_claude45_1457i_text20
Verbosity-ablation variant of synthetic-code-training/func_localize_claude45_1457i: the prose before every tool call is rewritten by nvidia/deepseek-ai/deepseek-v4-flash to about 20 tokens (accepted band 15-25 tokens of the Qwen3 tokenizer, up to 3 rounds; 178 of 26194 turns missed the band and keep their original prose).
Construction (shared by all _text* siblings): think and task_tracker turns and their result turns are
removed (trajectories… See the full description on the dataset page: https://huggingface.co/datasets/synthetic-code-training/func_localize_claude45_1457i_text20.localize-general-objects
Elliot Localize General Objects
Upstream training splits; known explicitly identified test/eval rows excluded. Cross-dataset benchmark overlap is not guaranteed. Published as a raw, manually gated release; source annotation caveats remain.
Task views reuse original image archives. Coordinates are normalized 0–1000. Within each query targets are sorted left-to-right then top-to-bottom.
HF preview configs contain 10 examples per view, not the complete training split. Full training… See the full description on the dataset page: https://huggingface.co/datasets/elliot-mllm/localize-general-objects.vqasynth_cauldron_localized_narratives_100_fullfunc_localize_deepseek_v4_flash_1374ilocalize-aerial
Hugging Face preview scope
The HF dataset viewer exposes 10 verified examples per stored view, in explicitly named *_preview configs and a preview split. These are not the full training split. Select a config to inspect original images, separate box overlays, full grouped queries, xyxy boxes on the 0–1000 grid, and record IDs. No annotations are truncated in the files. DOTA v1 and full v2 previews are alternatives, not additions to the default training mixture.
The complete… See the full description on the dataset page: https://huggingface.co/datasets/elliot-mllm/localize-aerial.caption-localized_narratives
Description
This dataset is a processed version of the Localized Narratives dataset by Pont-Tuset et al., particularly for a visual question answering task where answer is a caption.We converted the images to PIL format (image column of the dataset); created 64 different questions in French (question column); and finally translated the original captions from English to French and used them as answers (answer column).
Citation
Localized Narratives… See the full description on the dataset page: https://huggingface.co/datasets/lbourdois/caption-localized_narratives.swe_bench_localize_no_exec_498ilocalize-driving
Elliot Localize Driving / street scenes
Upstream training splits; known explicitly identified test/eval rows excluded. Cross-dataset benchmark overlap is not guaranteed. Published as a raw, manually gated release; source annotation caveats remain.
Task views reuse original image archives. Coordinates are normalized 0–1000. Within each query targets are sorted left-to-right then top-to-bottom.
HF preview configs contain 10 examples per view, not the complete training split. Full… See the full description on the dataset page: https://huggingface.co/datasets/elliot-mllm/localize-driving.func_localize_claude45_1457i_text0x
func_localize_claude45_1457i_text0x
Verbosity-ablation variant of synthetic-code-training/func_localize_claude45_1457i: every assistant turn other than think / task_tracker is its tool call only: the prose before the call is removed, while the thinking/planning steps stay.
Construction (shared by the _text0x/0.5x/2x/4x/8x siblings): think and task_tracker turns are kept
verbatim, as are the system prompt, task, tool calls and tool results. For the rephrased siblings, the… See the full description on the dataset page: https://huggingface.co/datasets/synthetic-code-training/func_localize_claude45_1457i_text0x.vamos_10pct_gpt5_mini_cocoqa_localized_narratives
vamos_10pct_gpt5_mini_cocoqa_localized_narratives
Description
VLN Navigation dataset with 100% of tartandrive data, 50% of scand data, 25% of coda data, 100% of in-domain spot data, and 10% of annotated/augmented data using gpt5-mini, and all of cocoqa and localized_narratives. Whenever daatsets aren't 100%, they are ranked by curvature and output of length 5.
Processing Parameters
{}
Dataset Configuration
Train dataset:
mixer:… See the full description on the dataset page: https://huggingface.co/datasets/mateoguaman/vamos_10pct_gpt5_mini_cocoqa_localized_narratives.swe_gym_raw_localize_no_exec_claude37_441i
