EmpathicRobotics/imagenet-variations-synth-pilot-v3
ImageNet Variations Synth Pilot v3.1 (diverse, 1K) Huu + Claude multimodal instruction pipeline — quality-fixed re-run. Pipeline Flux.1-schnell (Huu style phrasings + visual style boosts + aspect ratios) → Seed2 → Florence-2 grounding/caption → tracks A–F → USER / ASSISTANT. Tracks A programmatic QA (counts/style/spatial/absence) from detector B LLaVA-style (conversation / detailed description / complex reasoning) C Evol-Instruct (seed atypical Q →… See the full description on the dataset page: https://huggingface.co/datasets/EmpathicRobotics/imagenet-variations-synth-pilot-v3.
ImageNet Variations Synth Pilot v3.1 (diverse, 1K)
Huu + Claude multimodal instruction pipeline — quality-fixed re-run.
Pipeline
Flux.1-schnell (Huu style phrasings + visual style boosts + aspect ratios) → Seed2 → Florence-2 grounding/caption → tracks A–F → USER / ASSISTANT.
Tracks
- A programmatic QA (counts/style/spatial/absence) from detector
- B LLaVA-style (conversation / detailed description / complex reasoning)
- C Evol-Instruct (seed atypical Q → 1 evolve op)
- D mood / people / generation-from-image / math from verified counts only
- E contrastive pairs (2 styles, multi-
<seed2>) - F unanswerable / refusal (kept small; no hedge→F dump)
v3.1 fixes
- Hedge answers cleaned / fall back to A (not F)
- D1 math numbers only from Florence counts
- No “detected object” leak in questions
- Stronger Flux style boosts for webpage / book / poster
- E4 answers describe style difference (do not paste base_prompt)
Format
USER: <seed2> ... </seed2> {question} ASSISTANT: {answer}Files
data/pilot.jsonl— 1000 instruction records (id,text,image_path,metadata)images/— generated PNGs (incl.*_b.pngfor contrastive pairs)
