CoolFace
Datasetpublic

EmpathicRobotics/imagenet-variations-synth-pilot-v3

ImageNet Variations Synth Pilot v3.1 (diverse, 1K) Huu + Claude multimodal instruction pipeline — quality-fixed re-run. Pipeline Flux.1-schnell (Huu style phrasings + visual style boosts + aspect ratios) → Seed2 → Florence-2 grounding/caption → tracks A–F → USER / ASSISTANT. Tracks A programmatic QA (counts/style/spatial/absence) from detector B LLaVA-style (conversation / detailed description / complex reasoning) C Evol-Instruct (seed atypical Q →… See the full description on the dataset page: https://huggingface.co/datasets/EmpathicRobotics/imagenet-variations-synth-pilot-v3.

sourceHugging Faceapache-2.0updated 1mo agoView on Hugging Face
0likes55downloads
Dataset Card

ImageNet Variations Synth Pilot v3.1 (diverse, 1K)

Huu + Claude multimodal instruction pipeline — quality-fixed re-run.

Pipeline

Flux.1-schnell (Huu style phrasings + visual style boosts + aspect ratios) → Seed2 → Florence-2 grounding/caption → tracks A–F → USER / ASSISTANT.

Tracks

  • A programmatic QA (counts/style/spatial/absence) from detector
  • B LLaVA-style (conversation / detailed description / complex reasoning)
  • C Evol-Instruct (seed atypical Q → 1 evolve op)
  • D mood / people / generation-from-image / math from verified counts only
  • E contrastive pairs (2 styles, multi-<seed2>)
  • F unanswerable / refusal (kept small; no hedge→F dump)

v3.1 fixes

  • Hedge answers cleaned / fall back to A (not F)
  • D1 math numbers only from Florence counts
  • No “detected object” leak in questions
  • Stronger Flux style boosts for webpage / book / poster
  • E4 answers describe style difference (do not paste base_prompt)

Format

USER: <seed2> ... </seed2> {question} ASSISTANT: {answer}

Files

  • data/pilot.jsonl — 1000 instruction records (id, text, image_path, metadata)
  • images/ — generated PNGs (incl. *_b.png for contrastive pairs)