CoolFace
Datasetpublic

Pikaqiu0114/hico-det-llava-v1.6-13b-answers

HICO-DET LLaVA-1.6-13B action answers Per-image action lists for every HICO-DET image (38,118 train / 9,658 test), produced by LLaVA-1.6 (vicuna-13B) prompted with the 117 HICO-DET verb names and asked to list at most 7 valid actions visible in the picture. These are the text-side VLM answers consumed by UMI-HOI (Unified Multimodal Interaction HOI detection) at training and test time. No images are included; obtain HICO-DET separately and join on image. Files… See the full description on the dataset page: https://huggingface.co/datasets/Pikaqiu0114/hico-det-llava-v1.6-13b-answers.

sourceHugging Faceupdated 18d agoView on Hugging Face
0likes56downloads
Dataset Card

HICO-DET LLaVA-1.6-13B action answers

Per-image action lists for every HICO-DET image (38,118 train / 9,658 test), produced by LLaVA-1.6 (vicuna-13B) prompted with the 117 HICO-DET verb names and asked to list at most 7 valid actions visible in the picture. These are the text-side VLM answers consumed by UMI-HOI (Unified Multimodal Interaction HOI detection) at training and test time.

No images are included; obtain HICO-DET separately and join on image.

Files

FileContent
train.jsonl, test.jsonlone record per image: image, verbs, verb_ids, raw_answer
train.tar.gz, test.tar.gzthe original per-image .txt files in the layout UMI-HOI reads (<split>/<image_stem>.txt)

Each .txt file has three lines:

sit-at sit-on stand-on stand-under talk-on walk watch     # matched verb names, space separated
86 87 93 94 101 110 112                                   # HICO-DET verb indices (0..116)
based on the image provided, here are at most 7 valid actions ...   # raw model answer

Verb indices follow the HICO-DET verb order used by the hicodet loader of UMI-HOI (instances_*.json -> verbs); verb names use hyphens in place of underscores (stand-on, no-interaction).

Notes

  • —Decoding was greedy (temperature ~0). The prompt asked for at most 7 actions; the model returns 6.7 verbs per image on average and tends to fill the list, so recall is high and precision low (on the test split, against GT verbs: recall 52.6, precision 16.4).
  • —The train split covers all 38,118 HICO-DET train images, including the 485 images without any HOI annotation.
  • —The matching CLIP-tower vision features (.pt, 576x1024 per image, ~53 GB) are not part of this upload.
  • —Generated in June 2025 for UMI-HOI. To use with UMI-HOI, extract the tarballs and pass the directory as --llava-answer-path.