CoolFace
Datasetpublic

iLearn-Lab/MM26-ADGNet-AITIR-Text

πŸ“ AITIR: Asymmetric Image-Text Infrared Text Annotations Official asymmetric text annotations for infrared small target detection. " rel="nofollow"> πŸ“– Description This repository provides the official asymmetric text annotations used to construct the Asymmetric Image-Text Infrared (AITIR) dataset introduced in: ADGNet: Asymmetric Dual-text Guided Network for Infrared Small Target Detection, accepted by ACM Multimedia 2026. The released… See the full description on the dataset page: https://huggingface.co/datasets/iLearn-Lab/MM26-ADGNet-AITIR-Text.

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
0likes54downloads
Dataset Card

<h1>πŸ“ AITIR: Asymmetric Image-Text Infrared Text Annotations</h1>

<p> Official asymmetric text annotations for infrared small target detection. </p>

<p> <a href="https://github.com/iLearn-Lab/MM26-ADGNet"> <img src="https://img.shields.io/badge/GitHub-MM26--ADGNet-black?logo=github" alt="GitHub"> </a> <a href="<paper-link>"> <img src="https://img.shields.io/badge/ACM%20MM-2026-blue" alt="ACM MM 2026"> </a> </p>


πŸ“– Description

This repository provides the official asymmetric text annotations used to construct the Asymmetric Image-Text Infrared (AITIR) dataset introduced in:

ADGNet: Asymmetric Dual-text Guided Network for Infrared Small Target Detection, accepted by ACM Multimedia 2026.

The released annotations cover three infrared small target detection datasets:

  • β€”IRSTD-1K
  • β€”NUDT-SIRST
  • β€”SIRST

This repository contains text annotations only. The original infrared images and ground-truth masks are not included.


🏷️ Annotation Design

Each infrared image is associated with two asymmetric text prompts.

Fixed Target Prompt

A concise and image-independent target prompt is shared across all images:

text
an infrared image featuring one or multiple target

Detailed Background Prompt

Each image is assigned an image-dependent background prompt following the template:

text
an infrared [S] image with [C]

where:

  • β€”[S] describes the global infrared scene.
  • β€”[C] describes local structures, thermal clutter, and potential distractors.

Example:

text
an infrared cloudy sky image with faint dark cloud structures and bright illuminated building tops

πŸ“Š Annotation Coverage

DatasetFixed Target PromptDetailed Background Prompt
IRSTD-1Kβœ“βœ“
NUDT-SIRSTβœ“βœ“
SIRSTβœ“βœ“

πŸ“ Repository Structure

The asymmetric text annotations are organized by dataset and data split:

text
MM26-ADGNet-AITIR-Text/
β”œβ”€β”€ IRSTD-1K/
β”‚   └── text/
β”‚       β”œβ”€β”€ train_fg.json
β”‚       β”œβ”€β”€ train_bg.json
β”‚       β”œβ”€β”€ test_fg.json
β”‚       └── test_bg.json
β”œβ”€β”€ NUDT-SIRST/
β”‚   └── text/
β”‚       β”œβ”€β”€ train_fg.json
β”‚       β”œβ”€β”€ train_bg.json
β”‚       β”œβ”€β”€ test_fg.json
β”‚       └── test_bg.json
└── SIRST/
    └── text/
        β”œβ”€β”€ train_fg.json
        β”œβ”€β”€ train_bg.json
        β”œβ”€β”€ test_fg.json
        └── test_bg.json

🧾 Annotation Files

Each dataset contains four JSON annotation files:

FileDescription
train_fg.jsonTarget prompts for the training split
train_bg.jsonDetailed background prompts for the training split
test_fg.jsonTarget prompts for the testing split
test_bg.jsonDetailed background prompts for the testing split

Here, fg denotes the foreground or target prompt, while bg denotes the detailed background prompt.


πŸš€ Usage

Download the annotations and place the corresponding JSON files under the text/ directory of each original dataset:

text
datasets/
β”œβ”€β”€ IRSTD-1K/
β”‚   β”œβ”€β”€ images/
β”‚   β”œβ”€β”€ masks/
β”‚   β”œβ”€β”€ img_idx/
β”‚   └── text/
β”‚       β”œβ”€β”€ train_fg.json
β”‚       β”œβ”€β”€ train_bg.json
β”‚       β”œβ”€β”€ test_fg.json
β”‚       └── test_bg.json
β”œβ”€β”€ NUDT-SIRST/
β”‚   β”œβ”€β”€ images/
β”‚   β”œβ”€β”€ masks/
β”‚   β”œβ”€β”€ img_idx/
β”‚   └── text/
β”‚       β”œβ”€β”€ train_fg.json
β”‚       β”œβ”€β”€ train_bg.json
β”‚       β”œβ”€β”€ test_fg.json
β”‚       └── test_bg.json
└── SIRST/
    β”œβ”€β”€ images/
    β”œβ”€β”€ masks/
    β”œβ”€β”€ img_idx/
    └── text/
        β”œβ”€β”€ train_fg.json
        β”œβ”€β”€ train_bg.json
        β”œβ”€β”€ test_fg.json
        └── test_bg.json
[!NOTE] The JSON entries correspond to the samples in the respective training or testing split. Please keep the provided filenames and directory structure unchanged when using them with the official ADGNet implementation.

⚠️ Notes

  • β€”This repository provides text annotations only.
  • β€”The original infrared images and ground-truth masks must be obtained separately.
  • β€”The annotations are written in English.
  • β€”Users must comply with the licenses and terms of use of the original datasets.
  • β€”Annotation filenames must match the corresponding infrared image filenames.

πŸ“š Citation

If you find this project useful in your research, please consider citing our paper:

bibtex

Please also consider checking out and citing our other related work:

bibtex