LingoIITGN/Triveni
π¦ Pretraining Corpus π Dataset Overview This dataset combines data from two major sourcesβVaani and Flickr30kβto support multilingual and multimodal model pretraining. Source Languages Samples per Language Total Samples Vaani Hindi, English, Hinglish 30,195 90,585 Flickr30k Hindi, English, Hinglish 31,014 93,042 Total β β 183,627 π Dataset Sources π£οΈ Vaani Dataset License: CC-BY-4.0 Description: VAANI is anβ¦ See the full description on the dataset page: https://huggingface.co/datasets/LingoIITGN/Triveni.
π¦ Pretraining Corpus
π Dataset Overview
This dataset combines data from two major sourcesβVaani and Flickr30kβto support multilingual and multimodal model pretraining.
π Dataset Sources
π£οΈ Vaani Dataset
- License: CC-BY-4.0
- Description:
- VAANI is an India-representative multimodal, multilingual dataset.
- Phase 1 includes:
- \~16,000 hours of spontaneous, image-prompted speech.
- 9.6 million utterances from 84.6K speakers across 80 districts.
- 788.03 hours of transcribed text data, covering 54 Indian languages.
- Goal: Build inclusive language AI systems reflecting the full linguistic and cultural diversity of India.
- Initiated by IISc Bangalore and ARTPARK, under the National Language Translation Mission.
- Also available via Bhashini (MeitY).
πΌοΈ Flickr30k Dataset
- License: CC0: Public Domain
- Description:
- A standard benchmark for sentence-based image description.
- Includes 158,000 captions and 244,000 coreference chains.
- Annotated with 276,000 bounding boxes to enable grounded language understanding.
- Useful for tasks like image-sentence retrieval, textual grounding, and object localization.
- Flickr30k Entities enhances it with structured coreference and region-based annotations. ---
π‘ Use Case Highlights
- Pretraining for multilingual/multimodal models.
- Research in grounded image captioning, image-text retrieval, and speech-to-text systems.
- Study of language and regional diversity in Indian contexts.
π¦ Fine-Tuning Corpus
The Indic Multimodal Fine-Tuning Dataset is a multilingual, annotated dataset developed to fine-tune multimodal models in the Indian context. It supports captions in English, Hindi, and Hinglish, across diverse image categories sourced from Indian cities, culture, fashion, food, and more.
This dataset is specifically designed for fine-tuning LingoIITGN / Indic_multimodal_preTrain and similar multimodal models for tasks such as grounded image captioning and image-text retrieval.
π¦ Dataset Summary
- Languages: English, Hindi, Hinglish
- Use Cases: Grounded image captioning, image-text retrieval, regional language modeling
- Total Images: 11,406
- Annotations: Manual captions in 3 languages
- Curated by: LINGO Research Group, IIT Gandhinagar
π Dataset Structure
Each data point includes:
file_name: Name of the image fileimage: Raw image fileenglish_caption: Caption in Englishhindi_caption: Caption in Hindihinglish_caption: Caption in Hinglishoriginal_url: URL source of the image (where available)license: License type for the imagedataset_name: Subset name or category
π Image Categories
π Use Cases
- πΌοΈ Image Captioning: Multilingual caption generation for culturally rich images
- π Image-Text Retrieval: Multimodal alignment tasks using diverse Indian contexts
- π€ Speech-to-Text Grounding: Training models with visual context in Indian languages
- π Cross-lingual Modeling: Language transfer and Hinglish understanding
β οΈ License Notes
Each category may have different license terms. Users are expected to:
- Respect the individual license provided in the dataset.
- Avoid commercial use for images labeled with "N/A" or "Data files Β© Original Authors" without verifying rights.
- Attribute appropriately if reusing or modifying the dataset.
βοΈ Annotation and Curation
The dataset was annotated and verified by the LINGO Research Group at IIT Gandhinagar. Captions were manually written to reflect cultural and contextual accuracy in three languages.
π¬ Contact
For questions, suggestions, or contributions, reach out to the LINGO Group, IIT Gandhinagar.
π‘ Citation
If you use this dataset in your work, please cite the dataset and mention the LINGO Research Group, IIT Gandhinagar.
Made with β€οΈ for Indic multimodal research.
