CoolFace
Datasetpublic

LingoIITGN/Triveni

πŸ“¦ Pretraining Corpus πŸ“Š Dataset Overview This dataset combines data from two major sourcesβ€”Vaani and Flickr30kβ€”to support multilingual and multimodal model pretraining. Source Languages Samples per Language Total Samples Vaani Hindi, English, Hinglish 30,195 90,585 Flickr30k Hindi, English, Hinglish 31,014 93,042 Total β€” β€” 183,627 πŸ“ Dataset Sources πŸ—£οΈ Vaani Dataset License: CC-BY-4.0 Description: VAANI is an… See the full description on the dataset page: https://huggingface.co/datasets/LingoIITGN/Triveni.

sourceHugging Facemitupdated 1y agoView on Hugging Face
0likes224downloads
Dataset Card

πŸ“¦ Pretraining Corpus

πŸ“Š Dataset Overview

This dataset combines data from two major sourcesβ€”Vaani and Flickr30kβ€”to support multilingual and multimodal model pretraining.

SourceLanguagesSamples per LanguageTotal Samples
VaaniHindi, English, Hinglish30,19590,585
Flickr30kHindi, English, Hinglish31,01493,042
Totalβ€”β€”183,627

πŸ“ Dataset Sources

πŸ—£οΈ Vaani Dataset

  • β€”VAANI is an India-representative multimodal, multilingual dataset.
  • β€”Phase 1 includes:
  • β€”\~16,000 hours of spontaneous, image-prompted speech.
  • β€”9.6 million utterances from 84.6K speakers across 80 districts.
  • β€”788.03 hours of transcribed text data, covering 54 Indian languages.
  • β€”Goal: Build inclusive language AI systems reflecting the full linguistic and cultural diversity of India.
  • β€”Initiated by IISc Bangalore and ARTPARK, under the National Language Translation Mission.
  • β€”Also available via Bhashini (MeitY).

πŸ–ΌοΈ Flickr30k Dataset

  • β€”A standard benchmark for sentence-based image description.
  • β€”Includes 158,000 captions and 244,000 coreference chains.
  • β€”Annotated with 276,000 bounding boxes to enable grounded language understanding.
  • β€”Useful for tasks like image-sentence retrieval, textual grounding, and object localization.
  • β€”Flickr30k Entities enhances it with structured coreference and region-based annotations. ---

πŸ’‘ Use Case Highlights

  • β€”Pretraining for multilingual/multimodal models.
  • β€”Research in grounded image captioning, image-text retrieval, and speech-to-text systems.
  • β€”Study of language and regional diversity in Indian contexts.

πŸ“¦ Fine-Tuning Corpus

The Indic Multimodal Fine-Tuning Dataset is a multilingual, annotated dataset developed to fine-tune multimodal models in the Indian context. It supports captions in English, Hindi, and Hinglish, across diverse image categories sourced from Indian cities, culture, fashion, food, and more.

This dataset is specifically designed for fine-tuning LingoIITGN / Indic_multimodal_preTrain and similar multimodal models for tasks such as grounded image captioning and image-text retrieval.


πŸ“¦ Dataset Summary

  • β€”Languages: English, Hindi, Hinglish
  • β€”Use Cases: Grounded image captioning, image-text retrieval, regional language modeling
  • β€”Total Images: 11,406
  • β€”Annotations: Manual captions in 3 languages
  • β€”Curated by: LINGO Research Group, IIT Gandhinagar

πŸ“ Dataset Structure

Each data point includes:

  • β€”file_name: Name of the image file
  • β€”image: Raw image file
  • β€”english_caption: Caption in English
  • β€”hindi_caption: Caption in Hindi
  • β€”hinglish_caption: Caption in Hinglish
  • β€”original_url: URL source of the image (where available)
  • β€”license: License type for the image
  • β€”dataset_name: Subset name or category

πŸ“Š Image Categories

CategoryImage CountLicense
Indian city images7,500Apache License 2.0
Indian currency images100CC0: Public Domain
Indian dance images156N/A
Indian fashion images3,500N/A
Indian instrument images100Data files Β© Original Authors
Indian vehicle images500Data files Β© Original Authors
Indian food images500CC0: Public Domain
Indian temple images50N/A
Indian musical instrument images1661N/A

πŸ” Use Cases

  • β€”πŸ–ΌοΈ Image Captioning: Multilingual caption generation for culturally rich images
  • β€”πŸ” Image-Text Retrieval: Multimodal alignment tasks using diverse Indian contexts
  • β€”πŸŽ€ Speech-to-Text Grounding: Training models with visual context in Indian languages
  • β€”πŸŒ Cross-lingual Modeling: Language transfer and Hinglish understanding

⚠️ License Notes

Each category may have different license terms. Users are expected to:

  • β€”Respect the individual license provided in the dataset.
  • β€”Avoid commercial use for images labeled with "N/A" or "Data files Β© Original Authors" without verifying rights.
  • β€”Attribute appropriately if reusing or modifying the dataset.

✍️ Annotation and Curation

The dataset was annotated and verified by the LINGO Research Group at IIT Gandhinagar. Captions were manually written to reflect cultural and contextual accuracy in three languages.


πŸ“¬ Contact

For questions, suggestions, or contributions, reach out to the LINGO Group, IIT Gandhinagar.


πŸ’‘ Citation

If you use this dataset in your work, please cite the dataset and mention the LINGO Research Group, IIT Gandhinagar.


Made with ❀️ for Indic multimodal research.