datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
triveni-raw
📦 Pretraining Corpus
📊 Dataset Overview
This dataset combines data from two major sources—Vaani and Flickr30k—to support multilingual and multimodal model pretraining.
Source
Languages
Samples per Language
Total Samples
Vaani
Hindi, English, Hinglish
30,195
90,585
Flickr30k
Hindi, English, Hinglish
31,014
93,042
Total
—
—
183,627
📁 Dataset Sources
🗣️ Vaani Dataset
License: CC-BY-4.0
Description:
VAANI is an… See the full description on the dataset page: https://huggingface.co/datasets/LingoIITGN/triveni-raw.myawady-raw-dataset
Myawady Raw News Corpus 🇲🇲
This dataset contains over 59,000 full-text Burmese news articles scraped from the Myawady News Portal, the official media outlet of the Myanmar military government.
Unlike the title-only version, this dataset includes complete article content, with metadata fields such as category, publication date, and image URLs. It is intended for use in Myanmar NLP and AI research, including:
🧠 Language modeling
📰 Text summarization
🏷️ Named entity… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myawady-raw-dataset.
