CoolFace
Datasetpublic

LingoIITGN/Triveni

📦 Pretraining Corpus 📊 Dataset Overview This dataset combines data from two major sources—Vaani and Flickr30k—to support multilingual and multimodal model pretraining. Source Languages Samples per Language Total Samples Vaani Hindi, English, Hinglish 30,195 90,585 Flickr30k Hindi, English, Hinglish 31,014 93,042 Total — — 183,627 📁 Dataset Sources 🗣️ Vaani Dataset License: CC-BY-4.0 Description: VAANI is an… See the full description on the dataset page: https://huggingface.co/datasets/LingoIITGN/Triveni.

sourceHugging Facemitupdated 1y agoView on Hugging Face
0likes227downloads

LingoIITGN/Triveni · main · files are served by the source, never re-hosted here