CoolFace
Datasetpublicgated

ai4bharat/IndicCMix

IndicCMix Most Indic NLP data assumes people write in one script and one language at a time. Real chat looks nothing like that. You get Hindi words in Roman letters, English verbs in the middle of a Tamil sentence, and the same person switching scripts halfway through a paragraph. This dataset is an attempt to cover that actual messiness. For every English sentence, you get three different Indic renderings of it: one code-mixed in the native script, one clean native-script… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/IndicCMix.

sourceHugging Facemitupdated 5mo agoView on Hugging Face
1likes306downloads
../
filetrain-00000-of-00001.parquet36.6 MBdownload

ai4bharat/IndicCMix · main · files are served by the source, never re-hosted here

This repository is gated. The listing is public, but downloading a file means accepting the publisher’s terms at Hugging Face first — the links above take you there rather than around it.

ai4bharat/IndicCMix · CoolFace