CoolFace
Datasetpublicgated

ai4bharat/IndicCMix

IndicCMix Most Indic NLP data assumes people write in one script and one language at a time. Real chat looks nothing like that. You get Hindi words in Roman letters, English verbs in the middle of a Tamil sentence, and the same person switching scripts halfway through a paragraph. This dataset is an attempt to cover that actual messiness. For every English sentence, you get three different Indic renderings of it: one code-mixed in the native script, one clean native-script… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/IndicCMix.

sourceHugging Facemitupdated 5mo agoView on Hugging Face
1likes306downloads
fileas.parquet41.2 MBdownload
filebn.parquet39.7 MBdownload
filegu.parquet39.8 MBdownload
filehi.parquet39.5 MBdownload
fileka.parquet42.4 MBdownload
fileml.parquet43.2 MBdownload
filemr.parquet40.0 MBdownload
fileor.parquet40.7 MBdownload
filepa.parquet40.2 MBdownload
fileta.parquet42.5 MBdownload
filete.parquet40.9 MBdownload

ai4bharat/IndicCMix · main · files are served by the source, never re-hosted here

This repository is gated. The listing is public, but downloading a file means accepting the publisher’s terms at Hugging Face first — the links above take you there rather than around it.

ai4bharat/IndicCMix · CoolFace