CoolFace
Datasetpublic

zicsx/mC4-hindi

Dataset Card for "mC4-hindi" This dataset is a subset of the mC4 dataset, which is a multilingual colossal, cleaned version of Common Crawl's web crawl corpus. It contains natural text in 101 languages, including Hindi. This dataset is specifically focused on Hindi text, and contains a variety of different types of text, including news articles, blog posts, and social media posts. This dataset is intended to be used for training and evaluating natural language processing models… See the full description on the dataset page: https://huggingface.co/datasets/zicsx/mC4-hindi.

sourceHugging Faceapache-2.0updated 3y agoView on Hugging Face
0likes342downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
zicsx/mC4-hindi · CoolFace