TatarNLPWorld/tatar-web-corpus
Dataset Card for Tatar Web Corpus Dataset Details Dataset Description The largest open corpus for the Tatar language with over 1 million documents collected from news websites, social media, articles, books, and Wikipedia. Designed for various NLP tasks including language modeling, text classification, information extraction, and search. Curated by: TatarNLPWorld Community Language(s) (NLP): Tatar (tt) License: other – see Licensing & Legal Notice… See the full description on the dataset page: https://huggingface.co/datasets/TatarNLPWorld/tatar-web-corpus.
033
No card is published for this repository, or it could not be fetched from Hugging Face right now.
