LocalDoc/AzTC
AzTC (Azerbaijan Text Corpus) This is the first version of the largest text corpus in the Azerbaijani language. Overview The AzTC contains 51 million (approximately 1 billion tokens) non-recurring sentences. The data was collected from various resources such as websites, news, books, wikipedia, legislation, scientific articles and etc. License The AzTC licensed under the CC BY-NC-ND 4.0 license. What does this license allow? Attribution: You must… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/AzTC.
3129
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Upload dataset
initial commit
