CoolFace
Datasetpublic

LocalDoc/AzTC

AzTC (Azerbaijan Text Corpus) This is the first version of the largest text corpus in the Azerbaijani language. Overview The AzTC contains 51 million (approximately 1 billion tokens) non-recurring sentences. The data was collected from various resources such as websites, news, books, wikipedia, legislation, scientific articles and etc. License The AzTC licensed under the CC BY-NC-ND 4.0 license. What does this license allow? Attribution: You must… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/AzTC.

sourceHugging Facecc-by-nc-nd-4.0updated 1y agoView on Hugging Face
3likes129downloads
7 commits on main
303b5be1y ago

Update README.md

vrashad
f275a6b2y ago

Update README.md

vrashad
b6c6cd92y ago

Update README.md

vrashad
32f85cc2y ago

Update README.md

vrashad
a8d1a782y ago

Update README.md

vrashad
a695d982y ago

Upload dataset

vrashad
245cb912y ago

initial commit

vrashad