CoolFace
Datasetpublic

LocalDoc/AzTC

AzTC (Azerbaijan Text Corpus) This is the first version of the largest text corpus in the Azerbaijani language. Overview The AzTC contains 51 million (approximately 1 billion tokens) non-recurring sentences. The data was collected from various resources such as websites, news, books, wikipedia, legislation, scientific articles and etc. License The AzTC licensed under the CC BY-NC-ND 4.0 license. What does this license allow? Attribution: You must… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/AzTC.

sourceHugging Facecc-by-nc-nd-4.0updated 1y agoView on Hugging Face
3likes129downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face