CoolFace
Datasetpublic

LocalDoc/AzTC

AzTC (Azerbaijan Text Corpus) This is the first version of the largest text corpus in the Azerbaijani language. Overview The AzTC contains 51 million (approximately 1 billion tokens) non-recurring sentences. The data was collected from various resources such as websites, news, books, wikipedia, legislation, scientific articles and etc. License The AzTC licensed under the CC BY-NC-ND 4.0 license. What does this license allow? Attribution: You must… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/AzTC.

sourceHugging Facecc-by-nc-nd-4.0updated 1y agoView on Hugging Face
3likes129downloads
Dataset Card

AzTC (Azerbaijan Text Corpus)

This is the first version of the largest text corpus in the Azerbaijani language.

Overview

The AzTC contains 51 million (approximately 1 billion tokens) non-recurring sentences. The data was collected from various resources such as websites, news, books, wikipedia, legislation, scientific articles and etc.

License

The AzTC licensed under the CC BY-NC-ND 4.0 license. What does this license allow?

Attribution: You must give appropriate credit, provide a link to the license, and indicate if changes were made. Non-Commercial: You may not use the material for commercial purposes. No Derivatives: If you remix, transform, or build upon the material, you may not distribute the modified material.

For more information, please refer to the <a target="_blank" href="https://creativecommons.org/licenses/by-nc-nd/4.0/">CC BY-NC-ND 4.0 license</a>.

Contact

For more information, questions, or issues, please contact LocalDoc at [v.resad.89@gmail.com].

Citation

@misc{aztc_v1.0,
  title = {Azerbaijan Text Corpus v1.0 (AzTC)},
  author = {LocalDoc},
  year = {2024},
  howpublished = {\url{https://huggingface.co/datasets/LocalDoc/AzTC}},
  note = {Licensed under CC BY-NC-ND 4.0: Attribution-NonCommercial-NoDerivatives 4.0 International},
  url = {https://huggingface.co/datasets/LocalDoc/AzTC}
}