datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nso-gov-vn
nso-gov-vn — Vietnam NSO PX-Web mirror · Bản sao PX-Web của Tổng cục Thống kê
🇻🇳 Tóm tắt. Bản sao đầy đủ, từng bảng một, của cơ sở dữ liệu thống
kê PX-Web của Tổng cục Thống kê Việt Nam (NSO / GSO) tại
https://pxweb.nso.gov.vn. Mỗi ma trận PX-Web (multi-dimensional
data cube) được phơi ra cùng lúc ở (i) bản gốc với schema riêng và
(ii) định dạng long-format gộp chung, để bạn có thể chọn giữa
"nguyên trạng" hay "join sẵn".
🇬🇧 Summary. A complete, table-by-table mirror of the… See the full description on the dataset page: https://huggingface.co/datasets/tmquan/nso-gov-vn.mmlu-pro-augmentationrag-embeddings-and-textnchlt_speech_nso
NCHLT Speech Corpus -- Sepedi
This is the Sepedi language part of the NCHLT Speech Corpus of the South African languages.
Language code (ISO 639): nso
URI: https://hdl.handle.net/20.500.12185/270
Licence:
Creative Commons Attribution 3.0 Unported License (CC BY 3.0): http://creativecommons.org/licenses/by/3.0/legalcode
Attribution:
The Department of Arts and Culture of the government of the Republic of South Africa (DAC), Council for Scientific and Industrial… See the full description on the dataset page: https://huggingface.co/datasets/danielshaps/nchlt_speech_nso.m3-scientific-sft-mcqahf-dpo-pairs-collectiondpo-student-preference-pairswikiscrape1000-oaallenai-math_qa-1000yahma-alpaca-cleaned-nso
Terms of Use
This model is governed by a Apache 2.0 License.
allenai-qasc-1000commonsenseqa-1000MMLU-STEM-1000NSOARTC_Tanshi_TEST_0107wikiscrape1000-mcqopenai-subset-gsm1kopenlifescienceai-medmcqa-1000NSOARTC_Tanshi_1221babylm-nso
babylm-nso
Dataset Description
This dataset is part of the BabyLM multilingual collection.
Dataset Summary
Language: nso
Script: Latin
Number of Documents: 26772
Total Tokens: 1067761
Tokens Per Category
child-books: 122083 tokens
child-news: 130 tokens
educational: 92589 tokens
padding-mt: 206703 tokens
padding-news: 150960 tokens
padding-wikipedia: 495296 tokens
Data Fields
text: The document text
category: Type of content (e.g.… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/babylm-nso.rag-text-chunkscv24-nso-128-normalizedNSC_sample_0dpo-student-preferences
