datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
naijaweb
Naijaweb Dataset 🇳🇬
Naijaweb is a dataset that contains over 270,000+ documents, totaling approximately 230 million GPT-2 tokens. The data was web scraped from web pages popular among Nigerians, providing a rich resource for modeling Nigerian linguistic and cultural contexts.
Dataset Summary
Features
Data Types
text
string
link
string
token_count
int64
section
string
int_score
int64
language
string
language_probability
float64… See the full description on the dataset page: https://huggingface.co/datasets/saheedniyi/naijaweb.naijaweb-edu
Naijaweb Edu Dataset 🇳🇬
Naijaweb Edu is a subset of the naijaweb dataset with an educational score aboove 3 using the fineweb classifier. The initial fineweb dataset was web scraped from web pages popular among Nigerians, providing a rich resource for modeling Nigerian linguistic and cultural contexts.
Dataset Summary
Features
Data Types
text
string
link
string
token_count
int64
section
string
int_score
int64
language
string… See the full description on the dataset page: https://huggingface.co/datasets/saheedniyi/naijaweb-edu.naijaweb-edu2
Naijaweb Edu2 Dataset 🇳🇬
Naijaweb Edu 2 is a subset of the naijaweb dataset with an educational score aboove 2 using the fineweb classifier. The initial fineweb dataset was web scraped from web pages popular among Nigerians, providing a rich resource for modeling Nigerian linguistic and cultural contexts.
Dataset Summary
Features
Data Types
text
string
link
string
token_count
int64
section
string
int_score
int64
language
string… See the full description on the dataset page: https://huggingface.co/datasets/saheedniyi/naijaweb-edu2.
