datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
amharic-pretraining-corpusAmharic Pretraining Corpus is a large-scale dataset (~103M) for general amharic language pretraining tasks. It consists of diverse text sources, including news articles, books, social media posts, government documents, and web content, all written in Amharic.
You can load the dataset as follows
from datasets import load_dataset
ds = load_dataset("yordanoswuletaw/amharic-pretraining-corpus")
continued-pretraining-llama-format
Open Paws Continued Pretraining Llama Format
Overview
This dataset is part of the Open Paws initiative to develop AI training data aligned with animal liberation and advocacy principles. Created to train AI systems that understand and promote animal welfare, rights, and liberation.
Dataset Details
Dataset Type: Specialized Data
Format: CSV (Comma-separated values)
Languages: Multilingual (primarily English)
Focus: Animal advocacy and ethical reasoning… See the full description on the dataset page: https://huggingface.co/datasets/open-paws/continued-pretraining-llama-format.zulu-pretraining-datasetThis is IsiZulu Pretraining Dataset. The dataset was used to pre-train BafoGPT-3B
Books: Zulu-English Dictionary – A dictionary offering Zulu terms with English definitions, ideal for teaching basic word mappings.
Translation: South African Government Speeches – Official speeches in Zulu, which help the model understand structured Zulu sentences and phrases.
Transcription: Zulu Community Corpus – A collection of transcriptions, exposing the model to real-life conversational Zulu.
Document:… See the full description on the dataset page: https://huggingface.co/datasets/ChallengerSpaceShuttle/zulu-pretraining-dataset.
