datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Emakhuwa-Portuguese-News-MT
News Parallel Dataset for Emakhuwa of Mozambique
This repository contains releases of parallel data for machine translation in Mozambican languages.
Currently, it supports one language pair, Portuguese-Emakhuwa, Emakhuwa being the widely spoken language in Mozambique.
Dataset Details
Dataset Description
Funded by: This dataset was created with support from Lacuna Fund, the world’s first collaborative effort to provide data scientists, researchers, and… See the full description on the dataset page: https://huggingface.co/datasets/LIACC/Emakhuwa-Portuguese-News-MT.fineweb-portuguese-100k
FineWeb2 Portuguese 100k - Safety Classified
A 100,000-sample subset of FineWeb2 Portuguese web text, classified for content safety using Cohere Command A.
Dataset Description
Each record contains the original FineWeb2 text and metadata, plus a classification field with:
Field
Description
safety_rating
"safe" or "unsafe"
category
List of applicable harm categories (null if safe)
reason
Brief explanation of the classification
Safety Taxonomy… See the full description on the dataset page: https://huggingface.co/datasets/safety-aya/fineweb-portuguese-100k.Nemotron-Safety-Guard-Dataset-v3-portuguese
Nemotron Portuguese Safety (Translated)
Portuguese safety prompts/responses (translated from Spanish), with labels and categories.
Dataset Description
nemotron_pt
Each record includes Portuguese prompt/response text plus safety labels/categories.
Field
Description
id
Example id
prompt
Portuguese prompt text
response
Portuguese response text (may be null)
prompt_label
"safe" or "unsafe"
response_label
"safe" or "unsafe" (may be empty if… See the full description on the dataset page: https://huggingface.co/datasets/safety-aya/Nemotron-Safety-Guard-Dataset-v3-portuguese.wildguardtest-portuguese
