datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Emakhuwa-Portuguese-News-MT
News Parallel Dataset for Emakhuwa of Mozambique
This repository contains releases of parallel data for machine translation in Mozambican languages.
Currently, it supports one language pair, Portuguese-Emakhuwa, Emakhuwa being the widely spoken language in Mozambique.
Dataset Details
Dataset Description
Funded by: This dataset was created with support from Lacuna Fund, the world’s first collaborative effort to provide data scientists, researchers, and… See the full description on the dataset page: https://huggingface.co/datasets/LIACC/Emakhuwa-Portuguese-News-MT.alpaca-gpt4-portugueseThe dataset is used in the research related to MultilingualSIFT.
PortugueseDollyPortugueseDolly é uma tradição do Databricks Dolly 15k para português brasileiro (pt-br) utilizando o nllb 3.3b.
*Somente para demonstração e pesquisa. Proibido para uso comercial.
PortugueseDolly is a translation of the Databricks Dolly 15k into Brazilian Portuguese (pt-br) using GPT3.5 Turbo.
*For demonstration and research purposes only. Commercial use prohibited.
gsm8k_portuguese_traingsm8k-test-portugueseevol-instruct-portugueseThe dataset is used in the research related to MultilingualSIFT.
placeholder_tiebeportuguese-gpt3.5-fine-tuningbeavertails-portuguese
BeaverTails Portuguese
20,000 English prompt/response pairs from PKU-Alignment/BeaverTails translated into Portuguese using Cohere Command A, with multi-label safety categories.
Dataset Description
Each record contains the original English prompt/response, Portuguese translations, a boolean safety label, and a multi-label category dictionary.
Field
Description
_idx
Original dataset index
prompt
English prompt text
response
English response text
prompt_pt… See the full description on the dataset page: https://huggingface.co/datasets/safety-aya/beavertails-portuguese.trustllm_jailbreaktrigger-portuguesefineweb-portuguese-100k
FineWeb2 Portuguese 100k - Safety Classified
A 100,000-sample subset of FineWeb2 Portuguese web text, classified for content safety using Cohere Command A.
Dataset Description
Each record contains the original FineWeb2 text and metadata, plus a classification field with:
Field
Description
safety_rating
"safe" or "unsafe"
category
List of applicable harm categories (null if safe)
reason
Brief explanation of the classification
Safety Taxonomy… See the full description on the dataset page: https://huggingface.co/datasets/safety-aya/fineweb-portuguese-100k.math-portuguesewildjailbreak_harmful-portugueseuner_llm_inst_portuguese
Dataset Card for Universal NER v1 in the Aya format - Portuguese subset
This dataset is a format conversion for the Portuguese data in the original Universal NER v1 into the Aya instruction format and it's released here under the same CC-BY-SA 4.0 license and conditions.
The dataset contains different subsets and their dev/test/train splits, depending on language. For more details, please refer to:
Dataset Details
For the original Universal NER dataset v1 and more details… See the full description on the dataset page: https://huggingface.co/datasets/universalner/uner_llm_inst_portuguese.rebel_portugueseThis is a dataset that was created to re-train REBEL to work better for the Portuguese language.
This dataset was generated using CROCODILE, which was adapted to use a Portuguese specific model (pt_core_news_sm) instead of their default multi-language model (xx_ent_wiki_sm).
The dataset comes with a train, test, dev and train_dev splits. The train_dev split accounts for 80% of the dataset with the remaining 20% being the training data. The train and dev split was generated from the 80%… See the full description on the dataset page: https://huggingface.co/datasets/grsilva/rebel_portuguese.system_chat_portuguesesharegpt-portuguesePortuguese ShareGPT data translated by gpt-3.5-turbo.The dataset is used in the research related to MultilingualSIFT.
call_function_portuguesechat_template = """{% for message in messages %}{% if message['role'] == 'user' %}{{'<|im_start|>user\n' + message['content'] + eos_token + '\n'}}{% elif message['role'] == 'assistant' %}{{'<|im_start|>assistant\n' + message['content'] + eos_token + '\n' }}{% elif message['role'] == 'docs' %}{{'<|im_start|>docs\n' + message['content'] + eos_token + '\n' }}{% elif message['role'] == 'func_response' %}{{'<|im_start|>function_response\n' + message['content'] + eos_token + '\n' }}{% else %}{{… See the full description on the dataset page: https://huggingface.co/datasets/J-LAB/call_function_portuguese.portuguese-general-useNemotron-Safety-Guard-Dataset-v3-portuguese
Nemotron Portuguese Safety (Translated)
Portuguese safety prompts/responses (translated from Spanish), with labels and categories.
Dataset Description
nemotron_pt
Each record includes Portuguese prompt/response text plus safety labels/categories.
Field
Description
id
Example id
prompt
Portuguese prompt text
response
Portuguese response text (may be null)
prompt_label
"safe" or "unsafe"
response_label
"safe" or "unsafe" (may be empty if… See the full description on the dataset page: https://huggingface.co/datasets/safety-aya/Nemotron-Safety-Guard-Dataset-v3-portuguese.harmbench-portuguesewildguardtest-portuguesewildjailbreak_benign-portuguesewmdp-portuguesedo_anything_now-portuguesetoxigen_tiny-portuguesentx_llm_inst_portuguese
Dataset Card for NTX v1 in the Aya format - Portuguese subset
This dataset is a format conversion for the Portuguese data from the original NTX into the Aya instruction format and it's released here under the CC-BY-SA 4.0 license.
Dataset Details
For the original NTX dataset, the conversion to the Aya instructions format, or more details, please refer to the full dataset in instruction form (https://huggingface.co/datasets/tellarin-ai/ntx_llm_instructions) or to the paper… See the full description on the dataset page: https://huggingface.co/datasets/tellarin-ai/ntx_llm_inst_portuguese.openMath-portuguesecall_function_portuguese_nochattemplatexstest-portuguese
