datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
portuguese-male-voice-A-datasetBLUEX
BLUEX
There is a repository with the minimal code for using this dataset available here. If you use this dataset for research, please cite the paper:
@misc{almeida2023bluex,
title={BLUEX: A benchmark based on Brazilian Leading Universities Entrance eXams},
author={Thales Sales Almeida and Thiago Laitz and Giovana K. Bonás and Rodrigo Nogueira},
year={2023},
eprint={2307.05410},
archivePrefix={arXiv},
primaryClass={cs.CL}
}
math_dataset_portugueseTo run generation code within 'mathematics_dataset\mathematics_dataset':
Activate python venv .\.venv\Scripts\activate
Requirements defined in requires.txt
Run python generate_to_file.py --output_dir ds to generate dataset to directory \ds
Had to change enconding when opening files to utf-8 so that some characters are allowed (ã õ é)
To obtain dataset with the correct amount of rows:
python generate_to_file.py --output_dir ds --per_train_module 1999998 --per_test_module 10000
This… See the full description on the dataset page: https://huggingface.co/datasets/liaad/math_dataset_portuguese.cml_tts_dataset_portugueseraw_dataset_with_embeddings_bert-base-portuguese-cased-nli-assin-2
Dataset Card for "raw_dataset_with_embeddings_bert-base-portuguese-cased-nli-assin-2"
More Information needed
story_cloze_pt
Dataset Card for "story_cloze_pt"
This is a portuguese translation of the xstory_cloze dataset. The translation was performed using the Google Translate API.
This dataset follows the same structure as the original.
portuguese-ocr-datasettask_categories:
image-to-text
task_ids:
optical-character-recognition
text-recognition
Portuguese OCR Dataset
A comprehensive dataset for Portuguese OCR (Optical Character Recognition) generated from classic Portuguese literature with diverse fonts and visual styles.
Dataset Description
This dataset contains 20000 text images for OCR training, created from Portuguese books from Project Gutenberg. Each image contains a complete Portuguese sentence with proper… See the full description on the dataset page: https://huggingface.co/datasets/mazafard/portuguese-ocr-dataset.portuguese-speech-recognition-dataset
Portuguese Speech Dataset for recognition task
Dataset comprises 10+ hours of telephone dialogues in Portuguese, collected from 10+ native speakers across various topics and domains. It is a valuable resource for advancing speech recognition technology.
By utilizing this dataset, researchers and developers can advance their understanding and capabilities in automatic speech recognition (ASR) systems, transcribing audio, and natural language processing (NLP). - Get the data
The… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/portuguese-speech-recognition-dataset.placeholder_tiebeportuguese-speech-datasetbrazilian_european_portuguese_datasetcml_tts_dataset_portuguese-tokeniseddataset-portuguese-aira-v2-Gemma-formatDataset Aira para o formato do Modelo Gemma
Resumo do Dataset
Este conjunto de dados contém uma coleção de conversas individuais entre um assistente e um usuário.
As conversas foram geradas pelas interações do usuário com modelos já ajustados (ChatGPT, LLama 2, Open-Assistant, etc).
O conjunto de dados está disponível em português (tem a versão em Inglês que ainda não tratei). Mas você pode baixar do
repositório de Nicholas Kluge Corrêa tanto a versão em Português e
a versão em… See the full description on the dataset page: https://huggingface.co/datasets/EddyGiusepe/dataset-portuguese-aira-v2-Gemma-format.BLUEX_testportuguese-speech-recognition-dataset
Portuguese Telephone Dialogues Dataset - 10 Hours
Dataset comprises 10 hours of high-quality telephone audio recordings in Portuguese, featuring 20+ native speakers and achieving a 98% Word Accuracy Rate. Designed for advancing speech recognition models and language processing, this extensive speech data corpus covers diverse topics and domains, making it ideal for training robust automatic speech recognition (ASR) systems. - Get the data
Dataset characteristics:… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/portuguese-speech-recognition-dataset.Portuguese-Speech-Dataset
🎧 Portuguese Speech Dataset
The Portuguese Speech Dataset is a large-scale speech audio dataset designed to provide structured and high-quality audio data for modern AI and machine learning systems. It contains 195 hours of recorded speech data distributed across 894 files, available in MP3 and WAV formats, with a total size of 437 MB. This carefully curated audio dataset delivers diverse and representative voice data, with a balanced speaker distribution of 52% female and 48% male… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Portuguese-Speech-Dataset.placeholder_tiebe2orpheus-tts-portuguese-datasetqa-small-dataset-portugueseawesome-chatgpt-prompts-translated-to-Portuguese-5bbb
awesome-chatgpt-prompts-translated-to-Portuguese-5bbb
Note: This is an AI-generated dataset so its content may be inaccurate or false
Source of the data:
The dataset was generated using the Dataset ReWriter and meta-llama/Meta-Llama-3.1-8B-Instruct from the dataset fka/awesome-chatgpt-prompts and using the prompt 'Translate to pt-br':
Original Dataset: https://huggingface.co/datasets/fka/awesome-chatgpt-prompts
Model: https://huggingface.co/meta-llama/Meta-Llama-3.1-8B-Instruct… See the full description on the dataset page: https://huggingface.co/datasets/dataset-rewriter/awesome-chatgpt-prompts-translated-to-Portuguese-5bbb.xpaws_pt
Dataset Card for "xpaws_pt"
This is a portuguese translation of the x-paws dataset. The translation was performed using the Google Translate API.
This dataset follows the same structure as the original.
TuPi-Portuguese-Hate-Speech-Dataset
This dataset was moved to a new repo ✈️
The data and its repository have relocated to Silly-Machine/TuPyE-Dataset – they needed a change of scenery! Feel free to explore our other organizational projects while you're there.
portuguese_query_datasetgpt4v-dataset-portugueseportuguese-ocr-datasetNemotron-Safety-Guard-Dataset-v3-portuguese
Nemotron Portuguese Safety (Translated)
Portuguese safety prompts/responses (translated from Spanish), with labels and categories.
Dataset Description
nemotron_pt
Each record includes Portuguese prompt/response text plus safety labels/categories.
Field
Description
id
Example id
prompt
Portuguese prompt text
response
Portuguese response text (may be null)
prompt_label
"safe" or "unsafe"
response_label
"safe" or "unsafe" (may be empty if… See the full description on the dataset page: https://huggingface.co/datasets/safety-aya/Nemotron-Safety-Guard-Dataset-v3-portuguese.cml_tts_dataset_portuguese-multispeaker_tokenisedBLUEX_temp_placeholderNemotron-Safety-Guard-Dataset-v3-portuguese-safetycml_tts_dataset_portuguese-tokenised-dev
