Portuguese Data
mbart-large-50-finetuned-opus-en-pt-translation-finetuned-english-to-portuguese-handmade-datasetwav2vec2-large-100k-voxpopuli-ft-TTS-Dataset-plus-data-augmentation-portuguesewav2vec2-large-100k-voxpopuli-ft-Common-Voice_plus_TTS-Dataset-portuguesewav2vec2-large-100k-voxpopuli-ft-Common_Voice_plus_TTS-Dataset_plus_Data_Augmentation-portuguesewav2vec2-large-100k-voxpopuli-ft-TTS-Dataset-portugueseFine_tuned_model_with_portuguese_datasets
portuguese-male-voice-A-datasetBLUEX
BLUEX
There is a repository with the minimal code for using this dataset available here. If you use this dataset for research, please cite the paper:
@misc{almeida2023bluex,
title={BLUEX: A benchmark based on Brazilian Leading Universities Entrance eXams},
author={Thales Sales Almeida and Thiago Laitz and Giovana K. Bonás and Rodrigo Nogueira},
year={2023},
eprint={2307.05410},
archivePrefix={arXiv},
primaryClass={cs.CL}
}
math_dataset_portugueseTo run generation code within 'mathematics_dataset\mathematics_dataset':
Activate python venv .\.venv\Scripts\activate
Requirements defined in requires.txt
Run python generate_to_file.py --output_dir ds to generate dataset to directory \ds
Had to change enconding when opening files to utf-8 so that some characters are allowed (ã õ é)
To obtain dataset with the correct amount of rows:
python generate_to_file.py --output_dir ds --per_train_module 1999998 --per_test_module 10000
This… See the full description on the dataset page: https://huggingface.co/datasets/liaad/math_dataset_portuguese.cml_tts_dataset_portuguesestory_cloze_pt
Dataset Card for "story_cloze_pt"
This is a portuguese translation of the xstory_cloze dataset. The translation was performed using the Google Translate API.
This dataset follows the same structure as the original.
raw_dataset_with_embeddings_bert-base-portuguese-cased-nli-assin-2
Dataset Card for "raw_dataset_with_embeddings_bert-base-portuguese-cased-nli-assin-2"
More Information needed
