CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01pedrohavay /portuguese-male-voice-A-datasetaudio4 likes5.3k downloads3y agoHugging Face02portuguese-benchmark-datasets /BLUEX BLUEX There is a repository with the minimal code for using this dataset available here. If you use this dataset for research, please cite the paper: @misc{almeida2023bluex, title={BLUEX: A benchmark based on Brazilian Leading Universities Entrance eXams}, author={Thales Sales Almeida and Thiago Laitz and Giovana K. Bonás and Rodrigo Nogueira}, year={2023}, eprint={2307.05410}, archivePrefix={arXiv}, primaryClass={cs.CL} } text1K<n<10K11 likes736 downloads1y agoHugging Face03liaad /math_dataset_portugueseTo run generation code within 'mathematics_dataset\mathematics_dataset': Activate python venv .\.venv\Scripts\activate Requirements defined in requires.txt Run python generate_to_file.py --output_dir ds to generate dataset to directory \ds Had to change enconding when opening files to utf-8 so that some characters are allowed (ã õ é) To obtain dataset with the correct amount of rows: python generate_to_file.py --output_dir ds --per_train_module 1999998 --per_test_module 10000 This… See the full description on the dataset page: https://huggingface.co/datasets/liaad/math_dataset_portuguese.text1K<n<10K0 likes479 downloads3y agoHugging Face04freds0 /cml_tts_dataset_portugueseaudio10K<n<100K3 likes384 downloads2y agoHugging Face05iara-project /raw_dataset_with_embeddings_bert-base-portuguese-cased-nli-assin-2 Dataset Card for "raw_dataset_with_embeddings_bert-base-portuguese-cased-nli-assin-2" More Information needed text100K<n<1M0 likes196 downloads3y agoHugging Face06portuguese-benchmark-datasets /story_cloze_pt Dataset Card for "story_cloze_pt" This is a portuguese translation of the xstory_cloze dataset. The translation was performed using the Google Translate API. This dataset follows the same structure as the original. text1K<n<10K1 likes186 downloads3y agoHugging Face07mazafard /portuguese-ocr-datasettask_categories: image-to-text task_ids: optical-character-recognition text-recognition Portuguese OCR Dataset A comprehensive dataset for Portuguese OCR (Optical Character Recognition) generated from classic Portuguese literature with diverse fonts and visual styles. Dataset Description This dataset contains 20000 text images for OCR training, created from Portuguese books from Project Gutenberg. Each image contains a complete Portuguese sentence with proper… See the full description on the dataset page: https://huggingface.co/datasets/mazafard/portuguese-ocr-dataset.imageimage-to-textn<1K2 likes143 downloads1y agoHugging Face08UniDataPro /portuguese-speech-recognition-dataset Portuguese Speech Dataset for recognition task Dataset comprises 10+ hours of telephone dialogues in Portuguese, collected from 10+ native speakers across various topics and domains. It is a valuable resource for advancing speech recognition technology. By utilizing this dataset, researchers and developers can advance their understanding and capabilities in automatic speech recognition (ASR) systems, transcribing audio, and natural language processing (NLP). - Get the data The… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/portuguese-speech-recognition-dataset.audion<1K1 likes109 downloads1mo agoHugging Face09portuguese-benchmark-datasets /placeholder_tiebetext10K<n<100K0 likes55 downloads1y agoHugging Face10shunyalabs /portuguese-speech-datasetaudio1K<n<10K0 likes55 downloads1y agoHugging Face11joaosanches /brazilian_european_portuguese_datasettext100K<n<1M2 likes48 downloads2y agoHugging Face12lgris /cml_tts_dataset_portuguese-tokenised10K<n<100K0 likes34 downloads1y agoHugging Face13EddyGiusepe /dataset-portuguese-aira-v2-Gemma-formatDataset Aira para o formato do Modelo Gemma Resumo do Dataset Este conjunto de dados contém uma coleção de conversas individuais entre um assistente e um usuário. As conversas foram geradas pelas interações do usuário com modelos já ajustados (ChatGPT, LLama 2, Open-Assistant, etc). O conjunto de dados está disponível em português (tem a versão em Inglês que ainda não tratei). Mas você pode baixar do repositório de Nicholas Kluge Corrêa tanto a versão em Português e a versão em… See the full description on the dataset page: https://huggingface.co/datasets/EddyGiusepe/dataset-portuguese-aira-v2-Gemma-format.textquestion-answering10K<n<100K1 likes33 downloads2y agoHugging Face14portuguese-benchmark-datasets /BLUEX_testtext1K<n<10K0 likes29 downloads1y agoHugging Face15ud-nlp /portuguese-speech-recognition-dataset Portuguese Telephone Dialogues Dataset - 10 Hours Dataset comprises 10 hours of high-quality telephone audio recordings in Portuguese, featuring 20+ native speakers and achieving a 98% Word Accuracy Rate. Designed for advancing speech recognition models and language processing, this extensive speech data corpus covers diverse topics and domains, making it ideal for training robust automatic speech recognition (ASR) systems. - Get the data Dataset characteristics:… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/portuguese-speech-recognition-dataset.audioautomatic-speech-recognitionn<1K0 likes28 downloads10mo agoHugging Face16Speech-data /Portuguese-Speech-Dataset 🎧 Portuguese Speech Dataset The Portuguese Speech Dataset is a large-scale speech audio dataset designed to provide structured and high-quality audio data for modern AI and machine learning systems. It contains 195 hours of recorded speech data distributed across 894 files, available in MP3 and WAV formats, with a total size of 437 MB. This carefully curated audio dataset delivers diverse and representative voice data, with a balanced speaker distribution of 52% female and 48% male… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Portuguese-Speech-Dataset.audioautomatic-speech-recognitionn<1K0 likes28 downloads6mo agoHugging Face17portuguese-benchmark-datasets /placeholder_tiebe2text1K<n<10K0 likes26 downloads1y agoHugging Face18marcosremar2 /orpheus-tts-portuguese-datasettabular100K<n<1M0 likes26 downloads10mo agoHugging Face19Byzz-org /qa-small-dataset-portuguesetext100K<n<1M0 likes25 downloads5d agoHugging Face20dataset-rewriter /awesome-chatgpt-prompts-translated-to-Portuguese-5bbb awesome-chatgpt-prompts-translated-to-Portuguese-5bbb Note: This is an AI-generated dataset so its content may be inaccurate or false Source of the data: The dataset was generated using the Dataset ReWriter and meta-llama/Meta-Llama-3.1-8B-Instruct from the dataset fka/awesome-chatgpt-prompts and using the prompt 'Translate to pt-br': Original Dataset: https://huggingface.co/datasets/fka/awesome-chatgpt-prompts Model: https://huggingface.co/meta-llama/Meta-Llama-3.1-8B-Instruct… See the full description on the dataset page: https://huggingface.co/datasets/dataset-rewriter/awesome-chatgpt-prompts-translated-to-Portuguese-5bbb.textn<1K0 likes21 downloads2y agoHugging Face21portuguese-benchmark-datasets /xpaws_pt Dataset Card for "xpaws_pt" This is a portuguese translation of the x-paws dataset. The translation was performed using the Google Translate API. This dataset follows the same structure as the original. tabular1K<n<10K1 likes20 downloads3y agoHugging Face22FpOliveira /TuPi-Portuguese-Hate-Speech-Dataset This dataset was moved to a new repo ✈️ The data and its repository have relocated to Silly-Machine/TuPyE-Dataset – they needed a change of scenery! Feel free to explore our other organizational projects while you're there. text-classification10K<n<100K3 likes19 downloads3y agoHugging Face23rhaymison /portuguese_query_datasettext100K<n<1M3 likes16 downloads2y agoHugging Face24adalbertojunior /gpt4v-dataset-portugueseimage10K<n<100K1 likes13 downloads2y agoHugging Face25chitradrishti /portuguese-ocr-datasetimage10K<n<100K0 likes12 downloads10mo agoHugging Face26safety-aya /Nemotron-Safety-Guard-Dataset-v3-portuguese Nemotron Portuguese Safety (Translated) Portuguese safety prompts/responses (translated from Spanish), with labels and categories. Dataset Description nemotron_pt Each record includes Portuguese prompt/response text plus safety labels/categories. Field Description id Example id prompt Portuguese prompt text response Portuguese response text (may be null) prompt_label "safe" or "unsafe" response_label "safe" or "unsafe" (may be empty if… See the full description on the dataset page: https://huggingface.co/datasets/safety-aya/Nemotron-Safety-Guard-Dataset-v3-portuguese.tabulartext-classification10K<n<100K0 likes12 downloads6mo agoHugging Face27lgris /cml_tts_dataset_portuguese-multispeaker_tokenised10K<n<100K0 likes11 downloads1y agoHugging Face28portuguese-benchmark-datasets /BLUEX_temp_placeholdertext1K<n<10K0 likes8 downloads1y agoHugging Face29safety-aya /Nemotron-Safety-Guard-Dataset-v3-portuguese-safetytabular10K<n<100K0 likes7 downloads6mo agoHugging Face30lgris /cml_tts_dataset_portuguese-tokenised-dev1K<n<10K0 likes6 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.