CoolFace
20 results

pt_br

bastao /VeraCruz_PT-BR Dataset Summary The VeraCruz Dataset is a comprehensive collection of Portuguese language content, showcasing the linguistic and cultural diversity of of Portuguese-speaking regions. It includes around 190 million samples, organized by regional origin as indicated by URL metadata into primary categories. The primary categories are: Portugal (PT): Samples with content URLs indicating a clear Portuguese origin. Brazil (BR): Samples with content URLs indicating a clear Brazilian… See the full description on the dataset page: https://huggingface.co/datasets/bastao/VeraCruz_PT-BR.texttext-generation100M<n<1B17 likes53k downloads1y agoHugging Faceopedromartins /ASR-datasets-ptbr 📚 Datasets de Áudio em Português (PT-BR) Este repositório reúne diversos corpora públicos de fala em português do Brasil, combinados em um único dataset para facilitar treinamentos e pesquisas em ASR (Automatic Speech Recognition). O objetivo é fornecer um recurso amplo, padronizado e de fácil acesso para a comunidade. 📂 Datasets Integrados A tabela abaixo lista todos os datasets incluídos, com suas informações: Dataset Config Name TOTAL train test validation… See the full description on the dataset page: https://huggingface.co/datasets/opedromartins/ASR-datasets-ptbr.audioautomatic-speech-recognition1M<n<10M13 likes2.5k downloads1y agoHugging FaceMTEB-BR /mteb-pt-results 🇧🇷 MTEB-BR — Benchmark Results Canonical results store for MTEB-BR, a native Brazilian-Portuguese text-embedding benchmark. 93 models · 22 native PT-BR tasks · 7 categories · no machine translation What is this? This repository is the canonical, machine-readable results store for MTEB-BR — a benchmark that evaluates text-embedding models on native Brazilian Portuguese (data created or found in Portuguese; machine-translated corpora such as… See the full description on the dataset page: https://huggingface.co/datasets/MTEB-BR/mteb-pt-results.tabularfeature-extractionn<1K0 likes1.2k downloads2mo agoHugging Facedominguesm /mTEDx-ptbr Multilingual TEDx (Portuguese speech and transcripts) NOTE: This dataset contains only the Portuguese portion of the mTEDx dataset, already processed and segmented into parts. Multilingual TEDx (mTEDx) is a multilingual speech recognition and translation corpus to facilitate the training of ASR and SLT models in additional languages. The corpus comprises audio recordings and transcripts from TEDx Talks in 8 languages (Spanish, French, Portuguese, Italian, Russian, Greek, Arabic… See the full description on the dataset page: https://huggingface.co/datasets/dominguesm/mTEDx-ptbr.audioautomatic-speech-recognition1K<n<10K10 likes1k downloads3y agoHugging Facedominguesm /alpaca-data-pt-brNOTE: This is a machine translated version of the yahma/alpaca-cleaned dataset. Dataset Card for Alpaca-Cleaned Repository: https://github.com/gururise/AlpacaDataCleaned Dataset Description This is a cleaned version of the original Alpaca Dataset released by Stanford. The following issues have been identified in the original release and fixed in this dataset: Hallucinations: Many instructions in the original dataset had instructions referencing data on the internet… See the full description on the dataset page: https://huggingface.co/datasets/dominguesm/alpaca-data-pt-br.texttext-generation10K<n<100K35 likes595 downloads3y agoHugging FacePatoFlamejanteTV /geral-edu-ptbrdocument1K<n<10K0 likes527 downloads7mo agoHugging Face