CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01bastao /VeraCruz_PT-BR Dataset Summary The VeraCruz Dataset is a comprehensive collection of Portuguese language content, showcasing the linguistic and cultural diversity of of Portuguese-speaking regions. It includes around 190 million samples, organized by regional origin as indicated by URL metadata into primary categories. The primary categories are: Portugal (PT): Samples with content URLs indicating a clear Portuguese origin. Brazil (BR): Samples with content URLs indicating a clear Brazilian… See the full description on the dataset page: https://huggingface.co/datasets/bastao/VeraCruz_PT-BR.texttext-generation100M<n<1B17 likes60k downloads1y agoHugging Face02opedromartins /ASR-datasets-ptbr 📚 Datasets de Áudio em Português (PT-BR) Este repositório reúne diversos corpora públicos de fala em português do Brasil, combinados em um único dataset para facilitar treinamentos e pesquisas em ASR (Automatic Speech Recognition). O objetivo é fornecer um recurso amplo, padronizado e de fácil acesso para a comunidade. 📂 Datasets Integrados A tabela abaixo lista todos os datasets incluídos, com suas informações: Dataset Config Name TOTAL train test validation… See the full description on the dataset page: https://huggingface.co/datasets/opedromartins/ASR-datasets-ptbr.audioautomatic-speech-recognition1M<n<10M13 likes2.5k downloads1y agoHugging Face03MTEB-BR /mteb-pt-results 🇧🇷 MTEB-BR — Benchmark Results Canonical results store for MTEB-BR, a native Brazilian-Portuguese text-embedding benchmark. 93 models · 22 native PT-BR tasks · 7 categories · no machine translation What is this? This repository is the canonical, machine-readable results store for MTEB-BR — a benchmark that evaluates text-embedding models on native Brazilian Portuguese (data created or found in Portuguese; machine-translated corpora such as… See the full description on the dataset page: https://huggingface.co/datasets/MTEB-BR/mteb-pt-results.tabularfeature-extractionn<1K0 likes1.1k downloads2mo agoHugging Face04dominguesm /mTEDx-ptbr Multilingual TEDx (Portuguese speech and transcripts) NOTE: This dataset contains only the Portuguese portion of the mTEDx dataset, already processed and segmented into parts. Multilingual TEDx (mTEDx) is a multilingual speech recognition and translation corpus to facilitate the training of ASR and SLT models in additional languages. The corpus comprises audio recordings and transcripts from TEDx Talks in 8 languages (Spanish, French, Portuguese, Italian, Russian, Greek, Arabic… See the full description on the dataset page: https://huggingface.co/datasets/dominguesm/mTEDx-ptbr.audioautomatic-speech-recognition1K<n<10K10 likes1k downloads3y agoHugging Face05dominguesm /alpaca-data-pt-brNOTE: This is a machine translated version of the yahma/alpaca-cleaned dataset. Dataset Card for Alpaca-Cleaned Repository: https://github.com/gururise/AlpacaDataCleaned Dataset Description This is a cleaned version of the original Alpaca Dataset released by Stanford. The following issues have been identified in the original release and fixed in this dataset: Hallucinations: Many instructions in the original dataset had instructions referencing data on the internet… See the full description on the dataset page: https://huggingface.co/datasets/dominguesm/alpaca-data-pt-br.texttext-generation10K<n<100K35 likes593 downloads3y agoHugging Face06laicsiifes /coco-captions-pt-br 🎉 COCO Captions Dataset Translation for Portuguese Image Captioning 💾 Dataset Summary COCO Captions Portuguese Translation, a multimodal dataset for Portuguese image captioning with 123,287 images, each accompanied by five descriptive captions that have been generated by human annotators for every individual image. The original English captions were rendered into Portuguese through the utilization of the Google Translator API. 🧑‍💻 Hot to Get… See the full description on the dataset page: https://huggingface.co/datasets/laicsiifes/coco-captions-pt-br.imagetext-to-image100K<n<1M6 likes573 downloads4mo agoHugging Face07histlearn /notas-comunidade-ptbr Community Notes BR: Enriched Portuguese Dataset for NLP and Text Mining Community Notes BR is a curated Portuguese-language subset of X Community Notes enriched for Natural Language Processing (PLN), text mining, and research on collaborative misinformation moderation. It combines note-level metadata, topic and macrotheme labels, named entities, source-domain extraction, Matrix Factorization scoring fields, an optional Universal Dependencies syntax layer, and automatic… See the full description on the dataset page: https://huggingface.co/datasets/histlearn/notas-comunidade-ptbr.tabulartext-classification1M<n<10M0 likes553 downloads24d agoHugging Face08dominguesm /Canarim-Instruct-PTBR-Dataset 🐥 🇧🇷 Canarim Instruct Dataset [🐱 Github] What's Canarim? Canarim is a dataset with over 300,000 instructions in Portuguese, ranging from simple instructions like "Descreva os efeitos do aquecimento global" to more complex instructions like "Nesta tarefa, você precisa ser capaz de resumir uma determinada lista de pontos-chave" where additional context is provided. Why it's called Canarim? "Canarim" is spoken in some regions of Brazil (mainly by… See the full description on the dataset page: https://huggingface.co/datasets/dominguesm/Canarim-Instruct-PTBR-Dataset.text100K<n<1M46 likes510 downloads3y agoHugging Face09dominguesm /wikipedia-ptbr-20230601 Dataset Card for "wikipedia-ptbr-20230601" More Information needed text1M<n<10M11 likes457 downloads3y agoHugging Face10Fazzioni /recycling_the_web-pt-brAproximadamente 25GB de arquivos .parquet traduzidos com o openai/gpt-oss-120b a partir do dataset: #Facebook/recycling_the_web prompt utilizado: message = [{'role':"system",'content':"Traduza para português brasileiro"}, {'role':"user","content":sample['text']} ] Exemplo de uma amostra Original Daughtry's latest album, "Baptized", marks a significant departure from their previous work, as the band explores new sounds and… See the full description on the dataset page: https://huggingface.co/datasets/Fazzioni/recycling_the_web-pt-br.text10M<n<100M0 likes370 downloads1y agoHugging Face11brunosdorneles /common_voice_ptaudio100K<n<1M0 likes350 downloads2y agoHugging Face12Madras1 /corpus-ptbr-v1 🇧🇷 Corpus PT-BR v1 Um corpus em Português Brasileiro voltado para pré-treinamento e fine-tuning de LLMs. Combina dados reais curados com uma camada sintética construída para ampliar a diversidade estilística, lexical e discursiva em português. 🔄 Visão Geral do Pipeline 📊 Estatísticas Métrica Valor Total de documentos 8,399,857 Total de palavras ~4.84B Tokens estimados ~6.29B Tamanho (Parquet) ~17.9 GB Idioma Português Brasileiro… See the full description on the dataset page: https://huggingface.co/datasets/Madras1/corpus-ptbr-v1.tabulartext-generation1M<n<10M5 likes348 downloads5mo agoHugging Face13iara-project /news-articles-ptbr-dataset Dataset Card for "news-articles-ptbr-dataset" More Information needed text100K<n<1M4 likes341 downloads3y agoHugging Face14liaad /PtBrVId PtBrVId PtBrVId is a Portuguese Variety Identification corpus, built by combining pre-existing datasets originally created for different NLP tasks and released under permissive licenses. Our goal is to provide a large, diverse, and multi-domain resource for studying and improving automatic identification of European Portuguese (PT-PT) and Brazilian Portuguese (PT-BR). 📚 Data Sources The corpus is composed of datasets from various domains, each selected to ensure… See the full description on the dataset page: https://huggingface.co/datasets/liaad/PtBrVId.text1M<n<10M6 likes338 downloads1y agoHugging Face15DedeProGames /claude-code-traces-pt-brThis dataset was generated using teich by TeichAI claude Agent Traces This directory contains raw agent trace files generated by teich. JSONL files: 20 Training-ready tools Generated agent traces carry configured or recovered tool schemas so tools remain available for training even when a session did not call them. Native Claude Code imports recover schemas for Claude Code and Claude Desktop built-ins, plus conservative name-derived MCP schemas, when the raw… See the full description on the dataset page: https://huggingface.co/datasets/DedeProGames/claude-code-traces-pt-br.tabulartext-generationn<1K2 likes207 downloads2mo agoHugging Face16BrunoN-Dev /corpus-ptbr-v1 🇧🇷 Corpus PT-BR v1 Um corpus em Português Brasileiro voltado para pré-treinamento e fine-tuning de LLMs. Combina dados reais curados com uma camada sintética construída para ampliar a diversidade estilística, lexical e discursiva em português. 🔄 Visão Geral do Pipeline 📊 Estatísticas Métrica Valor Total de documentos 8,399,857 Total de palavras ~4.84B Tokens estimados ~6.29B Tamanho (Parquet) ~17.9 GB Idioma Português… See the full description on the dataset page: https://huggingface.co/datasets/BrunoN-Dev/corpus-ptbr-v1.tabulartext-generation1M<n<10M1 likes207 downloads2mo agoHugging Face17u1537782 /PtBrVIdtext1M<n<10M0 likes202 downloads2y agoHugging Face18lucianosb /cetacean-ptbrThis dataset is a merge of Open-Orca and Dolphin translated to portuguese. text1M<n<10M3 likes194 downloads2y agoHugging Face19bratao /corpus-ptbr-v1 🇧🇷 Corpus PT-BR v1 Um corpus em Português Brasileiro voltado para pré-treinamento e fine-tuning de LLMs. Combina dados reais curados com uma camada sintética construída para ampliar a diversidade estilística, lexical e discursiva em português. 🔄 Visão Geral do Pipeline 📊 Estatísticas Métrica Valor Total de documentos 8,399,857 Total de palavras ~4.84B Tokens estimados ~6.29B Tamanho (Parquet) ~17.9 GB Idioma Português Brasileiro (pt-br)… See the full description on the dataset page: https://huggingface.co/datasets/bratao/corpus-ptbr-v1.tabulartext-generation1M<n<10M0 likes192 downloads5mo agoHugging Face20opedromartins /asr-leaderboard-datasets-ptbraudio10K<n<100K0 likes188 downloads1mo agoHugging Face21recogna-nlp /ultra-alpaca-ptbrtext100K<n<1M5 likes181 downloads2y agoHugging Face22cnmoro /Instruct-PTBR-ENUS-11MThis dataset is a mix of multiple instruct datasets found on huggingface, while also including a bunch of other datasets (self-made) for tasks such as question-answering focused on RAG, summarization, keyword generation and others. Most of the original dataset was in the English language. I have translated most of it to Brazillian Portuguese. There is a “LANGUAGE” column, which indicates if its PT or EN. It is possible that the translation contains errors. For RAG, summarization and keyword… See the full description on the dataset page: https://huggingface.co/datasets/cnmoro/Instruct-PTBR-ENUS-11M.textquestion-answering1M<n<10M14 likes159 downloads3y agoHugging Face23cnmoro /AllTripletsMsMarco-PTBRNeed a huge dataset translated? Connect with me! text10M<n<100M7 likes159 downloads2y agoHugging Face24Madras1 /rag-qa-fulltext-ptbr RAG QA Full-Text PT-BR Mistral A large-scale dataset of Brazilian Portuguese RAG-style question-answer pairs with grounded evidence spans, generated from Madras1/corpus-ptbr-v1 documents using Mistral models. Every answer is anchored to literal quotations from the source text, making this dataset suitable for training and evaluating retrieval-augmented generation systems, extractive QA models, and reading comprehension benchmarks in Portuguese. Two configurations are available:… See the full description on the dataset page: https://huggingface.co/datasets/Madras1/rag-qa-fulltext-ptbr.tabularquestion-answering1M<n<10M0 likes159 downloads5mo agoHugging Face25adalbertojunior /punctuation-ptbrtext100K<n<1M0 likes152 downloads5y agoHugging Face26laicsiifes /flickr30k-pt-br 🎉 Flickr30K Translated for Portuguese Image Captioning 💾 Dataset Summary Flickr30K Portuguese Translated, a multimodal dataset for Portuguese image captioning with 31,014 images, each accompanied by five descriptive captions that have been generated by human annotators for every individual image. The original English captions were rendered into Portuguese through the utilization of the Google Translator API. The dataset is one of the results of work available at:… See the full description on the dataset page: https://huggingface.co/datasets/laicsiifes/flickr30k-pt-br.imagetext-generation10K<n<100K4 likes146 downloads1y agoHugging Face27tech4humans /Audio-Transcription-Models-Comparison-PT-BR Audio Transcription Models Comparison A dataset dedicated to comparing the performance of modern Speech-to-Text (STT) models, focusing exclusively on Brazilian Portuguese. About the Dataset This dataset was created to store and compare transcription results from different Artificial Intelligence models in challenging scenarios. Unlike generic benchmarks, this project focuses on the reality of usage in Brazil, covering: Regionalism: Local vocabulary, accents, and… See the full description on the dataset page: https://huggingface.co/datasets/tech4humans/Audio-Transcription-Models-Comparison-PT-BR.audioautomatic-speech-recognitionn<1K3 likes144 downloads8mo agoHugging Face28Tharyck /multispeaker-tts-ptbrDataset importado do https://gitlab.com/fb-audio-corpora audio100K<n<1M6 likes137 downloads1y agoHugging Face29marcosremar2 /pt-br-tts-synth pt-br-tts-synth 98813 frases PT-BR sintetizadas com Kokoro-82M (vozes pf_dora/pm_alex/pm_santa, speeds 0.9-1.1x), 16kHz mono WAV em 32 tar shards (WebDataset). Texto gerado por LLM (3 tiers de complexidade x 40 topicos); transcricao, tier, topico, voz e speed em metadata.jsonl. from datasets import load_dataset ds = load_dataset("webdataset", data_files="hf://datasets/marcosremar2/pt-br-tts-synth/shard_*.tar", split="train") audiotext-to-speech10K<n<100K0 likes135 downloads4mo agoHugging Face30adalbertojunior /punctuation-ptbr-lighttext10K<n<100K0 likes131 downloads5y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.