datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
VeraCruz_PT-BR
Dataset Summary
The VeraCruz Dataset is a comprehensive collection of Portuguese language content, showcasing the linguistic and cultural diversity of of Portuguese-speaking regions. It includes around 190 million samples, organized by regional origin as indicated by URL metadata into primary categories. The primary categories are:
Portugal (PT): Samples with content URLs indicating a clear Portuguese origin.
Brazil (BR): Samples with content URLs indicating a clear Brazilian… See the full description on the dataset page: https://huggingface.co/datasets/bastao/VeraCruz_PT-BR.ASR-datasets-ptbr
📚 Datasets de Áudio em Português (PT-BR)
Este repositório reúne diversos corpora públicos de fala em português do Brasil, combinados em um único dataset para facilitar treinamentos e pesquisas em ASR (Automatic Speech Recognition).
O objetivo é fornecer um recurso amplo, padronizado e de fácil acesso para a comunidade.
📂 Datasets Integrados
A tabela abaixo lista todos os datasets incluídos, com suas informações:
Dataset
Config Name
TOTAL
train
test
validation… See the full description on the dataset page: https://huggingface.co/datasets/opedromartins/ASR-datasets-ptbr.mteb-pt-results
🇧🇷 MTEB-BR — Benchmark Results
Canonical results store for MTEB-BR, a native Brazilian-Portuguese text-embedding benchmark.
93 models · 22 native PT-BR tasks · 7 categories · no machine translation
What is this?
This repository is the canonical, machine-readable results store for MTEB-BR — a benchmark that evaluates text-embedding models on native Brazilian Portuguese (data created or found in Portuguese; machine-translated corpora such as… See the full description on the dataset page: https://huggingface.co/datasets/MTEB-BR/mteb-pt-results.mTEDx-ptbr
Multilingual TEDx (Portuguese speech and transcripts)
NOTE: This dataset contains only the Portuguese portion of the mTEDx dataset, already processed and segmented into parts.
Multilingual TEDx (mTEDx) is a multilingual speech recognition and translation corpus to facilitate the training of ASR and SLT models in additional languages.
The corpus comprises audio recordings and transcripts from TEDx Talks in 8 languages (Spanish, French, Portuguese, Italian, Russian, Greek, Arabic… See the full description on the dataset page: https://huggingface.co/datasets/dominguesm/mTEDx-ptbr.alpaca-data-pt-brNOTE: This is a machine translated version of the yahma/alpaca-cleaned dataset.
Dataset Card for Alpaca-Cleaned
Repository: https://github.com/gururise/AlpacaDataCleaned
Dataset Description
This is a cleaned version of the original Alpaca Dataset released by Stanford. The following issues have been identified in the original release and fixed in this dataset:
Hallucinations: Many instructions in the original dataset had instructions referencing data on the internet… See the full description on the dataset page: https://huggingface.co/datasets/dominguesm/alpaca-data-pt-br.coco-captions-pt-br
🎉 COCO Captions Dataset Translation for Portuguese Image Captioning
💾 Dataset Summary
COCO Captions Portuguese Translation, a multimodal dataset for Portuguese image captioning with 123,287 images, each accompanied by five descriptive captions that have been
generated by human annotators for every individual image. The original English captions were rendered into Portuguese
through the utilization of the Google Translator API.
🧑💻 Hot to Get… See the full description on the dataset page: https://huggingface.co/datasets/laicsiifes/coco-captions-pt-br.notas-comunidade-ptbr
Community Notes BR: Enriched Portuguese Dataset for NLP and Text Mining
Community Notes BR is a curated Portuguese-language subset of X Community Notes enriched for Natural Language Processing (PLN), text mining, and research on collaborative misinformation moderation. It combines note-level metadata, topic and macrotheme labels, named entities, source-domain extraction, Matrix Factorization scoring fields, an optional Universal Dependencies syntax layer, and automatic… See the full description on the dataset page: https://huggingface.co/datasets/histlearn/notas-comunidade-ptbr.Canarim-Instruct-PTBR-Dataset
🐥 🇧🇷 Canarim Instruct Dataset
[🐱 Github]
What's Canarim?
Canarim is a dataset with over 300,000 instructions in Portuguese, ranging from simple instructions like "Descreva os efeitos do aquecimento global" to more complex instructions like "Nesta tarefa, você precisa ser capaz de resumir uma determinada lista de pontos-chave" where additional context is provided.
Why it's called Canarim?
"Canarim" is spoken in some regions of Brazil (mainly by… See the full description on the dataset page: https://huggingface.co/datasets/dominguesm/Canarim-Instruct-PTBR-Dataset.wikipedia-ptbr-20230601
Dataset Card for "wikipedia-ptbr-20230601"
More Information needed
recycling_the_web-pt-brAproximadamente 25GB de arquivos .parquet traduzidos com o openai/gpt-oss-120b a partir do dataset:
#Facebook/recycling_the_web
prompt utilizado:
message = [{'role':"system",'content':"Traduza para português brasileiro"},
{'role':"user","content":sample['text']}
]
Exemplo de uma amostra
Original
Daughtry's latest album, "Baptized", marks a significant departure from their previous work, as the band explores new sounds and… See the full description on the dataset page: https://huggingface.co/datasets/Fazzioni/recycling_the_web-pt-br.common_voice_ptcorpus-ptbr-v1
🇧🇷 Corpus PT-BR v1
Um corpus em Português Brasileiro voltado para pré-treinamento e fine-tuning de LLMs. Combina dados reais curados com uma camada sintética construída para ampliar a diversidade estilística, lexical e discursiva em português.
🔄 Visão Geral do Pipeline
📊 Estatísticas
Métrica
Valor
Total de documentos
8,399,857
Total de palavras
~4.84B
Tokens estimados
~6.29B
Tamanho (Parquet)
~17.9 GB
Idioma
Português Brasileiro… See the full description on the dataset page: https://huggingface.co/datasets/Madras1/corpus-ptbr-v1.news-articles-ptbr-dataset
Dataset Card for "news-articles-ptbr-dataset"
More Information needed
PtBrVId
PtBrVId
PtBrVId is a Portuguese Variety Identification corpus, built by combining pre-existing datasets originally created for different NLP tasks and released under permissive licenses.
Our goal is to provide a large, diverse, and multi-domain resource for studying and improving automatic identification of European Portuguese (PT-PT) and Brazilian Portuguese (PT-BR).
📚 Data Sources
The corpus is composed of datasets from various domains, each selected to ensure… See the full description on the dataset page: https://huggingface.co/datasets/liaad/PtBrVId.claude-code-traces-pt-brThis dataset was generated using teich by TeichAI
claude Agent Traces
This directory contains raw agent trace files generated by teich.
JSONL files: 20
Training-ready tools
Generated agent traces carry configured or recovered tool schemas so tools remain available for training even when a session did not call them.
Native Claude Code imports recover schemas for Claude Code and Claude Desktop built-ins, plus conservative name-derived MCP schemas, when the raw… See the full description on the dataset page: https://huggingface.co/datasets/DedeProGames/claude-code-traces-pt-br.corpus-ptbr-v1
🇧🇷 Corpus PT-BR v1
Um corpus em Português Brasileiro voltado para pré-treinamento e fine-tuning de LLMs. Combina dados reais curados com uma camada sintética construída para ampliar a diversidade estilística, lexical e discursiva em português.
🔄 Visão Geral do Pipeline
📊 Estatísticas
Métrica
Valor
Total de documentos
8,399,857
Total de palavras
~4.84B
Tokens estimados
~6.29B
Tamanho (Parquet)
~17.9 GB
Idioma
Português… See the full description on the dataset page: https://huggingface.co/datasets/BrunoN-Dev/corpus-ptbr-v1.PtBrVIdcetacean-ptbrThis dataset is a merge of Open-Orca and Dolphin translated to portuguese.
corpus-ptbr-v1
🇧🇷 Corpus PT-BR v1
Um corpus em Português Brasileiro voltado para pré-treinamento e fine-tuning de LLMs. Combina dados reais curados com uma camada sintética construída para ampliar a diversidade estilística, lexical e discursiva em português.
🔄 Visão Geral do Pipeline
📊 Estatísticas
Métrica
Valor
Total de documentos
8,399,857
Total de palavras
~4.84B
Tokens estimados
~6.29B
Tamanho (Parquet)
~17.9 GB
Idioma
Português Brasileiro (pt-br)… See the full description on the dataset page: https://huggingface.co/datasets/bratao/corpus-ptbr-v1.asr-leaderboard-datasets-ptbrultra-alpaca-ptbrInstruct-PTBR-ENUS-11MThis dataset is a mix of multiple instruct datasets found on huggingface, while also including a bunch of other datasets (self-made) for tasks such as question-answering focused on RAG, summarization, keyword generation and others.
Most of the original dataset was in the English language. I have translated most of it to Brazillian Portuguese. There is a “LANGUAGE” column, which indicates if its PT or EN. It is possible that the translation contains errors.
For RAG, summarization and keyword… See the full description on the dataset page: https://huggingface.co/datasets/cnmoro/Instruct-PTBR-ENUS-11M.AllTripletsMsMarco-PTBRNeed a huge dataset translated? Connect with me!
rag-qa-fulltext-ptbr
RAG QA Full-Text PT-BR Mistral
A large-scale dataset of Brazilian Portuguese RAG-style question-answer pairs
with grounded evidence spans, generated from Madras1/corpus-ptbr-v1 documents
using Mistral models. Every answer is anchored to literal quotations from the
source text, making this dataset suitable for training and evaluating
retrieval-augmented generation systems, extractive QA models, and reading
comprehension benchmarks in Portuguese.
Two configurations are available:… See the full description on the dataset page: https://huggingface.co/datasets/Madras1/rag-qa-fulltext-ptbr.punctuation-ptbrflickr30k-pt-br
🎉 Flickr30K Translated for Portuguese Image Captioning
💾 Dataset Summary
Flickr30K Portuguese Translated, a multimodal dataset for Portuguese image captioning with 31,014 images, each accompanied by five descriptive captions that have been
generated by human annotators for every individual image. The original English captions were rendered into Portuguese
through the utilization of the Google Translator API.
The dataset is one of the results of work available at:… See the full description on the dataset page: https://huggingface.co/datasets/laicsiifes/flickr30k-pt-br.Audio-Transcription-Models-Comparison-PT-BR
Audio Transcription Models Comparison
A dataset dedicated to comparing the performance of modern Speech-to-Text (STT) models, focusing exclusively on Brazilian Portuguese.
About the Dataset
This dataset was created to store and compare transcription results from different Artificial Intelligence models in challenging scenarios. Unlike generic benchmarks, this project focuses on the reality of usage in Brazil, covering:
Regionalism: Local vocabulary, accents, and… See the full description on the dataset page: https://huggingface.co/datasets/tech4humans/Audio-Transcription-Models-Comparison-PT-BR.multispeaker-tts-ptbrDataset importado do https://gitlab.com/fb-audio-corpora
pt-br-tts-synth
pt-br-tts-synth
98813 frases PT-BR sintetizadas com Kokoro-82M (vozes pf_dora/pm_alex/pm_santa, speeds 0.9-1.1x), 16kHz mono WAV em 32 tar shards (WebDataset). Texto gerado por LLM (3 tiers de complexidade x 40 topicos); transcricao, tier, topico, voz e speed em metadata.jsonl.
from datasets import load_dataset
ds = load_dataset("webdataset", data_files="hf://datasets/marcosremar2/pt-br-tts-synth/shard_*.tar", split="train")
punctuation-ptbr-light
