datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
BrazilianToxicTweetsClassification
BrazilianToxicTweetsClassification
An MTEB dataset
Massive Text Embedding Benchmark
ToLD-Br is the biggest dataset for toxic tweets in Brazilian Portuguese, crowdsourced by 42 annotators selected from
a pool of 129 volunteers. Annotators were selected aiming to create a plural group in terms of demographics (ethnicity,
sexual orientation, age, gender). Each tweet was labeled by three annotators in 6 possible categories: LGBTQ+phobia,
Xenophobia, Obscene, Insult… See the full description on the dataset page: https://huggingface.co/datasets/mteb/BrazilianToxicTweetsClassification.brazil-2026-electoral-divergence
AFOS — Brazil 2026 Electoral Divergence Dataset
🌐 English · Português · Español
English
Open, auditable daily dataset that cross-references prediction markets (Polymarket) × polling institutes (TSE-registered) × press coverage for Brazil's 2026 presidential cycle, with explicit divergence between sources instead of smoothed averages.
Maintained by AFOS Analytics — open-source civic infrastructure for electoral political-risk intelligence. This is the public… See the full description on the dataset page: https://huggingface.co/datasets/AFOS-Analytics1/brazil-2026-electoral-divergence.Nemotron-Personas-Brazil
Nemotron-Personas-Brazil
Abordagem de IA composta para geração de personas baseada em distribuições do mundo real
Visão Geral do Conjunto de Dados (Dataset Overview):
Nemotron-Personas-Brazil é um conjunto de dados (dataset) de código aberto (CC BY 4.0) composto por personas geradas sinteticamente e fundamentadas em distribuições demográficas, geográficas e traços de personalidade reais do Brasil, visando capturar a diversidade e a riqueza da população. Trata-se de… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Personas-Brazil.weather-rescue-brazil
Weather Rescue Brazil
Daily meteorological observations from Brazilian observatories, 1883–1890,
transcribed from printed nineteenth-century tables by an open 2B model running
offline, with every row carrying its provenance and a quality verdict.
This is an independent project. It is not affiliated with Zooniverse or with
the Weather Rescue / Rainfall Rescue projects, whose naming family it
gratefully follows.
Citing it
DOI: 10.5281/zenodo.22876699 — the concept… See the full description on the dataset page: https://huggingface.co/datasets/akagabi/weather-rescue-brazil.synthetic-printed-brazilian-passports
Brazilian passport dataset
The dataset comprises 5,000 high-resolution synthetic photos of Brazilian passports, designed to advance computer vision and identity verification systems. It provides a secure and ethical resource for training robust models for OCR (Optical Character Recognition), document analysis, and spoofing detection, all without exposing real personal data or sensitive personal information.
By utilizing this dataset, researchers and developers can enhance… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/synthetic-printed-brazilian-passports.brazil-ndviENADE_Brazilian_national_university_examination_MCQ_483brazilian_court_decisions
Dataset Card for predicting-brazilian-court-decisions
Dataset Summary
The dataset is a collection of 4043 Ementa (summary) court decisions and their metadata from
the Tribunal de Justiça de Alagoas (TJAL, the State Supreme Court of Alagoas (Brazil). The court decisions are labeled
according to 7 categories and whether the decisions were unanimous on the part of the judges or not. The dataset
supports the task of Legal Judgment Prediction.
Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/joelniklaus/brazilian_court_decisions.lpr-brazil-700knews_sources_brazil
news_sources_brazil
17,892 news outlets, one per row, keyed to the IBGE municipality. 16,298 come from atlas da notícia, the survey of local journalism that projor and volt data lab have run since 2017. the other 1,594 are the portals, wire agencies, fact-checkers and international references that truw, my news verification project, was already tracking. the municipality table ships alongside, with 2022 population, HDI, outlet count and the news desert flag, so the whole thing… See the full description on the dataset page: https://huggingface.co/datasets/brennercruvinel/news_sources_brazil.openai-terra-batch-wiki-brazil-1000-partial-20260724-01
OpenAI Terra Batch — Wikipédia PT-BR (run parcial)
Checkpoint publicável de uma execução real e interrompida do fluxo
document_task_matrix. A execução planejou gerar uma matriz de 1.000
documentos da Wikipédia em português por 25 tasks canônicas usando a Responses
API Batch e o modelo gpt-5.6-terra.
Este repositório não representa a conclusão dos 25.000 pares planejados. Ele
contém somente os 1.282 candidatos aceitos após a reconciliação offline de
todos os resultados Batch já… See the full description on the dataset page: https://huggingface.co/datasets/costadev00/openai-terra-batch-wiki-brazil-1000-partial-20260724-01.brazilian-customer-service-conversations
Brazilian Customer Service Conversations
Dataset de conversas de atendimento ao cliente em portugues brasileiro (PT-BR).
De um like me apoie em manter esse dataset!
Descricao
Conversas sinteticas de alta qualidade simulando interacoes reais entre clientes e atendentes em diversos setores da economia brasileira. Util para treinar e avaliar modelos de:
Chatbots de atendimento
Classificacao de intencao (intent classification)
Analise de sentimento em conversas
Geracao de… See the full description on the dataset page: https://huggingface.co/datasets/RichardSakaguchiMS/brazilian-customer-service-conversations.Brazilian-Forensic-Anomalies-Benchmark
BYTEN Forensic Anomalies Benchmark (v1.0)
Matriz de autoridade em engenharia forense e detecção de paradoxos temporais para o Judiciário Brasileiro.
Overview
Este dataset fornece a base lógica e semântica para a detecção de anomalias em evidências digitais. Ele foi projetado para calibrar modelos de linguagem (LLMs) e sistemas de RAG (Retrieval-Augmented Generation) na identificação de fraudes processuais, garantindo conformidade com o Art. 158-B do CPP e a ISO/IEC… See the full description on the dataset page: https://huggingface.co/datasets/agustinhoce/Brazilian-Forensic-Anomalies-Benchmark.solar-plants-brazil
🛰️ Solar Plants Brazil
Solar Plants Brazil is a geospatial dataset for binary semantic segmentation of photovoltaic (PV) solar power stations in satellite imagery. It consists of multi-spectral image tiles (including near-infrared) with pixel-level annotations indicating the presence of solar panels. This dataset enables training and evaluating deep learning models that automatically detect solar farm installations from overhead imagery, supporting applications in renewable energy… See the full description on the dataset page: https://huggingface.co/datasets/FederCO23/solar-plants-brazil.ipfs_brazil_laws_ir
Brazil legislation IR (CID-keyed sparse GraphRAG)
Research retrieval release of endomorphosis/ipfs_brazil_laws (revision 4da9b81784842bcce890bcdd4b837819e2f84fc2) packaged as
country-laws-ir-graphrag/v1 (layout family skillcenter-huggingface-release/v3 / publicus-ir).
Not legal advice. This is a research snapshot. The official gazette /
authentic source of Brazil prevails over this corpus. Retrieved documents
and graph edges are retrieval evidence only. No legal text was… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/ipfs_brazil_laws_ir.Brazilian_Bills_and_Invoices_Dataset
Brazilian Bills and Invoices Dataset
This dataset contains high-quality scanned and photographed images of Brazilian bills, invoices, and utility payment documents. It supports AI research in OCR, financial document understanding, and structured data extraction for Portuguese-language financial contexts.
Contact
For queries or collaborations related to this dataset, contact:
anoushka@kgen.io
abhishek.vadapalli@kgen.io
Supported Tasks
Task Categories:… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/Brazilian_Bills_and_Invoices_Dataset.bias-within-borders-brazil-images-ptdou-brazil-dataset
Dataset Card for Dataset Diário Oficial da União (DOU)
The Diário Oficial da União (DOU) is the official government gazette of Brazil, published by the National Press. It serves as the primary means of communication for federal government acts, including laws, decrees, ordinances, public notices, and other official decisions. The DOU ensures transparency and legal validity for government actions and is divided into three sections:
Section 1: Publishes laws, decrees, and… See the full description on the dataset page: https://huggingface.co/datasets/gerson-vfs/dou-brazil-dataset.brazilian-federal-law-corpus
Brazilian Federal Law Corpus (PT-BR)
41,526 passages of Brazilian federal legislation (laws and decrees — articles, titles, ementas) scraped from
planalto.gov.br. Official Brazilian legal texts are public domain (Lei 9.610/1998, art. 8). A clean corpus for
Portuguese legal IR / RAG / language modeling, with no LLM-generated content — pure law text — and
decontaminated against MTEB(por).
Companion to the synthetic QA pairs in
tardellirs/brazilian-legal-tax-qa-synthetic.… See the full description on the dataset page: https://huggingface.co/datasets/tardellirs/brazilian-federal-law-corpus.ipfs_brazil_laws
Brazil Federal Constitution and Laws (LexML / Planalto)
Research snapshot of official national legislation from LexML Brasil / Palacio do Planalto / Senado Dados Abertos.
Not legal advice. The official gazette / authentic source prevails over this corpus.
Snapshot
Field
Value
Snapshot date
2026-09-03
Coverage
snapshot
Source
LexML Brasil / Palacio do Planalto / Senado Dados Abertos
Collector
scrapers/collect_lexml.py
Laws / instruments
16,875… See the full description on the dataset page: https://huggingface.co/datasets/endomorphosis/ipfs_brazil_laws.brazilian-legal-tax-qa-synthetic
Brazilian Legal & Tax QA — synthetic retrieval pairs (PT-BR)
4,278 (query, positive-passage) pairs for training/evaluating Portuguese legal & tax retrieval. Queries are
controlled-diversity synthetic questions; positive passages are Brazilian federal law (public domain). Built for
the underserved Brazilian legal-IR setting, and decontaminated against MTEB(por) so it is safe to train on and
evaluate on common PT benchmarks.
How it was made
Positives: real… See the full description on the dataset page: https://huggingface.co/datasets/tardellirs/brazilian-legal-tax-qa-synthetic.Brazilian_Road_Signs_Dataset
Brazilian Road Signs Dataset
This dataset contains high-quality images of Brazilian road and traffic signs collected from various urban and rural environments. It supports AI research in computer vision, object detection, and autonomous driving systems adapted to Brazil’s signage standards and language.
Contact
For queries or collaborations related to this dataset, contact:
anoushka@kgen.io
abhishek.vadapalli@kgen.io
Supported Tasks
Task Categories:… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/Brazilian_Road_Signs_Dataset.sugarcane_dataset_northeast_sao_paulo_state_brazilThis is the first publicly available dataset that integrates sugarcane crop yield, production environment, meteorological records, and Sentinel-2 satellite imagery from
commercial fields in the northeast of São Paulo State, Brazil.
The description of this dataset is in the scientific paper https://www.sciencedirect.com/science/article/pii/S2352340926001022
Please acknowledge the use of this dataset by citing the paper above and the following reference:
@MISC{Barbosa2025-uq,
title =… See the full description on the dataset page: https://huggingface.co/datasets/lafbarbosa/sugarcane_dataset_northeast_sao_paulo_state_brazil.brazil-business-dataset
Brazil Business Email List — Market Intelligence Dataset
2,800,000 verified Brazil contacts are available from LeadsBlue →. This open dataset provides the aggregate market intelligence behind that database — contact volume, benchmark open/reply rates, send timing, and compliance for the Brazil segment.
At a glance: Market-intelligence reference data for the Brazil Business Email List audience — population scale, decision-maker structure, channel benchmarks, and compliance… See the full description on the dataset page: https://huggingface.co/datasets/emailmarketingdataset/brazil-business-dataset.brazilian-digital-identity-instagram-corpus
Brazilian Digital Identity on Instagram — Thematic Corpus
Authors: Étore Braga e Santos (Unicamp, ORCID 0009-0000-3502-705X) & Pâmella Fernandes de Sá (USP FEA-RP)License: CC BY 4.0 (dataset) — code at https://github.com/Etore-BeS/brazilian-digital-identity-paper (MIT)Version: 1.0.0 — Collection March 7–11, 2026
Dataset Description
Anonymized corpus of 17,287 Instagram comments on three international celebrity posts with substantial Brazilian engagement, annotated… See the full description on the dataset page: https://huggingface.co/datasets/Etore-BeS/brazilian-digital-identity-instagram-corpus.news-of-the-brazilian-newspaper
News of the Brazilian Newspaper
This repository contains a comprehensive dataset of news articles from a Brazilian newspaper, Folha de São Paulo (http://www.folha.uol.com.br/). The dataset includes 167,053 examples of news articles, comprising headlines, URLs of articles, complete articles, and their respective categories.
Dataset Creation
The headlines were initially gathered from Inshorts and were then used to scrape the complete news articles from Folha de São Paulo.… See the full description on the dataset page: https://huggingface.co/datasets/emdemor/news-of-the-brazilian-newspaper.medical-audio-sample-brazilian-portuguese
Julia's Data: Brazilian Portuguese Medical Audio Sample
Public sample of a Brazilian Portuguese medical audio dataset built for ASR,
TTS, and conversational AI evaluation. This repository contains deidentified
clinical source material transformed into five spoken content types and
recorded by a human speaker.
This sample includes 1 record, 20 aligned audio segments, 1 speaker, and about
5.26 minutes of audio.
Full dataset and commercial licensing: juliasdata.com
Commercial overview:… See the full description on the dataset page: https://huggingface.co/datasets/juliasdata/medical-audio-sample-brazilian-portuguese.central_bank_of_brazil
Dataset Summary
For dataset summary, please refer to https://huggingface.co/datasets/gtfintechlab/central_bank_of_brazil
Additional Information
This dataset is annotated across three different tasks: Stance Detection, Temporal Classification, and Uncertainty Estimation. The tasks have four, two, and two unique labels, respectively. This dataset contains 1,000 sentences taken from the meeting minutes of the Central Bank of Brazil.
Label Interpretation… See the full description on the dataset page: https://huggingface.co/datasets/gtfintechlab/central_bank_of_brazil.brazil-deputy-expensesbrazilian_european_portuguese_dataset
