datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
BrazilianToxicTweetsClassification
BrazilianToxicTweetsClassification
An MTEB dataset
Massive Text Embedding Benchmark
ToLD-Br is the biggest dataset for toxic tweets in Brazilian Portuguese, crowdsourced by 42 annotators selected from
a pool of 129 volunteers. Annotators were selected aiming to create a plural group in terms of demographics (ethnicity,
sexual orientation, age, gender). Each tweet was labeled by three annotators in 6 possible categories: LGBTQ+phobia,
Xenophobia, Obscene, Insult… See the full description on the dataset page: https://huggingface.co/datasets/mteb/BrazilianToxicTweetsClassification.brazil-2026-electoral-divergence
AFOS — Brazil 2026 Electoral Divergence Dataset
🌐 English · Português · Español
English
Open, auditable daily dataset that cross-references prediction markets (Polymarket) × polling institutes (TSE-registered) × press coverage for Brazil's 2026 presidential cycle, with explicit divergence between sources instead of smoothed averages.
Maintained by AFOS Analytics — open-source civic infrastructure for electoral political-risk intelligence. This is the public… See the full description on the dataset page: https://huggingface.co/datasets/AFOS-Analytics1/brazil-2026-electoral-divergence.Nemotron-Personas-Brazil
Nemotron-Personas-Brazil
Abordagem de IA composta para geração de personas baseada em distribuições do mundo real
Visão Geral do Conjunto de Dados (Dataset Overview):
Nemotron-Personas-Brazil é um conjunto de dados (dataset) de código aberto (CC BY 4.0) composto por personas geradas sinteticamente e fundamentadas em distribuições demográficas, geográficas e traços de personalidade reais do Brasil, visando capturar a diversidade e a riqueza da população. Trata-se de… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Personas-Brazil.weather-rescue-brazil
Weather Rescue Brazil
Daily meteorological observations from Brazilian observatories, 1883–1890,
transcribed from printed nineteenth-century tables by an open 2B model running
offline, with every row carrying its provenance and a quality verdict.
This is an independent project. It is not affiliated with Zooniverse or with
the Weather Rescue / Rainfall Rescue projects, whose naming family it
gratefully follows.
Citing it
DOI: 10.5281/zenodo.22876699 — the concept… See the full description on the dataset page: https://huggingface.co/datasets/akagabi/weather-rescue-brazil.ENADE_Brazilian_national_university_examination_MCQ_483brazilian_court_decisions
Dataset Card for predicting-brazilian-court-decisions
Dataset Summary
The dataset is a collection of 4043 Ementa (summary) court decisions and their metadata from
the Tribunal de Justiça de Alagoas (TJAL, the State Supreme Court of Alagoas (Brazil). The court decisions are labeled
according to 7 categories and whether the decisions were unanimous on the part of the judges or not. The dataset
supports the task of Legal Judgment Prediction.
Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/joelniklaus/brazilian_court_decisions.news_sources_brazil
news_sources_brazil
17,892 news outlets, one per row, keyed to the IBGE municipality. 16,298 come from atlas da notícia, the survey of local journalism that projor and volt data lab have run since 2017. the other 1,594 are the portals, wire agencies, fact-checkers and international references that truw, my news verification project, was already tracking. the municipality table ships alongside, with 2022 population, HDI, outlet count and the news desert flag, so the whole thing… See the full description on the dataset page: https://huggingface.co/datasets/brennercruvinel/news_sources_brazil.openai-terra-batch-wiki-brazil-1000-partial-20260724-01
OpenAI Terra Batch — Wikipédia PT-BR (run parcial)
Checkpoint publicável de uma execução real e interrompida do fluxo
document_task_matrix. A execução planejou gerar uma matriz de 1.000
documentos da Wikipédia em português por 25 tasks canônicas usando a Responses
API Batch e o modelo gpt-5.6-terra.
Este repositório não representa a conclusão dos 25.000 pares planejados. Ele
contém somente os 1.282 candidatos aceitos após a reconciliação offline de
todos os resultados Batch já… See the full description on the dataset page: https://huggingface.co/datasets/costadev00/openai-terra-batch-wiki-brazil-1000-partial-20260724-01.brazilian-customer-service-conversations
Brazilian Customer Service Conversations
Dataset de conversas de atendimento ao cliente em portugues brasileiro (PT-BR).
De um like me apoie em manter esse dataset!
Descricao
Conversas sinteticas de alta qualidade simulando interacoes reais entre clientes e atendentes em diversos setores da economia brasileira. Util para treinar e avaliar modelos de:
Chatbots de atendimento
Classificacao de intencao (intent classification)
Analise de sentimento em conversas
Geracao de… See the full description on the dataset page: https://huggingface.co/datasets/RichardSakaguchiMS/brazilian-customer-service-conversations.ipfs_brazil_laws_ir
Brazil legislation IR (CID-keyed sparse GraphRAG)
Research retrieval release of endomorphosis/ipfs_brazil_laws (revision 4da9b81784842bcce890bcdd4b837819e2f84fc2) packaged as
country-laws-ir-graphrag/v1 (layout family skillcenter-huggingface-release/v3 / publicus-ir).
Not legal advice. This is a research snapshot. The official gazette /
authentic source of Brazil prevails over this corpus. Retrieved documents
and graph edges are retrieval evidence only. No legal text was… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/ipfs_brazil_laws_ir.dou-brazil-dataset
Dataset Card for Dataset Diário Oficial da União (DOU)
The Diário Oficial da União (DOU) is the official government gazette of Brazil, published by the National Press. It serves as the primary means of communication for federal government acts, including laws, decrees, ordinances, public notices, and other official decisions. The DOU ensures transparency and legal validity for government actions and is divided into three sections:
Section 1: Publishes laws, decrees, and… See the full description on the dataset page: https://huggingface.co/datasets/gerson-vfs/dou-brazil-dataset.brazilian-federal-law-corpus
Brazilian Federal Law Corpus (PT-BR)
41,526 passages of Brazilian federal legislation (laws and decrees — articles, titles, ementas) scraped from
planalto.gov.br. Official Brazilian legal texts are public domain (Lei 9.610/1998, art. 8). A clean corpus for
Portuguese legal IR / RAG / language modeling, with no LLM-generated content — pure law text — and
decontaminated against MTEB(por).
Companion to the synthetic QA pairs in
tardellirs/brazilian-legal-tax-qa-synthetic.… See the full description on the dataset page: https://huggingface.co/datasets/tardellirs/brazilian-federal-law-corpus.ipfs_brazil_laws
Brazil Federal Constitution and Laws (LexML / Planalto)
Research snapshot of official national legislation from LexML Brasil / Palacio do Planalto / Senado Dados Abertos.
Not legal advice. The official gazette / authentic source prevails over this corpus.
Snapshot
Field
Value
Snapshot date
2026-09-03
Coverage
snapshot
Source
LexML Brasil / Palacio do Planalto / Senado Dados Abertos
Collector
scrapers/collect_lexml.py
Laws / instruments
16,875… See the full description on the dataset page: https://huggingface.co/datasets/endomorphosis/ipfs_brazil_laws.brazilian-legal-tax-qa-synthetic
Brazilian Legal & Tax QA — synthetic retrieval pairs (PT-BR)
4,278 (query, positive-passage) pairs for training/evaluating Portuguese legal & tax retrieval. Queries are
controlled-diversity synthetic questions; positive passages are Brazilian federal law (public domain). Built for
the underserved Brazilian legal-IR setting, and decontaminated against MTEB(por) so it is safe to train on and
evaluate on common PT benchmarks.
How it was made
Positives: real… See the full description on the dataset page: https://huggingface.co/datasets/tardellirs/brazilian-legal-tax-qa-synthetic.brazilian-digital-identity-instagram-corpus
Brazilian Digital Identity on Instagram — Thematic Corpus
Authors: Étore Braga e Santos (Unicamp, ORCID 0009-0000-3502-705X) & Pâmella Fernandes de Sá (USP FEA-RP)License: CC BY 4.0 (dataset) — code at https://github.com/Etore-BeS/brazilian-digital-identity-paper (MIT)Version: 1.0.0 — Collection March 7–11, 2026
Dataset Description
Anonymized corpus of 17,287 Instagram comments on three international celebrity posts with substantial Brazilian engagement, annotated… See the full description on the dataset page: https://huggingface.co/datasets/Etore-BeS/brazilian-digital-identity-instagram-corpus.news-of-the-brazilian-newspaper
News of the Brazilian Newspaper
This repository contains a comprehensive dataset of news articles from a Brazilian newspaper, Folha de São Paulo (http://www.folha.uol.com.br/). The dataset includes 167,053 examples of news articles, comprising headlines, URLs of articles, complete articles, and their respective categories.
Dataset Creation
The headlines were initially gathered from Inshorts and were then used to scrape the complete news articles from Folha de São Paulo.… See the full description on the dataset page: https://huggingface.co/datasets/emdemor/news-of-the-brazilian-newspaper.medical-audio-sample-brazilian-portuguese
Julia's Data: Brazilian Portuguese Medical Audio Sample
Public sample of a Brazilian Portuguese medical audio dataset built for ASR,
TTS, and conversational AI evaluation. This repository contains deidentified
clinical source material transformed into five spoken content types and
recorded by a human speaker.
This sample includes 1 record, 20 aligned audio segments, 1 speaker, and about
5.26 minutes of audio.
Full dataset and commercial licensing: juliasdata.com
Commercial overview:… See the full description on the dataset page: https://huggingface.co/datasets/juliasdata/medical-audio-sample-brazilian-portuguese.central_bank_of_brazil
Dataset Summary
For dataset summary, please refer to https://huggingface.co/datasets/gtfintechlab/central_bank_of_brazil
Additional Information
This dataset is annotated across three different tasks: Stance Detection, Temporal Classification, and Uncertainty Estimation. The tasks have four, two, and two unique labels, respectively. This dataset contains 1,000 sentences taken from the meeting minutes of the Central Bank of Brazil.
Label Interpretation… See the full description on the dataset page: https://huggingface.co/datasets/gtfintechlab/central_bank_of_brazil.brazil-deputy-expensesbrazilian_european_portuguese_datasetbrazil-pix-databrazil-solar-generation
Brazil Daily Solar Power Generation
Hourly generation records for solar ("Fotovoltaica") power plants across
Brazil's national grid, filtered from ONS's (Operador Nacional do Sistema
Elétrico) public per-plant generation dataset.
Columns
din_instante — timestamp of the hourly reading
nom_tipousina — plant type (filtered to solar/Fotovoltaica only)
nom_usina — power plant name
val_geracao — generation for that hour, in average MW (MWmed)
Coverage… See the full description on the dataset page: https://huggingface.co/datasets/J-Marcos-GS/brazil-solar-generation.Brazilian-Sign-Language-Alphabet-Datasetallenai_WildChat_Brazil_or_Portuguesebrazilian-cultural-video-dataset
Bamboo Data Brazilian Cultural Video Dataset (Sample)
⚠️ License Notice: Evaluation Only
This is a sample of the Bamboo Data brazilian cultural video dataset, provided for internal evaluation purposes ONLY. The use of this data is strictly limited by the license defined below.
Any use for training, fine-tuning, or inference of AI/ML models, or any commercial activity, is strictly prohibited with this sample.
Dataset Description
The Bamboo Data… See the full description on the dataset page: https://huggingface.co/datasets/bamboodata/brazilian-cultural-video-dataset.brazil-wildfire-geospatial-dataset
Banco Histórico de Incêndios no Brasil — 2018–2025
Banco de dados com 1,43 milhão de focos de calor registrados no Brasil entre 2018 e 2025, enriquecidos com dados meteorológicos (ERA5), cobertura do solo (MapBiomas) e altitude (SRTM).
Construído a partir de fontes públicas oficiais para suporte a pesquisas científicas sobre incêndios florestais.
Tabelas disponíveis
Tabela
Arquivos
Linhas
Descrição
focos_analise
focos_analise/*.parquet
1.430.756
Tabela… See the full description on the dataset page: https://huggingface.co/datasets/mateus-pcosta/brazil-wildfire-geospatial-dataset.adaption-brazil-crypto-regulatory-qa
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-brazil_crypto_regulatory_qa
This dataset contains prompt-completion pairs focused on Brazilian financial regulations regarding cryptocurrencies, tokens, and digital assets. The content specifically addresses the role of the CVM (Comissão de Valores Mobiliários), risk assessments for investors, and compliance with laws such as Lei 14.478/2022. Each entry provides structured JSON… See the full description on the dataset page: https://huggingface.co/datasets/Fernandosr85/adaption-brazil-crypto-regulatory-qa.Hourly-Electricity-Demand-Brazil-Dataset
🇧🇷 Hourly Load Curve - ONS (Brazil)
This dataset contains hourly electricity load data for Brazil, published by the ONS - National Electric System Operator. It spans from the year 2000 to the present (currently 2025), with continuous updates.
📌 Description
The data represents the hourly electricity demand profile across the Brazilian National Interconnected System (SIN). It is especially suitable for:
Electricity load forecasting
Energy demand pattern analysis
Time… See the full description on the dataset page: https://huggingface.co/datasets/SamuelM0422/Hourly-Electricity-Demand-Brazil-Dataset.Brazilian_CLT_preferencesDataset DescriptionThis dataset contains 736 validated human-preference entries designed to align language models with expert expectations for answering questions about Brazil’s Consolidation of Labor Laws (CLT). It was created to support Direct Preference Optimization (DPO) fine-tuning and evaluation of LLM-based legal assistants.
Intended Use
Primary Purpose: Training and evaluating models for legal question answering under the Brazilian CLT framework.
Target Users: Researchers… See the full description on the dataset page: https://huggingface.co/datasets/ai-eldorado/Brazilian_CLT_preferences.apple_detection_drone_brazil
Apple Detection Drone Brazil
A dataset for object detection of apples. The dataset contains 689 images with 2,471 bounding box annotations across 1 category.
This dataset is indexed on https://project-agml.github.io/ as part of the AgML python library.
Citation
@article{DBLP:journals/corr/abs-2110-12331,
author={Thiago T. Santos and Luciano Gebler},
title={A methodology for detection and localization of fruits in apples orchards from aerial images},
journal={CoRR}… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/apple_detection_drone_brazil.
