CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mteb /BrazilianToxicTweetsClassification BrazilianToxicTweetsClassification An MTEB dataset Massive Text Embedding Benchmark ToLD-Br is the biggest dataset for toxic tweets in Brazilian Portuguese, crowdsourced by 42 annotators selected from a pool of 129 volunteers. Annotators were selected aiming to create a plural group in terms of demographics (ethnicity, sexual orientation, age, gender). Each tweet was labeled by three annotators in 6 possible categories: LGBTQ+phobia, Xenophobia, Obscene, Insult… See the full description on the dataset page: https://huggingface.co/datasets/mteb/BrazilianToxicTweetsClassification.texttext-classification10K<n<100K1 likes6.5k downloads7mo agoHugging Face02AFOS-Analytics1 /brazil-2026-electoral-divergence AFOS — Brazil 2026 Electoral Divergence Dataset 🌐 English · Português · Español English Open, auditable daily dataset that cross-references prediction markets (Polymarket) × polling institutes (TSE-registered) × press coverage for Brazil's 2026 presidential cycle, with explicit divergence between sources instead of smoothed averages. Maintained by AFOS Analytics — open-source civic infrastructure for electoral political-risk intelligence. This is the public… See the full description on the dataset page: https://huggingface.co/datasets/AFOS-Analytics1/brazil-2026-electoral-divergence.tabular1K<n<10K2 likes5.1k downloads8h agoHugging Face03nvidia /Nemotron-Personas-Brazil Nemotron-Personas-Brazil Abordagem de IA composta para geração de personas baseada em distribuições do mundo real Visão Geral do Conjunto de Dados (Dataset Overview): Nemotron-Personas-Brazil é um conjunto de dados (dataset) de código aberto (CC BY 4.0) composto por personas geradas sinteticamente e fundamentadas em distribuições demográficas, geográficas e traços de personalidade reais do Brasil, visando capturar a diversidade e a riqueza da população. Trata-se de… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Personas-Brazil.texttext-generation1M<n<10M101 likes864 downloads8mo agoHugging Face04akagabi /weather-rescue-brazil Weather Rescue Brazil Daily meteorological observations from Brazilian observatories, 1883–1890, transcribed from printed nineteenth-century tables by an open 2B model running offline, with every row carrying its provenance and a quality verdict. This is an independent project. It is not affiliated with Zooniverse or with the Weather Rescue / Rainfall Rescue projects, whose naming family it gratefully follows. Citing it DOI: 10.5281/zenodo.22876699 — the concept… See the full description on the dataset page: https://huggingface.co/datasets/akagabi/weather-rescue-brazil.tabulartabular-to-text1K<n<10K0 likes388 downloads3d agoHugging Face05UniDataPro /synthetic-printed-brazilian-passports Brazilian passport dataset The dataset comprises 5,000 high-resolution synthetic photos of Brazilian passports, designed to advance computer vision and identity verification systems. It provides a secure and ethical resource for training robust models for OCR (Optical Character Recognition), document analysis, and spoofing detection, all without exposing real personal data or sensitive personal information. By utilizing this dataset, researchers and developers can enhance… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/synthetic-printed-brazilian-passports.imageimage-to-textn<1K1 likes360 downloads1mo agoHugging Face06riccardode /brazil-ndvi0 likes322 downloads10mo agoHugging Face07northern-64bit /ENADE_Brazilian_national_university_examination_MCQ_483textquestion-answeringn<1K0 likes298 downloads2y agoHugging Face08joelniklaus /brazilian_court_decisions Dataset Card for predicting-brazilian-court-decisions Dataset Summary The dataset is a collection of 4043 Ementa (summary) court decisions and their metadata from the Tribunal de Justiça de Alagoas (TJAL, the State Supreme Court of Alagoas (Brazil). The court decisions are labeled according to 7 categories and whether the decisions were unanimous on the part of the judges or not. The dataset supports the task of Legal Judgment Prediction. Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/joelniklaus/brazilian_court_decisions.texttext-classification1K<n<10K23 likes273 downloads4y agoHugging Face09guica /lpr-brazil-700kimage100K<n<1M4 likes233 downloads2y agoHugging Face10brennercruvinel /news_sources_brazil news_sources_brazil 17,892 news outlets, one per row, keyed to the IBGE municipality. 16,298 come from atlas da notícia, the survey of local journalism that projor and volt data lab have run since 2017. the other 1,594 are the portals, wire agencies, fact-checkers and international references that truw, my news verification project, was already tracking. the municipality table ships alongside, with 2022 population, HDI, outlet count and the news desert flag, so the whole thing… See the full description on the dataset page: https://huggingface.co/datasets/brennercruvinel/news_sources_brazil.tabular10K<n<100K0 likes174 downloads14d agoHugging Face11costadev00 /openai-terra-batch-wiki-brazil-1000-partial-20260724-01 OpenAI Terra Batch — Wikipédia PT-BR (run parcial) Checkpoint publicável de uma execução real e interrompida do fluxo document_task_matrix. A execução planejou gerar uma matriz de 1.000 documentos da Wikipédia em português por 25 tasks canônicas usando a Responses API Batch e o modelo gpt-5.6-terra. Este repositório não representa a conclusão dos 25.000 pares planejados. Ele contém somente os 1.282 candidatos aceitos após a reconciliação offline de todos os resultados Batch já… See the full description on the dataset page: https://huggingface.co/datasets/costadev00/openai-terra-batch-wiki-brazil-1000-partial-20260724-01.texttext-generation1K<n<10K0 likes171 downloads2mo agoHugging Face12RichardSakaguchiMS /brazilian-customer-service-conversations Brazilian Customer Service Conversations Dataset de conversas de atendimento ao cliente em portugues brasileiro (PT-BR). De um like me apoie em manter esse dataset! Descricao Conversas sinteticas de alta qualidade simulando interacoes reais entre clientes e atendentes em diversos setores da economia brasileira. Util para treinar e avaliar modelos de: Chatbots de atendimento Classificacao de intencao (intent classification) Analise de sentimento em conversas Geracao de… See the full description on the dataset page: https://huggingface.co/datasets/RichardSakaguchiMS/brazilian-customer-service-conversations.texttext-classificationn<1K5 likes135 downloads10mo agoHugging Face13agustinhoce /Brazilian-Forensic-Anomalies-Benchmark BYTEN Forensic Anomalies Benchmark (v1.0) Matriz de autoridade em engenharia forense e detecção de paradoxos temporais para o Judiciário Brasileiro. Overview Este dataset fornece a base lógica e semântica para a detecção de anomalias em evidências digitais. Ele foi projetado para calibrar modelos de linguagem (LLMs) e sistemas de RAG (Retrieval-Augmented Generation) na identificação de fraudes processuais, garantindo conformidade com o Art. 158-B do CPP e a ISO/IEC… See the full description on the dataset page: https://huggingface.co/datasets/agustinhoce/Brazilian-Forensic-Anomalies-Benchmark.0 likes135 downloads5mo agoHugging Face14FederCO23 /solar-plants-brazil 🛰️ Solar Plants Brazil Solar Plants Brazil is a geospatial dataset for binary semantic segmentation of photovoltaic (PV) solar power stations in satellite imagery. It consists of multi-spectral image tiles (including near-infrared) with pixel-level annotations indicating the presence of solar panels. This dataset enables training and evaluating deep learning models that automatically detect solar farm installations from overhead imagery, supporting applications in renewable energy… See the full description on the dataset page: https://huggingface.co/datasets/FederCO23/solar-plants-brazil.imageimage-segmentationn<1K3 likes118 downloads1y agoHugging Face15justicedao /ipfs_brazil_laws_ir Brazil legislation IR (CID-keyed sparse GraphRAG) Research retrieval release of endomorphosis/ipfs_brazil_laws (revision 4da9b81784842bcce890bcdd4b837819e2f84fc2) packaged as country-laws-ir-graphrag/v1 (layout family skillcenter-huggingface-release/v3 / publicus-ir). Not legal advice. This is a research snapshot. The official gazette / authentic source of Brazil prevails over this corpus. Retrieved documents and graph edges are retrieval evidence only. No legal text was… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/ipfs_brazil_laws_ir.tabulartext-retrieval1M<n<10M0 likes116 downloads3d agoHugging Face16HumynLabs /Brazilian_Bills_and_Invoices_Dataset Brazilian Bills and Invoices Dataset This dataset contains high-quality scanned and photographed images of Brazilian bills, invoices, and utility payment documents. It supports AI research in OCR, financial document understanding, and structured data extraction for Portuguese-language financial contexts. Contact For queries or collaborations related to this dataset, contact: anoushka@kgen.io abhishek.vadapalli@kgen.io Supported Tasks Task Categories:… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/Brazilian_Bills_and_Invoices_Dataset.imagen<1K2 likes99 downloads1y agoHugging Face17FelipeMahlow /bias-within-borders-brazil-images-ptimage1K<n<10K0 likes95 downloads3mo agoHugging Face18gerson-vfs /dou-brazil-dataset Dataset Card for Dataset Diário Oficial da União (DOU) The Diário Oficial da União (DOU) is the official government gazette of Brazil, published by the National Press. It serves as the primary means of communication for federal government acts, including laws, decrees, ordinances, public notices, and other official decisions. The DOU ensures transparency and legal validity for government actions and is divided into three sections: Section 1: Publishes laws, decrees, and… See the full description on the dataset page: https://huggingface.co/datasets/gerson-vfs/dou-brazil-dataset.imagetext-generation100K<n<1M0 likes87 downloads2y agoHugging Face19tardellirs /brazilian-federal-law-corpus Brazilian Federal Law Corpus (PT-BR) 41,526 passages of Brazilian federal legislation (laws and decrees — articles, titles, ementas) scraped from planalto.gov.br. Official Brazilian legal texts are public domain (Lei 9.610/1998, art. 8). A clean corpus for Portuguese legal IR / RAG / language modeling, with no LLM-generated content — pure law text — and decontaminated against MTEB(por). Companion to the synthetic QA pairs in tardellirs/brazilian-legal-tax-qa-synthetic.… See the full description on the dataset page: https://huggingface.co/datasets/tardellirs/brazilian-federal-law-corpus.texttext-retrieval10K<n<100K4 likes84 downloads4mo agoHugging Face20endomorphosis /ipfs_brazil_laws Brazil Federal Constitution and Laws (LexML / Planalto) Research snapshot of official national legislation from LexML Brasil / Palacio do Planalto / Senado Dados Abertos. Not legal advice. The official gazette / authentic source prevails over this corpus. Snapshot Field Value Snapshot date 2026-09-03 Coverage snapshot Source LexML Brasil / Palacio do Planalto / Senado Dados Abertos Collector scrapers/collect_lexml.py Laws / instruments 16,875… See the full description on the dataset page: https://huggingface.co/datasets/endomorphosis/ipfs_brazil_laws.texttext-retrieval100K<n<1M1 likes71 downloads22d agoHugging Face21tardellirs /brazilian-legal-tax-qa-synthetic Brazilian Legal & Tax QA — synthetic retrieval pairs (PT-BR) 4,278 (query, positive-passage) pairs for training/evaluating Portuguese legal & tax retrieval. Queries are controlled-diversity synthetic questions; positive passages are Brazilian federal law (public domain). Built for the underserved Brazilian legal-IR setting, and decontaminated against MTEB(por) so it is safe to train on and evaluate on common PT benchmarks. How it was made Positives: real… See the full description on the dataset page: https://huggingface.co/datasets/tardellirs/brazilian-legal-tax-qa-synthetic.texttext-retrieval1K<n<10K2 likes67 downloads4mo agoHugging Face22HumynLabs /Brazilian_Road_Signs_Dataset Brazilian Road Signs Dataset This dataset contains high-quality images of Brazilian road and traffic signs collected from various urban and rural environments. It supports AI research in computer vision, object detection, and autonomous driving systems adapted to Brazil’s signage standards and language. Contact For queries or collaborations related to this dataset, contact: anoushka@kgen.io abhishek.vadapalli@kgen.io Supported Tasks Task Categories:… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/Brazilian_Road_Signs_Dataset.imageimage-classificationn<1K0 likes62 downloads1y agoHugging Face23lafbarbosa /sugarcane_dataset_northeast_sao_paulo_state_brazilgatedThis is the first publicly available dataset that integrates sugarcane crop yield, production environment, meteorological records, and Sentinel-2 satellite imagery from commercial fields in the northeast of São Paulo State, Brazil. The description of this dataset is in the scientific paper https://www.sciencedirect.com/science/article/pii/S2352340926001022 Please acknowledge the use of this dataset by citing the paper above and the following reference: @MISC{Barbosa2025-uq, title =… See the full description on the dataset page: https://huggingface.co/datasets/lafbarbosa/sugarcane_dataset_northeast_sao_paulo_state_brazil.image1K<n<10K1 likes59 downloads15d agoHugging Face24emailmarketingdataset /brazil-business-dataset Brazil Business Email List — Market Intelligence Dataset 2,800,000 verified Brazil contacts are available from LeadsBlue →. This open dataset provides the aggregate market intelligence behind that database — contact volume, benchmark open/reply rates, send timing, and compliance for the Brazil segment. At a glance: Market-intelligence reference data for the Brazil Business Email List audience — population scale, decision-maker structure, channel benchmarks, and compliance… See the full description on the dataset page: https://huggingface.co/datasets/emailmarketingdataset/brazil-business-dataset.tabular-classificationn<1K0 likes54 downloads4mo agoHugging Face25Etore-BeS /brazilian-digital-identity-instagram-corpus Brazilian Digital Identity on Instagram — Thematic Corpus Authors: Étore Braga e Santos (Unicamp, ORCID 0009-0000-3502-705X) & Pâmella Fernandes de Sá (USP FEA-RP)License: CC BY 4.0 (dataset) — code at https://github.com/Etore-BeS/brazilian-digital-identity-paper (MIT)Version: 1.0.0 — Collection March 7–11, 2026 Dataset Description Anonymized corpus of 17,287 Instagram comments on three international celebrity posts with substantial Brazilian engagement, annotated… See the full description on the dataset page: https://huggingface.co/datasets/Etore-BeS/brazilian-digital-identity-instagram-corpus.tabulartext-classification10K<n<100K0 likes54 downloads6d agoHugging Face26emdemor /news-of-the-brazilian-newspaper News of the Brazilian Newspaper This repository contains a comprehensive dataset of news articles from a Brazilian newspaper, Folha de São Paulo (http://www.folha.uol.com.br/). The dataset includes 167,053 examples of news articles, comprising headlines, URLs of articles, complete articles, and their respective categories. Dataset Creation The headlines were initially gathered from Inshorts and were then used to scrape the complete news articles from Folha de São Paulo.… See the full description on the dataset page: https://huggingface.co/datasets/emdemor/news-of-the-brazilian-newspaper.texttext-classification100K<n<1M2 likes47 downloads2y agoHugging Face27juliasdata /medical-audio-sample-brazilian-portuguese Julia's Data: Brazilian Portuguese Medical Audio Sample Public sample of a Brazilian Portuguese medical audio dataset built for ASR, TTS, and conversational AI evaluation. This repository contains deidentified clinical source material transformed into five spoken content types and recorded by a human speaker. This sample includes 1 record, 20 aligned audio segments, 1 speaker, and about 5.26 minutes of audio. Full dataset and commercial licensing: juliasdata.com Commercial overview:… See the full description on the dataset page: https://huggingface.co/datasets/juliasdata/medical-audio-sample-brazilian-portuguese.audioautomatic-speech-recognitionn<1K1 likes46 downloads6mo agoHugging Face28gtfintechlab /central_bank_of_brazil Dataset Summary For dataset summary, please refer to https://huggingface.co/datasets/gtfintechlab/central_bank_of_brazil Additional Information This dataset is annotated across three different tasks: Stance Detection, Temporal Classification, and Uncertainty Estimation. The tasks have four, two, and two unique labels, respectively. This dataset contains 1,000 sentences taken from the meeting minutes of the Central Bank of Brazil. Label Interpretation… See the full description on the dataset page: https://huggingface.co/datasets/gtfintechlab/central_bank_of_brazil.tabulartext-classification1K<n<10K1 likes44 downloads1y agoHugging Face29itsalissonsilva /brazil-deputy-expensestabular100K<n<1M0 likes43 downloads1y agoHugging Face30joaosanches /brazilian_european_portuguese_datasettext100K<n<1M2 likes42 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.