datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
synthetic-printed-brazilian-passports
Brazilian passport dataset
The dataset comprises 5,000 high-resolution synthetic photos of Brazilian passports, designed to advance computer vision and identity verification systems. It provides a secure and ethical resource for training robust models for OCR (Optical Character Recognition), document analysis, and spoofing detection, all without exposing real personal data or sensitive personal information.
By utilizing this dataset, researchers and developers can enhance… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/synthetic-printed-brazilian-passports.lpr-brazil-700ksolar-plants-brazil
🛰️ Solar Plants Brazil
Solar Plants Brazil is a geospatial dataset for binary semantic segmentation of photovoltaic (PV) solar power stations in satellite imagery. It consists of multi-spectral image tiles (including near-infrared) with pixel-level annotations indicating the presence of solar panels. This dataset enables training and evaluating deep learning models that automatically detect solar farm installations from overhead imagery, supporting applications in renewable energy… See the full description on the dataset page: https://huggingface.co/datasets/FederCO23/solar-plants-brazil.Brazilian_Bills_and_Invoices_Dataset
Brazilian Bills and Invoices Dataset
This dataset contains high-quality scanned and photographed images of Brazilian bills, invoices, and utility payment documents. It supports AI research in OCR, financial document understanding, and structured data extraction for Portuguese-language financial contexts.
Contact
For queries or collaborations related to this dataset, contact:
anoushka@kgen.io
abhishek.vadapalli@kgen.io
Supported Tasks
Task Categories:… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/Brazilian_Bills_and_Invoices_Dataset.bias-within-borders-brazil-images-ptdou-brazil-dataset
Dataset Card for Dataset Diário Oficial da União (DOU)
The Diário Oficial da União (DOU) is the official government gazette of Brazil, published by the National Press. It serves as the primary means of communication for federal government acts, including laws, decrees, ordinances, public notices, and other official decisions. The DOU ensures transparency and legal validity for government actions and is divided into three sections:
Section 1: Publishes laws, decrees, and… See the full description on the dataset page: https://huggingface.co/datasets/gerson-vfs/dou-brazil-dataset.Brazilian_Road_Signs_Dataset
Brazilian Road Signs Dataset
This dataset contains high-quality images of Brazilian road and traffic signs collected from various urban and rural environments. It supports AI research in computer vision, object detection, and autonomous driving systems adapted to Brazil’s signage standards and language.
Contact
For queries or collaborations related to this dataset, contact:
anoushka@kgen.io
abhishek.vadapalli@kgen.io
Supported Tasks
Task Categories:… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/Brazilian_Road_Signs_Dataset.sugarcane_dataset_northeast_sao_paulo_state_brazilThis is the first publicly available dataset that integrates sugarcane crop yield, production environment, meteorological records, and Sentinel-2 satellite imagery from
commercial fields in the northeast of São Paulo State, Brazil.
The description of this dataset is in the scientific paper https://www.sciencedirect.com/science/article/pii/S2352340926001022
Please acknowledge the use of this dataset by citing the paper above and the following reference:
@MISC{Barbosa2025-uq,
title =… See the full description on the dataset page: https://huggingface.co/datasets/lafbarbosa/sugarcane_dataset_northeast_sao_paulo_state_brazil.Brazilian_Coffee_Scenes
Dataset Card for "Brazilian_Coffee_Scenes"
Licensing Information
[CC BY-NC]
Citation Information
Do deep features generalize from everyday objects to remote sensing and aerial scenes domains?
@inproceedings{penatti2015deep,
title = {Do deep features generalize from everyday objects to remote sensing and aerial scenes domains?},
author = {Penatti, Ot{\'a}vio AB and Nogueira, Keiller and Dos Santos, Jefersson A},
year = 2015… See the full description on the dataset page: https://huggingface.co/datasets/jonathan-roberts1/Brazilian_Coffee_Scenes.Brazilian-Sign-Language-Alphabet-DatasetBrazilianPortugueseOCRThis set was created and is expanding. The objective is for students to be able to use and learn about OCR.
Examples:
Llama Vision - Optical Character Recognition (OCR)
ExtractThinker using Ollama
ExtractThinker
Llama Vision
Dataset Description
This dataset contains images of three types of documents commonly used in Brazil: Electronic Invoices (NF-e), Medical Leaflets, and Telephone Directories. These documents are essential for various economic, medical, and communication purposes. The… See the full description on the dataset page: https://huggingface.co/datasets/sc0v0ne/BrazilianPortugueseOCR.brazilian-cultural-video-dataset
Bamboo Data Brazilian Cultural Video Dataset (Sample)
⚠️ License Notice: Evaluation Only
This is a sample of the Bamboo Data brazilian cultural video dataset, provided for internal evaluation purposes ONLY. The use of this data is strictly limited by the license defined below.
Any use for training, fine-tuning, or inference of AI/ML models, or any commercial activity, is strictly prohibited with this sample.
Dataset Description
The Bamboo Data… See the full description on the dataset page: https://huggingface.co/datasets/bamboodata/brazilian-cultural-video-dataset.soybean_leaf_disease_classification_brazil
Soybean Leaf Disease Classification Brazil
A dataset for disease classification of soybean leaves. The dataset contains 6,410 images across 3 classes: Caterpillar, Diabrotica speciosa, Healthy.Images per class:
Caterpillar: 3,309
Diabrotica speciosa: 2,205
Healthy: 896
This dataset is indexed on https://project-agml.github.io/ as part of the AgML python library.
Citation
@article{mignoni2022soybean,
title={Soybean images dataset for caterpillar and Diabrotica… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/soybean_leaf_disease_classification_brazil.soybean_weed_uav_brazil
Soybean Weed Uav Brazil
A dataset for image classification of Soybean Weed Uav Brazil. The dataset contains 15,336 images across 4 classes: broadleaf, grass, soil, soybean.Images per class:
broadleaf: 1,191
grass: 3,520
soil: 3,249
soybean: 7,376
This dataset is indexed on https://project-agml.github.io/ as part of the AgML python library.
Citation
dos Santos Ferreira, Alessandro; Pistori, Hemerson; Matte Freitas, Daniel; Gonçalves da Silva, Gercina (2017)… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/soybean_weed_uav_brazil.apple_detection_drone_brazil
Apple Detection Drone Brazil
A dataset for object detection of apples. The dataset contains 689 images with 2,471 bounding box annotations across 1 category.
This dataset is indexed on https://project-agml.github.io/ as part of the AgML python library.
Citation
@article{DBLP:journals/corr/abs-2110-12331,
author={Thiago T. Santos and Luciano Gebler},
title={A methodology for detection and localization of fruits in apples orchards from aerial images},
journal={CoRR}… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/apple_detection_drone_brazil.Brazilian_Cerrado-Savanna_Scenes
Dataset Card for "Brazilian_Cerrado-Savanna_Scenes"
Licensing Information
[CC BY-NC]
Citation Information
Towards vegetation species discrimination by using data-driven descriptors
@inproceedings{nogueira2016towards,
title = {Towards vegetation species discrimination by using data-driven descriptors},
author = {Nogueira, Keiller and Dos Santos, Jefersson A and Fornazari, Tamires and Silva, Thiago Sanna Freire and Morellato, Leonor Patricia and… See the full description on the dataset page: https://huggingface.co/datasets/jonathan-roberts1/Brazilian_Cerrado-Savanna_Scenes.Brazil-Medicine-Schools-Entrance-Exams-FAMERP-SANTA-CASAbrazilian-math-physics-qa-vision
Brazilian Math & Physics QA — Image Dependent
English | Português do Brasil
English
Summary
Brazilian Portuguese educational question-answer pairs whose problem statement or solution depends on one or more images.
Examples: 3,808
Referenced image URLs: 5,094 unique
Language: Brazilian Portuguese (pt-BR)
Schema
{"id":"vqa_...","subject":"matematica","category":"geometria","title":"...","messages":[{"role":"user","content":"...… See the full description on the dataset page: https://huggingface.co/datasets/artificialguybr/brazilian-math-physics-qa-vision.Brazilian_Item_Price_and_Description_Dataset
Brazilian Item Price and Description Dataset
This dataset contains high-resolution images and structured text data of product price tags and item descriptions collected from Brazilian retail stores and e-commerce platforms. It enables AI research in OCR, product recognition, and retail analytics for the Portuguese-speaking market.
Contact
For queries or collaborations related to this dataset, contact:
anoushka@kgen.io
abhishek.vadapalli@kgen.io
Supported… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/Brazilian_Item_Price_and_Description_Dataset.Mr.Porter.Product.prices.Brazil
Mr Porter web scraped data
About the website
The E-commerce industry in the Americas, particularly in Brazil, has observed a surge in its growth trajectory. The nation is a thriving hub for online businesses, with a notable player in this sector being Mr Porter, which operates prominently in the online luxury fashion retail segment. It has been observed that the dataset includes Ecommerce product-list page (PLP) data on Mr Porter in Brazil. This comprehensive data… See the full description on the dataset page: https://huggingface.co/datasets/DBQ/Mr.Porter.Product.prices.Brazil.Brazilian_Brand_Marketing_Banners_Dataset
Brazilian Brand Marketing Banners Dataset
This dataset contains high-quality images of Brazilian brand marketing banners collected from online and offline retail environments. It includes product advertisements, promotional offers, and digital marketing visuals designed for both Portuguese-speaking audiences and bilingual markets.
Contact
For queries or collaborations related to this dataset, contact:
anoushka@kgen.io
abhishek.vadapalli@kgen.io
Supported… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/Brazilian_Brand_Marketing_Banners_Dataset.synthetic-printed-brazilian-passports
Synthetic Passports Dataset - 5 000 passport photos
Dataset provides 5,000 files with high-resolution synthetic passport images with diverse angles, lighting, and backgrounds, designed for training OCR, computer vision, and identity verification models without exposing real personal data or sensitive information. - Get the data
Dataset characteristics:
Characteristic
Data
Description
Printed synthetic passport images for training ML models in PII extraction… See the full description on the dataset page: https://huggingface.co/datasets/ud-biometrics/synthetic-printed-brazilian-passports.bias-within-borders-brazil-imagesBrazilian_Dish_Photos_Dataset
