datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
weather-rescue-brazil
Weather Rescue Brazil
Daily meteorological observations from Brazilian observatories, 1883–1890,
transcribed from printed nineteenth-century tables by an open 2B model running
offline, with every row carrying its provenance and a quality verdict.
This is an independent project. It is not affiliated with Zooniverse or with
the Weather Rescue / Rainfall Rescue projects, whose naming family it
gratefully follows.
Citing it
DOI: 10.5281/zenodo.22876699 — the concept… See the full description on the dataset page: https://huggingface.co/datasets/akagabi/weather-rescue-brazil.ENADE_Brazilian_national_university_examination_MCQ_483brazilian_court_decisions
Dataset Card for predicting-brazilian-court-decisions
Dataset Summary
The dataset is a collection of 4043 Ementa (summary) court decisions and their metadata from
the Tribunal de Justiça de Alagoas (TJAL, the State Supreme Court of Alagoas (Brazil). The court decisions are labeled
according to 7 categories and whether the decisions were unanimous on the part of the judges or not. The dataset
supports the task of Legal Judgment Prediction.
Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/joelniklaus/brazilian_court_decisions.openai-terra-batch-wiki-brazil-1000-partial-20260724-01
OpenAI Terra Batch — Wikipédia PT-BR (run parcial)
Checkpoint publicável de uma execução real e interrompida do fluxo
document_task_matrix. A execução planejou gerar uma matriz de 1.000
documentos da Wikipédia em português por 25 tasks canônicas usando a Responses
API Batch e o modelo gpt-5.6-terra.
Este repositório não representa a conclusão dos 25.000 pares planejados. Ele
contém somente os 1.282 candidatos aceitos após a reconciliação offline de
todos os resultados Batch já… See the full description on the dataset page: https://huggingface.co/datasets/costadev00/openai-terra-batch-wiki-brazil-1000-partial-20260724-01.brazilian-customer-service-conversations
Brazilian Customer Service Conversations
Dataset de conversas de atendimento ao cliente em portugues brasileiro (PT-BR).
De um like me apoie em manter esse dataset!
Descricao
Conversas sinteticas de alta qualidade simulando interacoes reais entre clientes e atendentes em diversos setores da economia brasileira. Util para treinar e avaliar modelos de:
Chatbots de atendimento
Classificacao de intencao (intent classification)
Analise de sentimento em conversas
Geracao de… See the full description on the dataset page: https://huggingface.co/datasets/RichardSakaguchiMS/brazilian-customer-service-conversations.brazilian-federal-law-corpus
Brazilian Federal Law Corpus (PT-BR)
41,526 passages of Brazilian federal legislation (laws and decrees — articles, titles, ementas) scraped from
planalto.gov.br. Official Brazilian legal texts are public domain (Lei 9.610/1998, art. 8). A clean corpus for
Portuguese legal IR / RAG / language modeling, with no LLM-generated content — pure law text — and
decontaminated against MTEB(por).
Companion to the synthetic QA pairs in
tardellirs/brazilian-legal-tax-qa-synthetic.… See the full description on the dataset page: https://huggingface.co/datasets/tardellirs/brazilian-federal-law-corpus.brazilian-legal-tax-qa-synthetic
Brazilian Legal & Tax QA — synthetic retrieval pairs (PT-BR)
4,278 (query, positive-passage) pairs for training/evaluating Portuguese legal & tax retrieval. Queries are
controlled-diversity synthetic questions; positive passages are Brazilian federal law (public domain). Built for
the underserved Brazilian legal-IR setting, and decontaminated against MTEB(por) so it is safe to train on and
evaluate on common PT benchmarks.
How it was made
Positives: real… See the full description on the dataset page: https://huggingface.co/datasets/tardellirs/brazilian-legal-tax-qa-synthetic.adaption-brazil-crypto-regulatory-qa
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-brazil_crypto_regulatory_qa
This dataset contains prompt-completion pairs focused on Brazilian financial regulations regarding cryptocurrencies, tokens, and digital assets. The content specifically addresses the role of the CVM (Comissão de Valores Mobiliários), risk assessments for investors, and compliance with laws such as Lei 14.478/2022. Each entry provides structured JSON… See the full description on the dataset page: https://huggingface.co/datasets/Fernandosr85/adaption-brazil-crypto-regulatory-qa.Brazilian_CLT_preferencesDataset DescriptionThis dataset contains 736 validated human-preference entries designed to align language models with expert expectations for answering questions about Brazil’s Consolidation of Labor Laws (CLT). It was created to support Direct Preference Optimization (DPO) fine-tuning and evaluation of LLM-based legal assistants.
Intended Use
Primary Purpose: Training and evaluating models for legal question answering under the Brazilian CLT framework.
Target Users: Researchers… See the full description on the dataset page: https://huggingface.co/datasets/ai-eldorado/Brazilian_CLT_preferences.adaption-brazil-agri-qa-match
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-brazil_agri_qa_match
This dataset contains pairs of agricultural questions posed to the Brazilian Ministry of Agriculture and candidate answers sourced from government portals. Each sample includes a binary label indicating whether the provided answer correctly addresses the specific question asked. The content covers diverse topics such as family farming, traceability, ministerial… See the full description on the dataset page: https://huggingface.co/datasets/Fernandosr85/adaption-brazil-agri-qa-match.brazilian_court_decisions
Dataset Card for predicting-brazilian-court-decisions
Dataset Summary
The dataset is a collection of 4043 Ementa (summary) court decisions and their metadata from
the Tribunal de Justiça de Alagoas (TJAL, the State Supreme Court of Alagoas (Brazil). The court decisions are labeled
according to 7 categories and whether the decisions were unanimous on the part of the judges or not. The dataset
supports the task of Legal Judgment Prediction.
Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/charannatra/brazilian_court_decisions.adaption-brazilian-crypto-compliance
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-brazilian_crypto_compliance
This dataset contains structured compliance assessments for Brazilian crypto asset regulations, featuring prompts with specific scenarios and JSON completions detailing applicable laws, risk levels, and corrective actions. Each entry evaluates adherence to rules from authorities like the BCB, CVM, and COAF regarding issues such as asset segregation, KYC… See the full description on the dataset page: https://huggingface.co/datasets/Fernandosr85/adaption-brazilian-crypto-compliance.brazilian-math-physics-qa
Brazilian Math & Physics QA
English | Português do Brasil
English
Summary
Brazilian Portuguese question-answer pairs covering mathematics, physics, chemistry, and related educational subjects. Each record contains a user question and an assistant answer in chat/SFT format.
Examples: 19.082
Train: 18.148
Validation: 934
Language: Brazilian Portuguese (pt-BR)
Schema
{"id":"qa_...","subject":"fisica","category":"mecanica-geral"… See the full description on the dataset page: https://huggingface.co/datasets/artificialguybr/brazilian-math-physics-qa.wiki-brazil-instruct
wiki-brazil
Synthetic SFT dataset generated by sft-dataset-creator.
Examples: 20378
Generator: google/gemma-4-31B-it-qat-w4a16-ct
Evaluator: deterministic only
Configuration hash: 529d0d158caa1d5d3891f419c36eb3bce99159387592838295bf5527ada106d1
Review provenance, licenses, report.json, and the resolved configuration before public release.
brazilian-gov-formal-letters
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
brazilian_gov_formal_letters
This dataset contains pairs of prompts and completions for generating formal administrative communications in Portuguese directed at Brazilian federal institutions. The content includes complaints, suggestions, and official responses addressing issues such as missed deadlines, lack of transparency, and unresolved demands. Each sample demonstrates a formal, legalistic… See the full description on the dataset page: https://huggingface.co/datasets/Fernandosr85/brazilian-gov-formal-letters.adaption-brazilian-regulatory-filings
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-brazilian_regulatory_filings
This dataset contains samples of Brazilian regulatory filings (Fatos Relevantes and Market Notices) from publicly traded companies, presented in both Portuguese and English. Each sample includes a classification task where the text is analyzed to determine the event type, status, scope relative to a target entity, and market signal sentiment. The content… See the full description on the dataset page: https://huggingface.co/datasets/Fernandosr85/adaption-brazilian-regulatory-filings.brazilian-math-physics-qa-vision
Brazilian Math & Physics QA — Image Dependent
English | Português do Brasil
English
Summary
Brazilian Portuguese educational question-answer pairs whose problem statement or solution depends on one or more images.
Examples: 3,808
Referenced image URLs: 5,094 unique
Language: Brazilian Portuguese (pt-BR)
Schema
{"id":"vqa_...","subject":"matematica","category":"geometria","title":"...","messages":[{"role":"user","content":"...… See the full description on the dataset page: https://huggingface.co/datasets/artificialguybr/brazilian-math-physics-qa-vision.adaption-brazilian-civil-law-rulings
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
brazilian_civil_law_rulings
This dataset contains samples of Brazilian civil law court decisions, specifically focusing on appeals, special recourse agravo, and consumer protection cases handled by the Superior Court of Justice (STJ). The text includes detailed legal reasoning, case summaries, discussion of res judicata, contractual rescission, moral damages, and citations of relevant articles… See the full description on the dataset page: https://huggingface.co/datasets/Fernandosr85/adaption-brazilian-civil-law-rulings.adaption-brazil-health-guidance-pt
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-brazil_health_guidance_pt
This dataset contains pairs of citizen inquiries and institutional responses regarding Brazilian health insurance coverage and SUS access in Portuguese. Each entry follows a strict regulatory guide, providing formal, non-diagnostic administrative orientation on topics like the ANS Rol de Procedimentos, pre-existing conditions, and patient rights. The… See the full description on the dataset page: https://huggingface.co/datasets/Fernandosr85/adaption-brazil-health-guidance-pt.brazilian-passports
Disclaimer: All passport images and associated data in this dataset are synthetically generated and do not correspond to real individuals. Any names, numbers, or personal details are fictional and used solely for research and development purposes.
Introduction - Brazil
The Synthetic Brazil Passports Dataset brings together over 1,000 AI-generated passport images designed for training OCR and computer vision systems on identity documents. Every record is fully synthetic, so the… See the full description on the dataset page: https://huggingface.co/datasets/ud-synthetic/brazilian-passports.ENEM_Brazil_National_High_School_Exam_2022_to_2024brazilian-court-summaries
This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform.
brazilian_court_summaries
This dataset contains pairs of Brazilian court decision summaries in Portuguese, where the prompt includes the full text with the final ruling and the completion provides the text truncated just before the verdict. The samples cover various civil and criminal cases, such as appeals, habeas corpus, and indemnification actions, focusing on legal… See the full description on the dataset page: https://huggingface.co/datasets/Fernandosr85/brazilian-court-summaries.brazilian-gov-service-pairs
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
brazilian_gov_service_pairs
This dataset contains pairs of Brazilian government service categories and their corresponding citizen interaction types, such as complaints, requests, or communications. The prompts cover diverse sectors including transportation, telecommunications, education, and federal revenue. Each completion classifies the nature of the citizen's engagement with the specific… See the full description on the dataset page: https://huggingface.co/datasets/Fernandosr85/brazilian-gov-service-pairs.wiki-brazil-smoke-10pct-sft-d918a84e
wiki-brazil-smoke-10pct
Synthetic SFT dataset generated by sft-dataset-creator.
Examples: 60
Generator: google/gemma-4-26B-A4B-it
Evaluator: deterministic only
Configuration hash: 408df947799109f6826c54ae0cbc288bde514b695a9a9fe2c24f9ffd3053a640
Review provenance, licenses, report.json, and the resolved configuration before public release.
wiki-brazil
Wiki Brazil
Dataset criado a partir da API da Wikipédia em português.
Conteúdo
O dataset contém o artigo Brasil e os artigos em namespace principal
listados nos links internos desse artigo.
Arquivos
wiki_brazil.jsonl: registros em JSONL.
Schema
title: título resolvido do artigo na Wikipédia.
page_id_: identificador numérico da página na Wikipédia.
text: texto completo extraído em plaintext pela API MediaWiki.
Fonte… See the full description on the dataset page: https://huggingface.co/datasets/costadev00/wiki-brazil.wiki-brazil-instruct
wiki-brazil
Synthetic SFT dataset generated by sft-dataset-creator.
Examples: 20378
Generator: google/gemma-4-31B-it-qat-w4a16-ct
Evaluator: deterministic only
Configuration hash: 529d0d158caa1d5d3891f419c36eb3bce99159387592838295bf5527ada106d1
Review provenance, licenses, report.json, and the resolved configuration before public release.
wiki-brazil-smoke-20
wiki-brazil
Synthetic SFT dataset generated by sft-dataset-creator.
Examples: 18
Generator: google/gemma-4-31B-it-qat-w4a16-ct
Evaluator: deterministic only
Configuration hash: 83db6bf999d8eb49a516d262a38539a440ca941db9a36331dda906b101eac4b6
Review provenance, licenses, report.json, and the resolved configuration before public release.
wiki-brazil
Wiki Brazil
Dataset criado a partir da API da Wikipédia em português.
Conteúdo
O dataset contém o artigo Brasil e os artigos em namespace principal
listados nos links internos desse artigo.
Arquivos
wiki_brazil.jsonl: registros em JSONL.
Schema
title: título resolvido do artigo na Wikipédia.
page_id_: identificador numérico da página na Wikipédia.
text: texto completo extraído em plaintext pela API MediaWiki.
Fonte… See the full description on the dataset page: https://huggingface.co/datasets/br-llm-data/wiki-brazil.
