datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
claude-code-traces-pt-brThis dataset was generated using teich by TeichAI
claude Agent Traces
This directory contains raw agent trace files generated by teich.
JSONL files: 20
Training-ready tools
Generated agent traces carry configured or recovered tool schemas so tools remain available for training even when a session did not call them.
Native Claude Code imports recover schemas for Claude Code and Claude Desktop built-ins, plus conservative name-derived MCP schemas, when the raw… See the full description on the dataset page: https://huggingface.co/datasets/DedeProGames/claude-code-traces-pt-br.geo-ptbr
GEO-PTBR
The first Brazilian-Portuguese benchmark for Generative Engine Optimization
(GEO): 525 PT-BR queries with 2,625 source documents, plus
per-technique citation-visibility results measured on three generative
engines.
Companion artifact to the paper GEO-PTBR: A Brazilian-Portuguese Replication
of Generative Engine Optimization and the Engine-Dependence of Its Effects.
Experiment version: 0.3.0.
What this measures
Given a query and five already-retrieved… See the full description on the dataset page: https://huggingface.co/datasets/epicchi2103/geo-ptbr.wikipedia-pt-br-instruct-5k
Wikipedia PT-BR Instruct
wikipedia-pt-br-instruct is a synthetic supervised fine-tuning (SFT)
dataset in Brazilian Portuguese generated from Wikipedia-derived documents.
This release is an intermediate evaluation dataset produced with the
sft-dataset-creator pipeline from the run
wiki-ptbr-extract-calib-5kdocs-14tasks. It was generated from a fixed
revision of costadev00/wikipedia-pt-br-extract:
cdbd07dc4a3de6e64632c718710b3ae0ebaeb0ff
The dataset is intended for intermediate… See the full description on the dataset page: https://huggingface.co/datasets/costadev00/wikipedia-pt-br-instruct-5k.physics-ptbr
Tradução do Camel Pyysics dataset para Portuguese (PT-BR) usando NLLB 3.3b.
CAMEL: Communicative Agents for “Mind” Exploration of Large Scale Language Model Society
Github: https://github.com/lightaime/camel
Website: https://www.camel-ai.org/
Arxiv Paper: https://arxiv.org/abs/2303.17760
Dataset Summary
Physics dataset is composed of 20K problem-solution pairs obtained using gpt-4. The dataset problem-solutions pairs generating from 25 physics topics, 25… See the full description on the dataset page: https://huggingface.co/datasets/botbotrobotics/physics-ptbr.nsfw-pt_br
Dataset Card for The Pile
Dataset Summary
The NSFW is a 230K diverse, filtred and cleaned text from adult websites, high-quality
datasets combined together.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
This dataset is in Portuguese Brazil (pt_BR)
Dataset Structure
Data Instances
Data Fields
all
text (str): Text.
Dataset Creation
Curation Rationale
[More… See the full description on the dataset page: https://huggingface.co/datasets/MrAiran/nsfw-pt_br.tool-calling-traces-ptbr
Tool calling conversations in Portuguese
484 synthetic conversations that teach a model when to call a tool, which one to call and
with which arguments, and also when to answer directly, with no tool at all.
Each line of the file is a complete conversation: the user's question, the tool call, the
simulated return of that tool, and the final answer.
It was built because no dataset of tool calling in Portuguese with fictional tools existed.
The 30 tools and the user questions were… See the full description on the dataset page: https://huggingface.co/datasets/annajuliaasf/tool-calling-traces-ptbr.noticias-govbr-ptbr-1chemistry-ptbr
Tradução do Camel Chemisty dataset para Portuguese (PT-BR) usando NLLB 3.3b.
CAMEL: Communicative Agents for “Mind” Exploration of Large Scale Language Model Society
Github: https://github.com/lightaime/camel
Website: https://www.camel-ai.org/
Arxiv Paper: https://arxiv.org/abs/2303.17760
Dataset Summary
Chemistry dataset is composed of 20K problem-solution pairs obtained using gpt-4. The dataset problem-solutions pairs generating from 25 chemistry topics… See the full description on the dataset page: https://huggingface.co/datasets/botbotrobotics/chemistry-ptbr.FairytaleQA-translated-ptBR
Dataset Card for FairytaleQA-translated-ptBR
Dataset Summary
This repository contains the Brazilian Portuguese (pt-BR) machine-translated version of the original English FairytaleQA dataset (https://huggingface.co/datasets/WorkInTheDark/FairytaleQA). FairytaleQA is an open-source dataset designed to enhance comprehension of narratives, aimed at students from kindergarten to eighth grade. The dataset is meticulously annotated by education experts following an… See the full description on the dataset page: https://huggingface.co/datasets/benjleite/FairytaleQA-translated-ptBR.biology-ptbr
Tradução do Camel Biology dataset para Portuguese (PT-BR) usando NLLB 3.3b.
CAMEL: Communicative Agents for “Mind” Exploration of Large Scale Language Model Society
Github: https://github.com/lightaime/camel
Website: https://www.camel-ai.org/
Arxiv Paper: https://arxiv.org/abs/2303.17760
Dataset Summary
Biology dataset is composed of 20K problem-solution pairs obtained using gpt-4. The dataset problem-solutions pairs generating from 25 biology topics, 25… See the full description on the dataset page: https://huggingface.co/datasets/botbotrobotics/biology-ptbr.Lembron-Reason-24K-PTBR
Lembron-Reason-24K-PTBR
Lebron IA 24K PT-BR Think Distilled SFT é um dataset conversacional em português brasileiro, criado para Supervised Fine-Tuning (SFT), instruction tuning e treinamento de modelos focados em respostas detalhadas, raciocínio estruturado, comportamento de assistente e geração de conteúdo em PT-BR.
Este corpus foi estruturado como um dataset distilado do comportamento da Lebron IA, com foco em transferir para modelos menores ou médios um padrão de resposta… See the full description on the dataset page: https://huggingface.co/datasets/guell00/Lembron-Reason-24K-PTBR.gemini_orpo_dpo_ptbraurora-dataset-roleplay-ptbr
Aurora Dataset Roleplay 🌌 (PT-BR)
O que é É um dataset que contém mais de 2 mil diálogos em português do Brasil. Ainda que tenha sido gerado sinteticamente, foi utilizado apenas modelos SOTA, então os diálogos são muito próximos da naturalidade e espontaneidade de um ser humano. Foi feito pensando em roleplay, por isso os diálogos contém nuances psicológicas, cenários diversos e personagens com estilos de falas e motivações complexas.
Modos de Geração… See the full description on the dataset page: https://huggingface.co/datasets/wilsondesouza/aurora-dataset-roleplay-ptbr.wikipedia-pt-br-instructions-sft
wikipedia-pt-br-instructions-gemma-alpaca
Dataset de Instruction Following em formato Alpaca puro, com exatamente os campos instruction, input e output.
Origem
Derivado do dataset local instruction_following produzido por wiki-if-builder, por sua vez derivado de costadev00/wikipedia-pt-br-extract.
Campos
instruction: comando em português brasileiro.
input: contexto mínimo opcional.
output: resposta esperada.
Licença e limitações
A licença herdada é… See the full description on the dataset page: https://huggingface.co/datasets/costadev00/wikipedia-pt-br-instructions-sft.wikitext-ptbr-1pt-br-agentic-text-to-sql-distilled-trajectories
PT-BR Agentic Text-to-SQL Distilled Trajectories
This dataset contains message-only distilled trajectories for training tool-using Text-to-SQL agents in Brazilian Portuguese. The trajectories were selected from LLM-judged correct conversations and preserve the agent protocol used in the released code.
Code and reproducibility repository:
https://github.com/Boakpe/distilled-slms-for-text-to-sql-pt-br
Related collection:… See the full description on the dataset page: https://huggingface.co/datasets/Boakpe/pt-br-agentic-text-to-sql-distilled-trajectories.wikipedia-pt-br-instruct-5k
Wikipedia PT-BR Instruct
wikipedia-pt-br-instruct is a synthetic supervised fine-tuning (SFT)
dataset in Brazilian Portuguese generated from Wikipedia-derived documents.
This release is an intermediate evaluation dataset produced with the
sft-dataset-creator pipeline from the run
wiki-ptbr-extract-calib-5kdocs-14tasks. It was generated from a fixed
revision of costadev00/wikipedia-pt-br-extract:
cdbd07dc4a3de6e64632c718710b3ae0ebaeb0ff
The dataset is intended for intermediate… See the full description on the dataset page: https://huggingface.co/datasets/br-llm-data/wikipedia-pt-br-instruct-5k.ptbr-human-preferences
🇧🇷 HUBX Human Preference Dataset (PT-BR)
The largest Portuguese-Brazilian human preference dataset for RLHF/DPO training.
📊 Dataset Statistics
Metric
Value
Total Annotations
314,757
Unique Tasks
450
Human Annotators
~600
Avg. Votes per Task
~699
Language
Portuguese (Brazil)
Domain
Communication Quality & Tone
🎯 Why This Dataset?
🇧🇷 Native PT-BR: Collected from Brazilian Portuguese speakers - not translated
👥 Real Humans:… See the full description on the dataset page: https://huggingface.co/datasets/Hub-Ai/ptbr-human-preferences.wikipedia-pt-br-instructions
wikipedia-pt-br-instructions-gemma
Dataset sintético de Instruction Following em português brasileiro, derivado de artigos da Wikipedia pt-BR.
Origem
Os exemplos derivam de costadev00/wikipedia-pt-br-extract e preservam source_page_id, source_title, source_dataset e a licença herdada cc-by-sa-3.0.
Processo
A geração é inspirada em Alpaca e Self-Instruct: cada artigo válido passa por uma chamada de analista documental que produz candidatos de… See the full description on the dataset page: https://huggingface.co/datasets/br-llm-data/wikipedia-pt-br-instructions.wikipedia-pt-br-instructions
wikipedia-pt-br-instructions-gemma
Dataset sintético de Instruction Following em português brasileiro, derivado de artigos da Wikipedia pt-BR.
Origem
Os exemplos derivam de costadev00/wikipedia-pt-br-extract e preservam source_page_id, source_title, source_dataset e a licença herdada cc-by-sa-3.0.
Processo
A geração é inspirada em Alpaca e Self-Instruct: cada artigo válido passa por uma chamada de analista documental que produz candidatos de instrução ancorados… See the full description on the dataset page: https://huggingface.co/datasets/costadev00/wikipedia-pt-br-instructions.
