datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
corpus-ptbr-v1
🇧🇷 Corpus PT-BR v1
Um corpus em Português Brasileiro voltado para pré-treinamento e fine-tuning de LLMs. Combina dados reais curados com uma camada sintética construída para ampliar a diversidade estilística, lexical e discursiva em português.
🔄 Visão Geral do Pipeline
📊 Estatísticas
Métrica
Valor
Total de documentos
8,399,857
Total de palavras
~4.84B
Tokens estimados
~6.29B
Tamanho (Parquet)
~17.9 GB
Idioma
Português Brasileiro… See the full description on the dataset page: https://huggingface.co/datasets/Madras1/corpus-ptbr-v1.claude-code-traces-pt-brThis dataset was generated using teich by TeichAI
claude Agent Traces
This directory contains raw agent trace files generated by teich.
JSONL files: 20
Training-ready tools
Generated agent traces carry configured or recovered tool schemas so tools remain available for training even when a session did not call them.
Native Claude Code imports recover schemas for Claude Code and Claude Desktop built-ins, plus conservative name-derived MCP schemas, when the raw… See the full description on the dataset page: https://huggingface.co/datasets/DedeProGames/claude-code-traces-pt-br.corpus-ptbr-v1
🇧🇷 Corpus PT-BR v1
Um corpus em Português Brasileiro voltado para pré-treinamento e fine-tuning de LLMs. Combina dados reais curados com uma camada sintética construída para ampliar a diversidade estilística, lexical e discursiva em português.
🔄 Visão Geral do Pipeline
📊 Estatísticas
Métrica
Valor
Total de documentos
8,399,857
Total de palavras
~4.84B
Tokens estimados
~6.29B
Tamanho (Parquet)
~17.9 GB
Idioma
Português… See the full description on the dataset page: https://huggingface.co/datasets/BrunoN-Dev/corpus-ptbr-v1.corpus-ptbr-v1
🇧🇷 Corpus PT-BR v1
Um corpus em Português Brasileiro voltado para pré-treinamento e fine-tuning de LLMs. Combina dados reais curados com uma camada sintética construída para ampliar a diversidade estilística, lexical e discursiva em português.
🔄 Visão Geral do Pipeline
📊 Estatísticas
Métrica
Valor
Total de documentos
8,399,857
Total de palavras
~4.84B
Tokens estimados
~6.29B
Tamanho (Parquet)
~17.9 GB
Idioma
Português Brasileiro (pt-br)… See the full description on the dataset page: https://huggingface.co/datasets/bratao/corpus-ptbr-v1.corpus-ptbr-v2
Corpus PT-BR v2
Um corpus em Portugues Brasileiro voltado para pre-treinamento, continuacao de pre-treinamento e fine-tuning de LLMs. Esta versao mantem a base e o pipeline geral do Madras1/corpus-ptbr-v1, com uma expansao adicional da camada sintetica gerada principalmente por modelos Mistral.
O que mudou na v2
A v2 preserva o desenho da v1 e adiciona um novo bloco sintetico local:
Componente novo
Valor
Documentos sinteticos adicionados
371,002
Palavras… See the full description on the dataset page: https://huggingface.co/datasets/Madras1/corpus-ptbr-v2.ptbr-books-publicos
PT-BR Public-Domain Books
Part of the MagTina350m pretrain corpus release by Dataseek
under the Magestic.ai brand. This is one of nine silver-layer datasets that fed
dataseek/magtina350m-base.
Summary
28 K curated Brazilian-Portuguese public-domain books. Used at ~2 epochs in MagTina350m pretrain (79 M unique tokens sampled to 158 M consumed). Useful as a small but high-quality literary slice.
Source and collection method
Curated public-domain Brazilian books… See the full description on the dataset page: https://huggingface.co/datasets/dataseek/ptbr-books-publicos.ptbr-gov-legal
PT-BR Legal & Government Documents
Part of the MagTina350m pretrain corpus release by Dataseek
under the Magestic.ai brand. This is one of nine silver-layer datasets that fed
dataseek/magtina350m-base.
Summary
935 K Brazilian legal and government documents: federal/state laws, court decisions, regulatory acts, official communications. Mixed corpus combining eduagarcia/LegalPT_dedup (HuggingFace) with a Kaggle Brazilian-legal-proceedings dump.
Source and collection… See the full description on the dataset page: https://huggingface.co/datasets/dataseek/ptbr-gov-legal.ptbr-wiki
PT-BR Wikipedia (cleaned, encyclopedia only)
Part of the MagTina350m pretrain corpus release by Dataseek
under the Magestic.ai brand. This is one of nine silver-layer datasets that fed
dataseek/magtina350m-base.
Summary
1.08 M Brazilian-Portuguese Wikipedia articles, cleaned of wiki-markup, templates, user-talk welcome banners ("Bem-vindo, X!"), VfD voting discussions, eliminação notification templates, JS/CSS gadget pages and talk-page conversations (detected by inline… See the full description on the dataset page: https://huggingface.co/datasets/dataseek/ptbr-wiki.corpus-ptbr-v1
🇧🇷 Corpus PT-BR v1
Um corpus em Português Brasileiro voltado para pré-treinamento e fine-tuning de LLMs. Combina dados reais curados com uma camada sintética construída para ampliar a diversidade estilística, lexical e discursiva em português.
🔄 Visão Geral do Pipeline
📊 Estatísticas
Métrica
Valor
Total de documentos
8,399,857
Total de palavras
~4.84B
Tokens estimados
~6.29B
Tamanho (Parquet)
~17.9 GB
Idioma
Português Brasileiro… See the full description on the dataset page: https://huggingface.co/datasets/strak2005/corpus-ptbr-v1.ptbr-dou
Diário Oficial da União 2025-2026 (DO1 + DO1E)
Part of the MagTina350m pretrain corpus release by Dataseek
under the Magestic.ai brand. This is one of nine silver-layer datasets that fed
dataseek/magtina350m-base.
Summary
20 K articles from the Brazilian Official Federal Gazette covering 2025-2026 — sections DO1 (regular edition) and DO1E (extra edition). Rich structured metadata: publication date, edition, page, PDF URL, ementa (summary), full text. Small corpus but high… See the full description on the dataset page: https://huggingface.co/datasets/dataseek/ptbr-dou.ptbr-blogs
PT-BR Blogs (long-form, C4-derived)
Part of the MagTina350m pretrain corpus release by Dataseek
under the Magestic.ai brand. This is one of nine silver-layer datasets that fed
dataseek/magtina350m-base.
Summary
185 K long-form Brazilian-Portuguese blog posts (≥ 5 K words each) extracted from C4 by filtering Blogspot, WordPress, Medium and similar platform domains. Higher per-document quality than generic web; useful for stylistic diversity and long-context training.… See the full description on the dataset page: https://huggingface.co/datasets/dataseek/ptbr-blogs.ptbr-academic
PT-BR Academic Corpus (CAPES theses + SciELO abstracts/books)
Part of the MagTina350m pretrain corpus release by Dataseek
under the Magestic.ai brand. This is one of nine silver-layer datasets that fed
dataseek/magtina350m-base.
Summary
1.86 M Brazilian-Portuguese academic documents combining CAPES thesis and dissertation abstracts, SciELO article abstracts and SciELO open-access books. Complements the fulltext SciELO articles corpus with broader coverage at shorter… See the full description on the dataset page: https://huggingface.co/datasets/dataseek/ptbr-academic.ptbr-creative-cpt-qwen35-08b-v02
PT-BR Creative CPT — Qwen3.5-0.8B data-prep v0.2
This repository is a derived, model/tokenizer-specific training artifact for continued pretraining experiments.
It is not the canonical text corpus.
Canonical source:
oliveirabruno01/ptbr-creative-cpt
Canonical corpus fingerprint:
21f72f64b3b73425bc78d91046a52aefddb8413b747d69f3422c31da8f536840
Identity
Model/tokenizer: Qwen/Qwen3.5-0.8B-Base
Context length: 2048
Data-prep version: v0.2
Primary split policy:… See the full description on the dataset page: https://huggingface.co/datasets/oliveirabruno01/ptbr-creative-cpt-qwen35-08b-v02.ptbr-creative-cpt
PT-BR Creative Corpus v0.1.0
A curated Brazilian-Portuguese creative-writing corpus for continued pretraining / midtraining research.
Status
This is the canonical corpus freeze, not a final model-specific training build.
Canonical text units: 1,354
Document/edition entities: 803
Characters: 82,538,439
Words (whitespace count): 13,929,410
Historical project estimate: 18,339,188 chars/4.5 tokens, retained only in the audit_metrics config.
The canonical corpus… See the full description on the dataset page: https://huggingface.co/datasets/oliveirabruno01/ptbr-creative-cpt.Kimi-K3-And-DeepSeek-V4-Pro-0813-Distillation-in-PT-BR
🇧 Destilação PT-BR com Raciocínio (Chain-of-Thought)
Este dataset contém exemplos de alta qualidade gerados através da destilação de modelos de ponta (Teacher Models) disponíveis via NVIDIA NIM, focados em instrução, raciocínio lógico e naturalidade em Português Brasileiro (PT-BR).
O grande diferencial deste dataset é a inclusão explícita do processo de pensamento (Chain-of-Thought / thinking) dos modelos professores, permitindo treinar modelos menores (Student Models) não… See the full description on the dataset page: https://huggingface.co/datasets/Davizig10jojo/Kimi-K3-And-DeepSeek-V4-Pro-0813-Distillation-in-PT-BR.ptbr-scielo
PT-BR SciELO Articles (Brazilian Open-Access Research)
Part of the MagTina350m pretrain corpus release by Dataseek
under the Magestic.ai brand. This is one of nine silver-layer datasets that fed
dataseek/magtina350m-base.
Summary
154 K full-text Brazilian Portuguese research and review articles from SciELO Brazil (post-2010), spanning health sciences, social sciences, humanities and engineering. Avg ~34 KB per article — the largest academic-prose corpus in the MagTina350m… See the full description on the dataset page: https://huggingface.co/datasets/dataseek/ptbr-scielo.stratasynth-termination-negotiations-pt-br
stratasynth-termination-negotiations-pt-br
Synthetic employment-termination negotiations between an employer / HR
representative and an employee who is being let go. Generated with StrataSynth's
dimensional engine, psychometrically and culturally conditioned to Brazil
(BR).
Every conversation is a two-party exchange grounded in the same scenario
(PRO-05 — employment termination): delivering the decision, stated reasons
(restructuring, performance, budget), severance and final… See the full description on the dataset page: https://huggingface.co/datasets/StrataSynth/stratasynth-termination-negotiations-pt-br.wikipedia-pt-br-instructions
wikipedia-pt-br-instructions-gemma
Dataset sintético de Instruction Following em português brasileiro, derivado de artigos da Wikipedia pt-BR.
Origem
Os exemplos derivam de costadev00/wikipedia-pt-br-extract e preservam source_page_id, source_title, source_dataset e a licença herdada cc-by-sa-3.0.
Processo
A geração é inspirada em Alpaca e Self-Instruct: cada artigo válido passa por uma chamada de analista documental que produz candidatos de… See the full description on the dataset page: https://huggingface.co/datasets/br-llm-data/wikipedia-pt-br-instructions.wikipedia-pt-br-instructions
wikipedia-pt-br-instructions-gemma
Dataset sintético de Instruction Following em português brasileiro, derivado de artigos da Wikipedia pt-BR.
Origem
Os exemplos derivam de costadev00/wikipedia-pt-br-extract e preservam source_page_id, source_title, source_dataset e a licença herdada cc-by-sa-3.0.
Processo
A geração é inspirada em Alpaca e Self-Instruct: cada artigo válido passa por uma chamada de analista documental que produz candidatos de instrução ancorados… See the full description on the dataset page: https://huggingface.co/datasets/costadev00/wikipedia-pt-br-instructions.
