CoolFace
19 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Madras1 /corpus-ptbr-v1 🇧🇷 Corpus PT-BR v1 Um corpus em Português Brasileiro voltado para pré-treinamento e fine-tuning de LLMs. Combina dados reais curados com uma camada sintética construída para ampliar a diversidade estilística, lexical e discursiva em português. 🔄 Visão Geral do Pipeline 📊 Estatísticas Métrica Valor Total de documentos 8,399,857 Total de palavras ~4.84B Tokens estimados ~6.29B Tamanho (Parquet) ~17.9 GB Idioma Português Brasileiro… See the full description on the dataset page: https://huggingface.co/datasets/Madras1/corpus-ptbr-v1.tabulartext-generation1M<n<10M5 likes352 downloads5mo agoHugging Face02DedeProGames /claude-code-traces-pt-brThis dataset was generated using teich by TeichAI claude Agent Traces This directory contains raw agent trace files generated by teich. JSONL files: 20 Training-ready tools Generated agent traces carry configured or recovered tool schemas so tools remain available for training even when a session did not call them. Native Claude Code imports recover schemas for Claude Code and Claude Desktop built-ins, plus conservative name-derived MCP schemas, when the raw… See the full description on the dataset page: https://huggingface.co/datasets/DedeProGames/claude-code-traces-pt-br.tabulartext-generationn<1K2 likes208 downloads2mo agoHugging Face03BrunoN-Dev /corpus-ptbr-v1 🇧🇷 Corpus PT-BR v1 Um corpus em Português Brasileiro voltado para pré-treinamento e fine-tuning de LLMs. Combina dados reais curados com uma camada sintética construída para ampliar a diversidade estilística, lexical e discursiva em português. 🔄 Visão Geral do Pipeline 📊 Estatísticas Métrica Valor Total de documentos 8,399,857 Total de palavras ~4.84B Tokens estimados ~6.29B Tamanho (Parquet) ~17.9 GB Idioma Português… See the full description on the dataset page: https://huggingface.co/datasets/BrunoN-Dev/corpus-ptbr-v1.tabulartext-generation1M<n<10M1 likes202 downloads2mo agoHugging Face04bratao /corpus-ptbr-v1 🇧🇷 Corpus PT-BR v1 Um corpus em Português Brasileiro voltado para pré-treinamento e fine-tuning de LLMs. Combina dados reais curados com uma camada sintética construída para ampliar a diversidade estilística, lexical e discursiva em português. 🔄 Visão Geral do Pipeline 📊 Estatísticas Métrica Valor Total de documentos 8,399,857 Total de palavras ~4.84B Tokens estimados ~6.29B Tamanho (Parquet) ~17.9 GB Idioma Português Brasileiro (pt-br)… See the full description on the dataset page: https://huggingface.co/datasets/bratao/corpus-ptbr-v1.tabulartext-generation1M<n<10M0 likes191 downloads5mo agoHugging Face05Madras1 /corpus-ptbr-v2 Corpus PT-BR v2 Um corpus em Portugues Brasileiro voltado para pre-treinamento, continuacao de pre-treinamento e fine-tuning de LLMs. Esta versao mantem a base e o pipeline geral do Madras1/corpus-ptbr-v1, com uma expansao adicional da camada sintetica gerada principalmente por modelos Mistral. O que mudou na v2 A v2 preserva o desenho da v1 e adiciona um novo bloco sintetico local: Componente novo Valor Documentos sinteticos adicionados 371,002 Palavras… See the full description on the dataset page: https://huggingface.co/datasets/Madras1/corpus-ptbr-v2.tabulartext-generation100K<n<1M0 likes121 downloads4mo agoHugging Face06dataseek /ptbr-books-publicos PT-BR Public-Domain Books Part of the MagTina350m pretrain corpus release by Dataseek under the Magestic.ai brand. This is one of nine silver-layer datasets that fed dataseek/magtina350m-base. Summary 28 K curated Brazilian-Portuguese public-domain books. Used at ~2 epochs in MagTina350m pretrain (79 M unique tokens sampled to 158 M consumed). Useful as a small but high-quality literary slice. Source and collection method Curated public-domain Brazilian books… See the full description on the dataset page: https://huggingface.co/datasets/dataseek/ptbr-books-publicos.tabulartext-generation10K<n<100K0 likes92 downloads5mo agoHugging Face07dataseek /ptbr-gov-legal PT-BR Legal & Government Documents Part of the MagTina350m pretrain corpus release by Dataseek under the Magestic.ai brand. This is one of nine silver-layer datasets that fed dataseek/magtina350m-base. Summary 935 K Brazilian legal and government documents: federal/state laws, court decisions, regulatory acts, official communications. Mixed corpus combining eduagarcia/LegalPT_dedup (HuggingFace) with a Kaggle Brazilian-legal-proceedings dump. Source and collection… See the full description on the dataset page: https://huggingface.co/datasets/dataseek/ptbr-gov-legal.tabulartext-generation100K<n<1M0 likes87 downloads5mo agoHugging Face08dataseek /ptbr-wiki PT-BR Wikipedia (cleaned, encyclopedia only) Part of the MagTina350m pretrain corpus release by Dataseek under the Magestic.ai brand. This is one of nine silver-layer datasets that fed dataseek/magtina350m-base. Summary 1.08 M Brazilian-Portuguese Wikipedia articles, cleaned of wiki-markup, templates, user-talk welcome banners ("Bem-vindo, X!"), VfD voting discussions, eliminação notification templates, JS/CSS gadget pages and talk-page conversations (detected by inline… See the full description on the dataset page: https://huggingface.co/datasets/dataseek/ptbr-wiki.tabulartext-generation1M<n<10M0 likes86 downloads5mo agoHugging Face09strak2005 /corpus-ptbr-v1 🇧🇷 Corpus PT-BR v1 Um corpus em Português Brasileiro voltado para pré-treinamento e fine-tuning de LLMs. Combina dados reais curados com uma camada sintética construída para ampliar a diversidade estilística, lexical e discursiva em português. 🔄 Visão Geral do Pipeline 📊 Estatísticas Métrica Valor Total de documentos 8,399,857 Total de palavras ~4.84B Tokens estimados ~6.29B Tamanho (Parquet) ~17.9 GB Idioma Português Brasileiro… See the full description on the dataset page: https://huggingface.co/datasets/strak2005/corpus-ptbr-v1.tabulartext-generation1M<n<10M0 likes80 downloads4mo agoHugging Face10dataseek /ptbr-dou Diário Oficial da União 2025-2026 (DO1 + DO1E) Part of the MagTina350m pretrain corpus release by Dataseek under the Magestic.ai brand. This is one of nine silver-layer datasets that fed dataseek/magtina350m-base. Summary 20 K articles from the Brazilian Official Federal Gazette covering 2025-2026 — sections DO1 (regular edition) and DO1E (extra edition). Rich structured metadata: publication date, edition, page, PDF URL, ementa (summary), full text. Small corpus but high… See the full description on the dataset page: https://huggingface.co/datasets/dataseek/ptbr-dou.tabulartext-generation10K<n<100K0 likes70 downloads5mo agoHugging Face11dataseek /ptbr-blogs PT-BR Blogs (long-form, C4-derived) Part of the MagTina350m pretrain corpus release by Dataseek under the Magestic.ai brand. This is one of nine silver-layer datasets that fed dataseek/magtina350m-base. Summary 185 K long-form Brazilian-Portuguese blog posts (≥ 5 K words each) extracted from C4 by filtering Blogspot, WordPress, Medium and similar platform domains. Higher per-document quality than generic web; useful for stylistic diversity and long-context training.… See the full description on the dataset page: https://huggingface.co/datasets/dataseek/ptbr-blogs.tabulartext-generation100K<n<1M1 likes67 downloads5mo agoHugging Face12dataseek /ptbr-academic PT-BR Academic Corpus (CAPES theses + SciELO abstracts/books) Part of the MagTina350m pretrain corpus release by Dataseek under the Magestic.ai brand. This is one of nine silver-layer datasets that fed dataseek/magtina350m-base. Summary 1.86 M Brazilian-Portuguese academic documents combining CAPES thesis and dissertation abstracts, SciELO article abstracts and SciELO open-access books. Complements the fulltext SciELO articles corpus with broader coverage at shorter… See the full description on the dataset page: https://huggingface.co/datasets/dataseek/ptbr-academic.tabulartext-generation1M<n<10M0 likes57 downloads5mo agoHugging Face13oliveirabruno01 /ptbr-creative-cpt-qwen35-08b-v02 PT-BR Creative CPT — Qwen3.5-0.8B data-prep v0.2 This repository is a derived, model/tokenizer-specific training artifact for continued pretraining experiments. It is not the canonical text corpus. Canonical source: oliveirabruno01/ptbr-creative-cpt Canonical corpus fingerprint: 21f72f64b3b73425bc78d91046a52aefddb8413b747d69f3422c31da8f536840 Identity Model/tokenizer: Qwen/Qwen3.5-0.8B-Base Context length: 2048 Data-prep version: v0.2 Primary split policy:… See the full description on the dataset page: https://huggingface.co/datasets/oliveirabruno01/ptbr-creative-cpt-qwen35-08b-v02.tabulartext-generation1K<n<10K0 likes56 downloads2d agoHugging Face14oliveirabruno01 /ptbr-creative-cpt PT-BR Creative Corpus v0.1.0 A curated Brazilian-Portuguese creative-writing corpus for continued pretraining / midtraining research. Status This is the canonical corpus freeze, not a final model-specific training build. Canonical text units: 1,354 Document/edition entities: 803 Characters: 82,538,439 Words (whitespace count): 13,929,410 Historical project estimate: 18,339,188 chars/4.5 tokens, retained only in the audit_metrics config. The canonical corpus… See the full description on the dataset page: https://huggingface.co/datasets/oliveirabruno01/ptbr-creative-cpt.tabulartext-generation1K<n<10K0 likes54 downloads2d agoHugging Face15Davizig10jojo /Kimi-K3-And-DeepSeek-V4-Pro-0813-Distillation-in-PT-BR 🇧 Destilação PT-BR com Raciocínio (Chain-of-Thought) Este dataset contém exemplos de alta qualidade gerados através da destilação de modelos de ponta (Teacher Models) disponíveis via NVIDIA NIM, focados em instrução, raciocínio lógico e naturalidade em Português Brasileiro (PT-BR). O grande diferencial deste dataset é a inclusão explícita do processo de pensamento (Chain-of-Thought / thinking) dos modelos professores, permitindo treinar modelos menores (Student Models) não… See the full description on the dataset page: https://huggingface.co/datasets/Davizig10jojo/Kimi-K3-And-DeepSeek-V4-Pro-0813-Distillation-in-PT-BR.tabulartext-generationn<1K0 likes53 downloads11d agoHugging Face16dataseek /ptbr-scielo PT-BR SciELO Articles (Brazilian Open-Access Research) Part of the MagTina350m pretrain corpus release by Dataseek under the Magestic.ai brand. This is one of nine silver-layer datasets that fed dataseek/magtina350m-base. Summary 154 K full-text Brazilian Portuguese research and review articles from SciELO Brazil (post-2010), spanning health sciences, social sciences, humanities and engineering. Avg ~34 KB per article — the largest academic-prose corpus in the MagTina350m… See the full description on the dataset page: https://huggingface.co/datasets/dataseek/ptbr-scielo.tabulartext-generation100K<n<1M0 likes50 downloads5mo agoHugging Face17StrataSynth /stratasynth-termination-negotiations-pt-br stratasynth-termination-negotiations-pt-br Synthetic employment-termination negotiations between an employer / HR representative and an employee who is being let go. Generated with StrataSynth's dimensional engine, psychometrically and culturally conditioned to Brazil (BR). Every conversation is a two-party exchange grounded in the same scenario (PRO-05 — employment termination): delivering the decision, stated reasons (restructuring, performance, budget), severance and final… See the full description on the dataset page: https://huggingface.co/datasets/StrataSynth/stratasynth-termination-negotiations-pt-br.tabulartext-generation1K<n<10K0 likes40 downloads3d agoHugging Face18br-llm-data /wikipedia-pt-br-instructionsgated wikipedia-pt-br-instructions-gemma Dataset sintético de Instruction Following em português brasileiro, derivado de artigos da Wikipedia pt-BR. Origem Os exemplos derivam de costadev00/wikipedia-pt-br-extract e preservam source_page_id, source_title, source_dataset e a licença herdada cc-by-sa-3.0. Processo A geração é inspirada em Alpaca e Self-Instruct: cada artigo válido passa por uma chamada de analista documental que produz candidatos de… See the full description on the dataset page: https://huggingface.co/datasets/br-llm-data/wikipedia-pt-br-instructions.tabulartext-generationn<1K0 likes6 downloads3mo agoHugging Face19costadev00 /wikipedia-pt-br-instructionsgated wikipedia-pt-br-instructions-gemma Dataset sintético de Instruction Following em português brasileiro, derivado de artigos da Wikipedia pt-BR. Origem Os exemplos derivam de costadev00/wikipedia-pt-br-extract e preservam source_page_id, source_title, source_dataset e a licença herdada cc-by-sa-3.0. Processo A geração é inspirada em Alpaca e Self-Instruct: cada artigo válido passa por uma chamada de analista documental que produz candidatos de instrução ancorados… See the full description on the dataset page: https://huggingface.co/datasets/costadev00/wikipedia-pt-br-instructions.tabulartext-generationn<1K0 likes3 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.