strak2005/corpus-ptbr-v1
🇧🇷 Corpus PT-BR v1 Um corpus em PortuguĂŞs Brasileiro voltado para prĂ©-treinamento e fine-tuning de LLMs. Combina dados reais curados com uma camada sintĂ©tica construĂda para ampliar a diversidade estilĂstica, lexical e discursiva em portuguĂŞs. 🔄 VisĂŁo Geral do Pipeline 📊 EstatĂsticas MĂ©trica Valor Total de documentos 8,399,857 Total de palavras ~4.84B Tokens estimados ~6.29B Tamanho (Parquet) ~17.9 GB Idioma PortuguĂŞs… See the full description on the dataset page: https://huggingface.co/datasets/strak2005/corpus-ptbr-v1.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face