CoolFace
Modelpublic

AIInstituteLNCC/manaca-1b-base

sourceHugging Facecc-by-4.0updated 24d agoView on Hugging Face
37likes2.3kdownloads
Model Card

Manacá-1B (base)

<p align="center"> <img src="manacaidentity.png" width="340" alt="Manacá — Tibouchina mutabilis: os três estágios florais como metáfora do treinamento do LLM / the three flowering colours as a metaphor for language-model maturation"/> </p>

<p align="center"><em>LLM base aberto e reprodutível para o português do Brasil · Open, reproducible Brazilian-Portuguese base LLM</em></p>

<p align="center"> <a href="https://creativecommons.org/licenses/by/4.0/"><img src="https://img.shields.io/badge/License-CC%20BY%204.0-lightgrey.svg" alt="License: CC BY 4.0"></a> <a href="https://github.com/Instituto-IA-LNCC/manaca-1b-base"><img src="https://img.shields.io/badge/Code-GitHub-181717.svg?logo=github" alt="Code"></a> <img src="https://img.shields.io/badge/Language-PT--BR-009c3b.svg" alt="PT-BR"> <img src="https://img.shields.io/badge/Params-1.72B-purple.svg" alt="1.72B"> <img src="https://img.shields.io/badge/Base-pretrained-8A2BE2.svg" alt="base"> </p>

[🇧🇷 Português](#português) · [🇬🇧 English](#english)

Manacá-1B é um modelo de linguagem decoder-only de ~1,72 bilhão de parâmetros, treinado do zero para o português do Brasil. Modelo base (pré-treinado, sem ajuste de instrução). Manacá-1B is a ~1.72B-parameter decoder-only language model trained from scratch for Brazilian Portuguese. This is a base (pretrained, non-instruction-tuned) model.

  • —Código, logs e pipeline reprodutível | Code, logs, and reproducible pipeline: https://github.com/Instituto-IA-LNCC/manaca-1b-base
  • —Cooperação | Cooperation: LNCC (Instituto de IA) × NII/LLM-jp

Português

Descrição do modelo

Manacá-1B é um transformer decoder-only no estilo Llama-3, treinado do zero para o português do Brasil com um pipeline totalmente containerizado e reprodutível. É um modelo base: aprendeu a modelar a língua por pré-treino, mas não passou por ajuste de instrução (SFT) nem alinhamento (DPO/RLHF). Ele completa e continua texto; não segue instruções nem conversa.

ItemValor
Parâmetros1.722.951.680 (~1,72B)
Camadas / dim / FFN24 / 2048 / 8192 (SwiGLU)
Cabeças / grupos KV / head dim32 / 8 (GQA) / 64
Posição / norma / biasRoPE (θ=500000) / RMSNorm (eps 1e-5) / bias em todas as lineares
Embeddingsdesamarradas (entrada e saída separadas)
Contexto / precisão4096 / bfloat16
Vocabulário64.000 (padded para 64.128)
FrameworkMegatron-LM (fork LLM-jp), otimizador distribuído
Referência de arquiteturaLLM-jp-3.1-1.8B

Tokenizador (leia isto)

O tokenizador é um SentencePiece unigram de 64k com a normalização nmt_nfkc_cf (NFKC + case folding). O modelo é lowercase por construção: o texto de entrada é normalizado para minúsculas antes da segmentação. Os arquivos de tokenizador publicados aqui já incluem a correção que reproduz exatamente a tokenização de treino (normalizador Sequence([NFKC, Lowercase])). Use sempre o AutoTokenizer deste repositório; um tokenizador sem esse normalizador tokeniza maiúsculas por byte-fallback e degrada os resultados de forma invisível.

Dados de treino

Corpus curado, aberto, de ~20,1B tokens únicos (pré-treino de ~2 épocas ⇒ ~41,9B tokens de atualização, ~24 tokens/parâmetro, próximo do compute-optimal de Chinchilla; repetição até ~4 épocas é quase tão boa quanto dados frescos, Muennighoff et al., 2023).

FonteDomínioTokensLicença
GigaVerbo (TucanoBR) — subamostraWeb geral PT-BR~14,7 B (73,1%)Apache-2.0
Ulysses Tesemõ (USP)Jurídico / legislativo~4,8 B (23,9%)Domínio público
Wikipedia-pt (2023-11)Enciclopédico~0,6 B (3,1%)CC BY-SA 4.0

Detalhes de treino

20.000 passos, batch global 512, sequência 4096. Adam (β 0,9/0,999, ε 1e-8), weight decay 0,1, clip de gradiente 1,0, learning rate 3e-4 → 3e-5 (cosseno, warmup de 2.000), z-loss, bf16, recompute full, seed 1234. Treinado em 2 GPUs de 24 GB (sem NVLink) com paralelismo de dados (ZeRO-1). A corrida foi estável: 0 iterações puladas e 0 NaN; loss de treino 11,41 → 2,48; loss de validação 2,07 nats.

<p align="center"> <img src="trainingdynamicsen.png" width="760" alt="Dinâmica de pré-treino do Manacá-1B: loss de treino e validação, norma do gradiente e learning rate ao longo de ~42B tokens"/> </p>

Avaliação

Acurácia (%) em quatro benchmarks de português, mesmo harness. CALAME-PT por geração da última palavra; ARC-Challenge-PT (25-shot), HellaSwag-PT (10-shot) e LAMBADA-PT (0-shot) por log-verossimilhança.

BenchmarkManacá-1BMelhor par sub-7BSabiá-7B
CALAME-PT60,6360,39 (GlórIA-1b3)63,23
LAMBADA-PT45,3137,38 (mGPT-1b3)63,67
HellaSwag-PT41,6148,63 (Tucano-2b4)64,55
ARC-Ch-PT27,1830,85 (Tucano-2b4)46,67

O Manacá-1B é o modelo mais forte abaixo de 7B na predição da última palavra (LAMBADA-PT), empata com os melhores modelos de 1 a 2 B no CALAME-PT, e fica na faixa de acaso em ARC-Challenge-PT (como todo modelo base nessa escala). Detalhes, testes pareados de McNemar e validação do harness no repositório e no preprint.

<p align="center"> <img src="benchmarkspaperen.png" width="640" alt="Manacá-1B em quatro benchmarks de português, acurácia vs. parâmetros (escala log), com IC95%"/> </p>

Uso

python
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

model_id = "menezesbruno/manaca-1b-base"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id, torch_dtype=torch.bfloat16, device_map="auto")

# O modelo é lowercase por construção (o tokenizador normaliza a entrada).
prompt = "a capital do brasil é"
inputs = tok(prompt, return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=40, do_sample=False)
print(tok.decode(out[0], skip_special_tokens=True))

Uso pretendido e limitações

Pesquisa e uso como modelo base para português do Brasil: modelagem de língua, continuação de texto, e ponto de partida para ajuste fino/instrução. Não é um modelo de instrução nem de chat; não segue comandos sem ajuste posterior. É um modelo base de ~1,72B: erra fatos, pode gerar conteúdo incorreto ou enviesado refletindo os dados de treino, e tem raciocínio limitado (próximo do acaso em múltipla escolha). Avalie antes de qualquer uso sensível.


English

Model description

Manacá-1B is a Llama-3-style decoder-only transformer trained from scratch for Brazilian Portuguese with a fully containerized, reproducible pipeline. It is a base model: it learned to model the language through pretraining but has not been instruction-tuned (SFT) or aligned (DPO/RLHF). It completes and continues text; it does not follow instructions or chat.

ItemValue
Parameters1,722,951,680 (~1.72B)
Layers / model dim / FFN24 / 2048 / 8192 (SwiGLU)
Heads / KV groups / head dim32 / 8 (GQA) / 64
Position / norm / biasRoPE (θ=500000) / RMSNorm (eps 1e-5) / all linear layers
Embeddingsuntied (separate input and output)
Context length / precision4096 / bfloat16
Vocabulary64,000 (padded to 64,128)
FrameworkMegatron-LM (LLM-jp fork), distributed optimizer
Architecture referenceLLM-jp-3.1-1.8B

Tokenizer (read this)

The tokenizer is a 64k SentencePiece unigram with nmt_nfkc_cf normalization (NFKC + case folding). The model is lowercase by construction: input text is lowercased before segmentation. The tokenizer files published here already include the fix that reproduces the training tokenization exactly (normalizer Sequence([NFKC, Lowercase])). Always use the AutoTokenizer from this repository; a tokenizer without that normalizer routes capitalized text to byte-fallback and silently degrades results.

Training data

A curated, openly licensed corpus of ~20.1B unique tokens (pretraining for ~2 epochs ⇒ ~41.9B update tokens, ~24 tokens/parameter, close to the Chinchilla compute-optimal point; repeating data up to ~4 epochs is nearly as good as fresh data, Muennighoff et al., 2023).

SourceDomainTokensLicense
GigaVerbo (TucanoBR) — subsampleGeneral web PT-BR~14.7B (73.1%)Apache-2.0
Ulysses Tesemõ (USP)Legal / legislative~4.8B (23.9%)Public domain
Wikipedia-pt (2023-11)Encyclopedic~0.6B (3.1%)CC BY-SA 4.0

Training details

20,000 steps, global batch 512, sequence length 4096. Adam (β 0.9/0.999, ε 1e-8), weight decay 0.1, gradient clip 1.0, learning rate 3e-4 → 3e-5 (cosine, 2,000-step warmup), z-loss, bf16, full recompute, seed 1234. Trained on 2 GPUs of 24 GB (no NVLink) with data parallelism (ZeRO-1). The run was stable: 0 skipped and 0 NaN steps; training loss 11.41 → 2.48; validation loss 2.07 nats.

<p align="center"> <img src="trainingdynamicsen.png" width="760" alt="Manacá-1B pretraining dynamics: training and validation loss, gradient norm, and learning rate over ~42B tokens"/> </p>

Evaluation

Accuracy (%) on four Portuguese benchmarks, one harness. CALAME-PT by last-word generation; ARC-Challenge-PT (25-shot), HellaSwag-PT (10-shot), and LAMBADA-PT (0-shot) by log-likelihood.

BenchmarkManacá-1BBest sub-7B peerSabiá-7B
CALAME-PT60.6360.39 (GlórIA-1b3)63.23
LAMBADA-PT45.3137.38 (mGPT-1b3)63.67
HellaSwag-PT41.6148.63 (Tucano-2b4)64.55
ARC-Ch-PT27.1830.85 (Tucano-2b4)46.67

Manacá-1B is the strongest model below 7B on last-word prediction (LAMBADA-PT), ties the best 1-to-2B models on CALAME-PT, and sits near chance on ARC-Challenge-PT (as does every base model at this scale). Details, paired McNemar tests, and harness validation are in the repository and the preprint.

<p align="center"> <img src="benchmarkspaperen.png" width="640" alt="Manacá-1B across four Portuguese benchmarks, accuracy vs. parameters (log scale), with 95% CIs"/> </p>

Usage

python
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

model_id = "menezesbruno/manaca-1b-base"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id, torch_dtype=torch.bfloat16, device_map="auto")

# The model is lowercase by construction (the tokenizer normalizes the input).
prompt = "a capital do brasil é"
inputs = tok(prompt, return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=40, do_sample=False)
print(tok.decode(out[0], skip_special_tokens=True))

Intended use and limitations

Research and use as a base model for Brazilian Portuguese: language modeling, text continuation, and a starting point for fine-tuning/instruction-tuning. It is not an instruction or chat model and will not follow commands without further tuning. As a ~1.72B base model it makes factual mistakes, may produce incorrect or biased content reflecting its training data, and has limited reasoning (near chance on multiple choice). Evaluate before any sensitive use.


Citation

bibtex
@misc{manaca1b2026,
  title  = {Manac\'a-1B: An Open, Reproducible Brazilian-Portuguese Language Model
            and a Tokenizer-Aware, Paired Evaluation},
  author = {Menezes, Bruno Leonardo Santos and Cardoso, Carlos Leonardo Souza and
            Porto, Fabio Andr\'e Machado},
  year   = {2026},
  note   = {LNCC (AI Institute) $\times$ NII/LLM-jp},
  url    = {https://github.com/Instituto-IA-LNCC/manaca-1b-base}
}

Acknowledgments

Developed at the Artificial Intelligence Institute of the National Laboratory for Scientific Computing (LNCC), Brazil, in cooperation with the National Institute of Informatics (NII), Japan, and the LLM-jp project, whose open methodology, tooling (the LLM-jp fork of Megatron-LM), and the LLM-jp-3.1-1.8B recipe this model builds on.

License: Creative Commons Attribution 4.0 International (CC BY 4.0).