datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PT-HF500B
PT-HF500B (FinePhrase)
Overview
FinePhrase is a large-scale synthetic dataset designed for high-quality language modeling, reasoning, and instruction-following tasks. It transforms raw educational web data into structured, instruction-rich formats suitable for training advanced language models.
This dataset has been extensively used in the pre-training pipeline of TNSA models, including:
NGen-3
NGen-4
NGen-4-OW
It plays a critical role in improving reasoning ability… See the full description on the dataset page: https://huggingface.co/datasets/TNSA/PT-HF500B.claude-code-traces-pt-brThis dataset was generated using teich by TeichAI
claude Agent Traces
This directory contains raw agent trace files generated by teich.
JSONL files: 20
Training-ready tools
Generated agent traces carry configured or recovered tool schemas so tools remain available for training even when a session did not call them.
Native Claude Code imports recover schemas for Claude Code and Claude Desktop built-ins, plus conservative name-derived MCP schemas, when the raw… See the full description on the dataset page: https://huggingface.co/datasets/DedeProGames/claude-code-traces-pt-br.pt_exams
PHEB - Portuguese High School Exams MCQ
MCQ set of PHEB a collection of Portuguese exam questions for evaluating language models on academic knowledge on the Portuguese curriculum.
For more details, see the PHEB paper.
This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite for evaluating large language models on European Portuguese.
Citation
If you use this dataset or AMALIA in your work… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/pt_exams.Kimi-K3-And-DeepSeek-V4-Pro-0813-Distillation-in-PT-BR
🇧 Destilação PT-BR com Raciocínio (Chain-of-Thought)
Este dataset contém exemplos de alta qualidade gerados através da destilação de modelos de ponta (Teacher Models) disponíveis via NVIDIA NIM, focados em instrução, raciocínio lógico e naturalidade em Português Brasileiro (PT-BR).
O grande diferencial deste dataset é a inclusão explícita do processo de pensamento (Chain-of-Thought / thinking) dos modelos professores, permitindo treinar modelos menores (Student Models) não… See the full description on the dataset page: https://huggingface.co/datasets/Davizig10jojo/Kimi-K3-And-DeepSeek-V4-Pro-0813-Distillation-in-PT-BR.wildguardmix-ptpt
WildGuardTest-PT
Portuguese machine translation of WildGuardTest, a benchmark for evaluating safety guardrails in language models.
Translated using Gemma-4 31B-It.
Original Dataset: https://huggingface.co/datasets/walledai/WildGuardTest
Note: This dataset is machine translated and may contain translation errors or artifacts.
This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite for evaluating large… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/wildguardmix-ptpt.stratasynth-termination-negotiations-pt-br
stratasynth-termination-negotiations-pt-br
Synthetic employment-termination negotiations between an employer / HR
representative and an employee who is being let go. Generated with StrataSynth's
dimensional engine, psychometrically and culturally conditioned to Brazil
(BR).
Every conversation is a two-party exchange grounded in the same scenario
(PRO-05 — employment termination): delivering the decision, stated reasons
(restructuring, performance, budget), severance and final… See the full description on the dataset page: https://huggingface.co/datasets/StrataSynth/stratasynth-termination-negotiations-pt-br.toxichat-ptpt
ToxicChat-PT
Portuguese machine translation of ToxicChat, a benchmark for detecting toxic content in conversational AI.
Translated using Gemma-4 31B-It.
Original Dataset: https://huggingface.co/datasets/lmsys/toxic-chat
Note: This dataset is machine translated and may contain translation errors or artifacts.
This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite for evaluating large language models on… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/toxichat-ptpt.aime-1983-2024-ptpt
AIME-PT (1983-2024)
Portuguese translation of problems from the American Invitational Mathematics Examination (AIME) spanning 1983-2024.
Translated using Gemma-4 31B-It.
Original Dataset: https://artofproblemsolving.com/wiki/index.php/AIME_Problems_and_Solutions
Note: This dataset is machine translated and may contain translation errors or artifacts.
This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/aime-1983-2024-ptpt.amalia-cita-legal
AMALIA cita-legal — grounded legal citation SFT (RAG-first)
(question + real source excerpts → answer that cites the exact article, or a refusal when the excerpts don't answer the question). Built for
specializing AMALIA-9B
toward Portuguese legal text, as the grounded-answering counterpart to
teex-pt/amalia-sum-dre.
Derived from
teex-pt/leis-pt-consolidada
by teex-pt/pt-amalia.
Why this exists (RAG-first, not closed-book)
leis-pt's own project spec concludes that… See the full description on the dataset page: https://huggingface.co/datasets/teex-pt/amalia-cita-legal.wikipedia-pt-br-instructions
wikipedia-pt-br-instructions-gemma
Dataset sintético de Instruction Following em português brasileiro, derivado de artigos da Wikipedia pt-BR.
Origem
Os exemplos derivam de costadev00/wikipedia-pt-br-extract e preservam source_page_id, source_title, source_dataset e a licença herdada cc-by-sa-3.0.
Processo
A geração é inspirada em Alpaca e Self-Instruct: cada artigo válido passa por uma chamada de analista documental que produz candidatos de… See the full description on the dataset page: https://huggingface.co/datasets/br-llm-data/wikipedia-pt-br-instructions.anneal_pt_v3_clean_v2_reshard_7168
anneal_pt_v3 clean v2 — balanced 7,168 shards
This is the cleaned and deterministically re-sharded training split of
anneal_pt_v3. It contains 2,103,321,326 rows in 7,168 Parquet files
(3,376,320,935,167 stored bytes).
The shard count is 8 × lcm(128, 448), so it divides evenly for both intended
data-parallel configurations:
Data ranks
Shards per rank
Maximum relative load deviation
128
56
1.4942%
448
16
0.3888%
The files are ordered by a fixed-capacity LPT… See the full description on the dataset page: https://huggingface.co/datasets/rides21/anneal_pt_v3_clean_v2_reshard_7168.wikipedia-pt-br-instructions
wikipedia-pt-br-instructions-gemma
Dataset sintético de Instruction Following em português brasileiro, derivado de artigos da Wikipedia pt-BR.
Origem
Os exemplos derivam de costadev00/wikipedia-pt-br-extract e preservam source_page_id, source_title, source_dataset e a licença herdada cc-by-sa-3.0.
Processo
A geração é inspirada em Alpaca e Self-Instruct: cada artigo válido passa por uma chamada de analista documental que produz candidatos de instrução ancorados… See the full description on the dataset page: https://huggingface.co/datasets/costadev00/wikipedia-pt-br-instructions.financial-sentiment-pt
Financial Sentiments (PT-BR)
A dataset containing ~63k rows of translated Financial Sentiment data in Portuguese.
Pipeline
Dataset cleaning (dedup and benchmark decontamination) -> merge -> Translation Pipeline -> MiniCPM5 2B -> Kiwi COMET as scorer -> drop below a threshold when doing sft (default = 0.5). Exact code at: https://github.com/mansa-team/musa
Data
n_scored=64685
mean=0.6893
n_below=6975
