datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
llm_pt_leaderboard_resultsllm0to1-pt-aihub-624
AI-Hub 웹데이터 기반 한국어 말뭉치 (전처리)
LLM0to1-10b (SmolLM3 기반 10B 한/영 이중언어 LLM) 사전학습 코퍼스의 일부. huggingface.co/izlley
개요
카테고리: korean
원본 출처: AI-Hub 데이터셋 624 (웹데이터 기반 한국어 말뭉치)
라이선스: AI-Hub 이용약관(재배포 허가 확인)
토큰 수(우리 토크나이저 vocab 160k): 3.777B / 문서 120,000건
600B 믹스 내 역할: korean 카테고리(목표 25% = 150B). 카테고리 unique 62.5B 중 이 소스 6.0%(~9.07B 기여), 카테고리 전체 약 2.40 epoch 반복
전처리·필터링
zip 스트리밍 추출→한글비율≥0.25·길이≥40 필터→문서 exact-dedup(md5)→PII 스크럽(주민번호·전화·이메일)
토크나이저:… See the full description on the dataset page: https://huggingface.co/datasets/izlley/llm0to1-pt-aihub-624.mc4-pt-cleaned
Description
This is a clenned version of AllenAI mC4 PtBR section. The original dataset can be found here https://huggingface.co/datasets/allenai/c4
Clean procedure
We applied the same clenning procedure as explained here: https://gitlab.com/yhavinga/c4nlpreproc.git
The repository offers two strategies. The first one, found in the main.py file, uses pyspark to create a dataframe that can both clean the text and create a
pseudo mix on the entire dataset. We found this… See the full description on the dataset page: https://huggingface.co/datasets/thegoodfellas/mc4-pt-cleaned.calame-pt
CALAME-PT
Context-Aware LAnguage Modeling Evaluation for Portuguese
CALAME-PT is a PT benchmark composed of small texts (contexts) and their respective last words.
These contexts should, in theory, contain enough information so that a human or a model is capable of guessing its last word - without being too specific and/or too ambiguous.
Composition
CALAME-PT is composed of 2 "sets" of data - handwritten and generated.
Handwritten Set: contains 406… See the full description on the dataset page: https://huggingface.co/datasets/NOVA-vision-language/calame-pt.ovos-stt-bench-mls-pt-PT
OVOS stt bench — mls-pt-PT
Per-clip transcripts predictions of the registered
OVOS Plugin Arena
stt fighters over
facebook/multilingual_librispeech.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble workflow turns… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-mls-pt-PT.claude-code-traces-pt-brThis dataset was generated using teich by TeichAI
claude Agent Traces
This directory contains raw agent trace files generated by teich.
JSONL files: 20
Training-ready tools
Generated agent traces carry configured or recovered tool schemas so tools remain available for training even when a session did not call them.
Native Claude Code imports recover schemas for Claude Code and Claude Desktop built-ins, plus conservative name-derived MCP schemas, when the raw… See the full description on the dataset page: https://huggingface.co/datasets/DedeProGames/claude-code-traces-pt-br.bigbenchhard-mt-pt
BBH-PT (Big-Bench Hard)
Portuguese machine translation of BIG-Bench Hard, a challenging subset of the BIG-Bench benchmark covering diverse reasoning tasks.
Translated using a Finetuned GemmaX2-9B for pt-PT with rule-based adaptations.
Note: Some tasks (e.g., hyperbaton) are not translated as they do not transfer meaningfully to Portuguese.
Original Dataset: https://github.com/suzgunmirac/BIG-Bench-Hard
Note: This dataset is machine translated and may contain… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/bigbenchhard-mt-pt.Kayla-ptpytorch tensor version of Kayla! made with love by netcat420 :3
netcat7 <-- discord, add me! :P
pt_exams
PHEB - Portuguese High School Exams MCQ
MCQ set of PHEB a collection of Portuguese exam questions for evaluating language models on academic knowledge on the Portuguese curriculum.
For more details, see the PHEB paper.
This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite for evaluating large language models on European Portuguese.
Citation
If you use this dataset or AMALIA in your work… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/pt_exams.alpaca-cleaned-pt
Data Description
This HF data repository contains the Portuguese Alpaca dataset used in our study of monolingual versus multilingual instruction tuning.
GitHub
Paper
Creation
Machine-translated from yahma/alpaca-cleaned into Portuguese.
Usage
This data is intended to be used for Portuguese instruction tuning.
The dataset has roughly 52K instances in the JSON format.
Each instance has an instruction, an output, and an optional input. An example is shown… See the full description on the dataset page: https://huggingface.co/datasets/pinzhenchen/alpaca-cleaned-pt.amazon_massive_intent_pt-PTsorrel-T-gemma-3-27b-pt-seed0-documentsMaths-Grade-SchoolMaths-Grade-School
I am releasing large Grade School level Mathematics datatset.
This extensive dataset, comprising nearly one million instructions in JSON format, encapsulates a diverse array of topics fundamental to building a strong mathematical foundation.
This dataset is in instruction format so that model developers, researchers etc. can easily use this dataset.
Following Fields & sub Fields are covered:
Calculus
Probability
Algebra
Liner Algebra
Trigonometry
Differential Equations… See the full description on the dataset page: https://huggingface.co/datasets/pt-sk/Maths-Grade-School.MATH-PT
Math-PT: A Math Reasoning Benchmark for European and Brazilian Portuguese
Math-PT is a high-quality evaluation dataset designed to measure the mathematical reasoning capabilities of Large Language Models (LLMs) in Portuguese.
Unlike many existing benchmarks that rely on English translations, Math-PT uses native-language problems sourced from prestigious academic competitions and national exams in both Portugal and Brazil.
Dataset Details
Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/tiagoteixeira03/MATH-PT.Vision-COTovos-vad-bench-speech-vs-nonspeech-pt-PT
OVOS vad bench — speech-vs-nonspeech-pt-PT
Per-clip speech / non-speech decisions predictions of the registered
OVOS Plugin Arena
vad fighters over
PolyAI/minds14.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-vad-bench-speech-vs-nonspeech-pt-PT.ptpt-failure-set-gate
Where a local 27B actually breaks against a frontier model — a European-Portuguese failure-set gate
On broad everyday tasks, a clean local 27B is near-indistinguishable from a frontier model under blind judging. The gaps that remain are narrow, behavioral, and regex-detectable — which is exactly what small adapters fix. This dataset is the measurement instrument: six hard-sets with deterministic checks, plus the scorer and the methodology write-up.
Dataset summary… See the full description on the dataset page: https://huggingface.co/datasets/JigSawPT/ptpt-failure-set-gate.ovos-stt-bench-vocatives-pt-PT
OVOS stt bench — vocatives-pt-PT
Per-clip transcripts predictions of the registered
OVOS Plugin Arena
stt fighters over
Jarbas/VocativesEuropeanPortuguese.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble workflow… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-vocatives-pt-PT.social_i_qa_pt
SocialIQa dataset v1.4 (PT)
This is translation to the portuguese language of the dataset allenai/social_i_qa.Translations were done using three independent models:
Helsinki-NLP/opus-mt-tc-big-en-pt
unicamp-dl/translation-en-pt-t5
facebook/nllb-200-distilled-1.3B
Translations were evaluated using the evaluation metric GEMBA - GPT Estimation Metric Based Assessment
(from the article Large Language Models Are State-of-the-Art Evaluators of Translation Quality) using… See the full description on the dataset page: https://huggingface.co/datasets/fabiogr/social_i_qa_pt.MSCOCO_PT-BRwikipedia-pt-br-instruct-5k
Wikipedia PT-BR Instruct
wikipedia-pt-br-instruct is a synthetic supervised fine-tuning (SFT)
dataset in Brazilian Portuguese generated from Wikipedia-derived documents.
This release is an intermediate evaluation dataset produced with the
sft-dataset-creator pipeline from the run
wiki-ptbr-extract-calib-5kdocs-14tasks. It was generated from a fixed
revision of costadev00/wikipedia-pt-br-extract:
cdbd07dc4a3de6e64632c718710b3ae0ebaeb0ff
The dataset is intended for intermediate… See the full description on the dataset page: https://huggingface.co/datasets/costadev00/wikipedia-pt-br-instruct-5k.ovos-intent-bench-speech-massive-pt-PT
ovos-intent-bench-speech-massive-pt-PT
An OVOS Plugin Arena benchmark repository. The arena's prediction runner
publishes each plugin's raw output on the speech-massive-pt-PT dataset here, one JSON-lines
file per plugin under predictions/<lang>/, and where a sample-set manifest
governs scoring it lives under sample_sets/. The public leaderboard at
https://openvoiceos.github.io/ovos-plugin-arena/ is computed from these rows,
and anyone can recompute an entry from them without… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-intent-bench-speech-massive-pt-PT.dolly-15k-libretranslate-pt
Summary
databricks-dolly-15k ( https://huggingface.co/datasets/databricks/databricks-dolly-15k/ ) is an open source dataset of instruction-following records generated by thousands of Databricks employees in several of the behavioral categories outlined in the InstructGPT paper, including brainstorming, classification, closed QA, generation, information extraction, open QA, and summarization.
This is a portuguese translation done with libretranslate (… See the full description on the dataset page: https://huggingface.co/datasets/Gustrd/dolly-15k-libretranslate-pt.ovos-wake-word-bench-mlsw-negatives-pt-PT
ovos-wake-word-bench-mlsw-negatives-pt-PT
A sample-set manifest for the OVOS Plugin Arena, not an audio corpus. It holds
a seeded, reproducible list of clip identifiers selected from the
Multilingual Spoken Words Corpus
(MLCommons, CC-BY-4.0), with the seed and the source row count recorded, so
every wake-word plugin scored against these negatives is scored on exactly the
same clips and false-accept rates are comparable across plugins.
The audio is not redistributed here; it… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-mlsw-negatives-pt-PT.massive-pt-br
MASSIVE pt-BR — localização brasileira
Localização para português do Brasil do split pt-PT do
MASSIVE (Amazon; CC-BY-4.0) — o benchmark
original não tem locale pt-BR. 15.852 frases (train 10.948 / validation 1.929 /
test 2.930 aceitas), com anotação de slots preservada ([campo : valor]):
a frase limpa é derivada da anotada por construção, então os valores de slot são
literais por construção.
Método: professor LLM recebe apenas a frase ANOTADA e devolve a versão pt-BR anotada;… See the full description on the dataset page: https://huggingface.co/datasets/Magurofg/massive-pt-br.bifonia-pt-homographs
bifonia — Portuguese Heterophonic Homograph Disambiguation
Labelled European-Portuguese (pt-PT) sentences for 27 heterophonic homographs —
words with identical spelling whose pronunciation (IPA) depends on part of speech
or meaning, e.g. para (preposition ˈpɐɾɐ vs verb ˈpaɾɐ), molho
(sauce ˈmoʎu vs bundle ˈmɔʎu), corte (royal court ˈkoɾtɨ vs cut ˈkɔɾtɨ).
Useful for grapheme-to-phoneme / TTS front-ends and for POS disambiguation.
Schema
field
description… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/bifonia-pt-homographs.Astra-Orbis-5.2K-PT-BR
Astra 5.2K PT-BR Extra High
Dataset conversacional em português brasileiro, curado para instruction tuning, SFT e treinamento de modelos com foco em respostas detalhadas, raciocínio, explicações passo a passo e comportamento de assistente. Contém 3 milhões de tolkens no total do dataset.
O corpus contém conversas no formato chat, exemplos de programação, tarefas analíticas, perguntas educacionais, resolução de problemas, matemática, escrita, tradução, produtividade e respostas… See the full description on the dataset page: https://huggingface.co/datasets/guell00/Astra-Orbis-5.2K-PT-BR.PT-Culture_Data
Portuguese Cultural SFT Dataset
A supervised fine-tuning dataset for European Portuguese (PT-PT) cultural knowledge, built to teach models the traditions, figures, places, and expressions of Portuguese culture.
The dataset contains 216,832 examples across two versions in conversational SFT format, organised into ten cultural domains:
Domain
v1 (95,815)
v2 (121,017)
Total
Personalities
36,370
62,788
99,158
Audiovisual
15,786
26,857
42,643
Heritage
8,020… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/PT-Culture_Data.ifeval-pt
IFEval-pt
Benchmark de seguimento de instrução em português do Brasil, com 308 itens cuja
verificação é feita por código — sem juiz LLM, sem GPU, sem opinião. Reproduzir custa
alguns minutos de CPU.
O que ele tem de diferente
Já existem versões do IFEval em português: amalia-llm/IFEval-mt-pt (541 itens, pt-PT),
Polygl0t/IFEval-PT (300), a fatia pt do facebook/Multi-IF (524) e
akawaai/Portuguese-IFEval (618). Todas são tradução do conjunto em inglês — medem… See the full description on the dataset page: https://huggingface.co/datasets/Magurofg/ifeval-pt.smoltalk2_everyday_conv_pt
SMOL Everyday Conversation PT
This dataset consists of a Portuguese version of the smoltalk_smollm3_everyday_conversations_no_think split of HuggingFaceTB/smoltalk2. The first two user turns were translated as well as the first assistant turn, the continuation of the conversation was generated using Gemma 3-27B.
Note: This dataset is machine translated and may contain translation errors or artifacts.
This dataset is provided as part of the AMALIA project and is… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/smoltalk2_everyday_conv_pt.
