datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PT-HF500B
PT-HF500B (FinePhrase)
Overview
FinePhrase is a large-scale synthetic dataset designed for high-quality language modeling, reasoning, and instruction-following tasks. It transforms raw educational web data into structured, instruction-rich formats suitable for training advanced language models.
This dataset has been extensively used in the pre-training pipeline of TNSA models, including:
NGen-3
NGen-4
NGen-4-OW
It plays a critical role in improving reasoning ability… See the full description on the dataset page: https://huggingface.co/datasets/TNSA/PT-HF500B.laion400m-ptmteb-pt-results
🇧🇷 MTEB-BR — Benchmark Results
Canonical results store for MTEB-BR, a native Brazilian-Portuguese text-embedding benchmark.
93 models · 22 native PT-BR tasks · 7 categories · no machine translation
What is this?
This repository is the canonical, machine-readable results store for MTEB-BR — a benchmark that evaluates text-embedding models on native Brazilian Portuguese (data created or found in Portuguese; machine-translated corpora such as… See the full description on the dataset page: https://huggingface.co/datasets/MTEB-BR/mteb-pt-results.ClassiCC-PT
📚 ClassiCC-PT: Classified Common Crawl Corpus for Portuguese
📖 Overview
ClassiCC-PT (Classified Common Crawl – Portuguese) is a large-scale web corpus containing ~120B Portuguese tokens extracted from Common Crawl snapshots. It is specifically curated for training large language models in Portuguese, with a focus on data quality, language specificity, and targeted filtering.
This corpus was created as part of a study on continued pretraining for adapting English-trained… See the full description on the dataset page: https://huggingface.co/datasets/ClassiCC-Corpus/ClassiCC-PT.in1k_clip_qwen25vl_3b_224res_64tokens_new_ptin1k_clip_qwen25vl_3b_448res_256tokens_new_merged_ptGSA-PT-Qwen2-7B-Instruct-chunk4-chunk4-data
GSA-PT-Qwen2-7B-Instruct-chunk4-chunk4-data
This is the continue pretraining dataset used for training GSA (Gist Sparse Attention) models based on Qwen2-7B-Instruct with chunk size chunk4-chunk4.
Each sample is tokenized and formatted with GSA gist tokens for continue pretraining.
Paper
GSA: Gist Sparse Attention via Learnable Compression and Selective Unfolding
Related Model
yuzhenm/GSA-PT-Qwen2-7B-Instruct-chunk4-chunk4 — model trained on this dataset
eval_record-pick-and-place-pos5-so101_pt-ft-3epThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 36,
"total_frames": 43756,
"total_tasks": 2,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:36"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/hbseong/eval_record-pick-and-place-pos5-so101_pt-ft-3ep.eval_record-pick-and-place-so101_pt-ft-3epThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 94,
"total_frames": 105687,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:94"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/hbseong/eval_record-pick-and-place-so101_pt-ft-3ep.generations-simnpo_gemma-3-12b-pt_20260416_171305-corpus_sweep_post_evalovos-stt-bench-mls-pt-PT
OVOS stt bench — mls-pt-PT
Per-clip transcripts predictions of the registered
OVOS Plugin Arena
stt fighters over
facebook/multilingual_librispeech.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble workflow turns… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-mls-pt-PT.claude-code-traces-pt-brThis dataset was generated using teich by TeichAI
claude Agent Traces
This directory contains raw agent trace files generated by teich.
JSONL files: 20
Training-ready tools
Generated agent traces carry configured or recovered tool schemas so tools remain available for training even when a session did not call them.
Native Claude Code imports recover schemas for Claude Code and Claude Desktop built-ins, plus conservative name-derived MCP schemas, when the raw… See the full description on the dataset page: https://huggingface.co/datasets/DedeProGames/claude-code-traces-pt-br.mteb-pt-results
Note: This dataset has moved to MTEB-BR/mteb-pt-results. This legacy copy remains only to preserve its archival DOI (10.57967/hf/9377); the maintained version lives under the MTEB-BR organization.
🇧🇷 MTEB-BR — Benchmark Results
Canonical results store for MTEB-BR, a native Brazilian-Portuguese text-embedding benchmark.
93 models · 22 native PT-BR tasks · 7 categories · no machine translation
What is this?
This repository is the canonical… See the full description on the dataset page: https://huggingface.co/datasets/mteb-pt/mteb-pt-results.GSA-PT-Qwen2-7B-Instruct-chunk16-data
GSA-PT-Qwen2-7B-Instruct-chunk16-data
This is the continue pretraining dataset used for training GSA (Gist Sparse Attention) models based on Qwen2-7B-Instruct with chunk size chunk16.
Each sample is tokenized and formatted with GSA gist tokens for continue pretraining.
Paper
GSA: Gist Sparse Attention via Learnable Compression and Selective Unfolding
Related Model
yuzhenm/GSA-PT-Qwen2-7B-Instruct-chunk16 — model trained on this dataset
pt_exams
PHEB - Portuguese High School Exams MCQ
MCQ set of PHEB a collection of Portuguese exam questions for evaluating language models on academic knowledge on the Portuguese curriculum.
For more details, see the PHEB paper.
This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite for evaluating large language models on European Portuguese.
Citation
If you use this dataset or AMALIA in your work… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/pt_exams.ToxSyn-PT
Dataset Summary
ToxSyn-PT is a large-scale synthetic dataset designed for fine-grained hate speech detection in Brazilian Portuguese. It comprises 53,274 sentences equally balanced between toxic and non-toxic labels, covering nine legally protected minority groups (including Black, Women, LGBTQIA+, Native Brazilian, Muslim, Jewish, and Elderly).
Unlike most existing datasets that only capture hostile mentions, ToxSyn-PT systematically includes non-toxic counterexamples (benign… See the full description on the dataset page: https://huggingface.co/datasets/AKCIT/ToxSyn-PT.ovos-vad-bench-speech-vs-nonspeech-pt-PT
OVOS vad bench — speech-vs-nonspeech-pt-PT
Per-clip speech / non-speech decisions predictions of the registered
OVOS Plugin Arena
vad fighters over
PolyAI/minds14.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-vad-bench-speech-vs-nonspeech-pt-PT.eval_record-pick-and-place-ez2-so101_pt-ft-3epThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 21,
"total_frames": 13660,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:21"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/hbseong/eval_record-pick-and-place-ez2-so101_pt-ft-3ep.glue-ptptGLUE-PTPT is an European Portuguese translation of the GLUE benchmark using DeepL Pro.ovos-stt-bench-vocatives-pt-PT
OVOS stt bench — vocatives-pt-PT
Per-clip transcripts predictions of the registered
OVOS Plugin Arena
stt fighters over
Jarbas/VocativesEuropeanPortuguese.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble workflow… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-vocatives-pt-PT.fine_web2_odia_ptmultieurlex21-pt-semantic-cache
MultiEURLEX-21 PT — frozen semantic chunk embeddings
Precomputed, unit-normalised chunk embeddings of the official MTEB
MultiEURLEXMultilabelClassification Portuguese data, produced once so that
downstream experiments (frozen MaleCNS connectome reservoir, label-free
controls, ablations) never pay the sentence encoders again. No benchmark labels
are stored or used here.
source dataset: mteb/eurlex-multilingual revision 2aea5a6dc8fdcfeca41d0fb963c0a338930bde5c, subset pt
splits:… See the full description on the dataset page: https://huggingface.co/datasets/franklinbaldo/multieurlex21-pt-semantic-cache.mmlu-pt-undergraduate-fullGSA-PT-Qwen2-7B-Instruct-chunk8-data
GSA-PT-Qwen2-7B-Instruct-chunk8-data
This is the continue pretraining dataset used for training GSA (Gist Sparse Attention) models based on Qwen2-7B-Instruct with chunk size chunk8.
Each sample is tokenized and formatted with GSA gist tokens for continue pretraining.
Paper
GSA: Gist Sparse Attention via Learnable Compression and Selective Unfolding
Related Model
yuzhenm/GSA-PT-Qwen2-7B-Instruct-chunk8 — model trained on this dataset
GSA-PT-Qwen2-7B-Instruct-chunk8-chunk4-data
GSA-PT-Qwen2-7B-Instruct-chunk8-chunk4-data
This is the continue pretraining dataset used for training GSA (Gist Sparse Attention) models based on Qwen2-7B-Instruct with chunk size chunk8-chunk4.
Each sample is tokenized and formatted with GSA gist tokens for continue pretraining.
Paper
GSA: Gist Sparse Attention via Learnable Compression and Selective Unfolding
Related Model
yuzhenm/GSA-PT-Qwen2-7B-Instruct-chunk8-chunk4 — model trained on this dataset
tweets_pt_sentiment_analysis
Dataset Card for "tweets_pt_sentiment_analysis"
More Information needed
mmlu-pt-undergraduate-hardovos-intent-bench-speech-massive-pt-PT
ovos-intent-bench-speech-massive-pt-PT
An OVOS Plugin Arena benchmark repository. The arena's prediction runner
publishes each plugin's raw output on the speech-massive-pt-PT dataset here, one JSON-lines
file per plugin under predictions/<lang>/, and where a sample-set manifest
governs scoring it lives under sample_sets/. The public leaderboard at
https://openvoiceos.github.io/ovos-plugin-arena/ is computed from these rows,
and anyone can recompute an entry from them without… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-intent-bench-speech-massive-pt-PT.tweet_topic_ERNIE-4.5-0.3B-PT_vocab_2000_lastmozilla-common-voice-converted-to-parquet-pt
