datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
VeraCruz_PT-BR
Dataset Summary
The VeraCruz Dataset is a comprehensive collection of Portuguese language content, showcasing the linguistic and cultural diversity of of Portuguese-speaking regions. It includes around 190 million samples, organized by regional origin as indicated by URL metadata into primary categories. The primary categories are:
Portugal (PT): Samples with content URLs indicating a clear Portuguese origin.
Brazil (BR): Samples with content URLs indicating a clear Brazilian… See the full description on the dataset page: https://huggingface.co/datasets/bastao/VeraCruz_PT-BR.ruri-dataset-v2-ptWIP: 正式公開準備中
各データセットのライセンスは元データセットに従います。
c4-pt-randMore35M-part04-deduplicated-128000-no-digit-split-mask-train-15003771-lines
Dataset Card for "c4-pt-randMore35M-part04-deduplicated-128000-no-digit-split-mask-train-15003771-lines"
More Information needed
msmarco
Dataset Card for "msmarco"
More Information needed
Pt-Corpus-Instruct
Portuguese-Corpus Instruct
Dataset Summary
Portuguese-Corpus Instruct is a concatenation of several portions of Brazilian Portuguese datasets found in the Hub.
In a tokenized format, the dataset (uncompressed) weighs 80 GB and has approximately 6.2B tokens. This version of the corpus (Pt-Corpus-Instruct) includes several instances of conversational and general instructional data, allowing trained models to go through preference pre-training during their initial… See the full description on the dataset page: https://huggingface.co/datasets/nicholasKluge/Pt-Corpus-Instruct.llm_pt_leaderboard_resultsPT-HF500B
PT-HF500B (FinePhrase)
Overview
FinePhrase is a large-scale synthetic dataset designed for high-quality language modeling, reasoning, and instruction-following tasks. It transforms raw educational web data into structured, instruction-rich formats suitable for training advanced language models.
This dataset has been extensively used in the pre-training pipeline of TNSA models, including:
NGen-3
NGen-4
NGen-4-OW
It plays a critical role in improving reasoning ability… See the full description on the dataset page: https://huggingface.co/datasets/TNSA/PT-HF500B.qrecc-corpus
Dataset Card for "qrecc"
More Information needed
laion400m-ptmteb-pt-results
🇧🇷 MTEB-BR — Benchmark Results
Canonical results store for MTEB-BR, a native Brazilian-Portuguese text-embedding benchmark.
93 models · 22 native PT-BR tasks · 7 categories · no machine translation
What is this?
This repository is the canonical, machine-readable results store for MTEB-BR — a benchmark that evaluates text-embedding models on native Brazilian Portuguese (data created or found in Portuguese; machine-translated corpora such as… See the full description on the dataset page: https://huggingface.co/datasets/MTEB-BR/mteb-pt-results.sentence-pt-enzh-tr原数据集:https://huggingface.co/datasets/TigerResearch/pretrain_en、https://huggingface.co/datasets/TigerResearch/pretrain_zh
这个数据集就是把数据切断成句了,没想到行数居然接近1:1哦
中文数据总行数: 434818374
英文数据总行数: 450173542
中文比例: 0.4913, 英文比例: 0.5087
merge是中英混合后的版本(每个文件都有中英文且尽可能保持一样的比例后内部打乱,混训的直接按量节选即可),处理程序遗漏了最后几批数据,导致只有210个4M条的文件(69GB)。
imdb_ptLarge Movie Review Dataset.
This is a dataset for binary sentiment classification containing substantially more data than previous benchmark datasets. We provide a set of 25,000 highly polar movie reviews for training, and 25,000 for testing. There is additional unlabeled data for use as well.\mc4-pt
MC4-PT
MC4-PT is the is the portuguese subset from MC4.
MC4 is a multilingual colossal, cleaned version of Common Crawl's web crawl corpus. Based on Common Crawl dataset: "https://commoncrawl.org".
This is the raw version. Deduplicated version is available here.
alpaca-data-pt-brNOTE: This is a machine translated version of the yahma/alpaca-cleaned dataset.
Dataset Card for Alpaca-Cleaned
Repository: https://github.com/gururise/AlpacaDataCleaned
Dataset Description
This is a cleaned version of the original Alpaca Dataset released by Stanford. The following issues have been identified in the original release and fixed in this dataset:
Hallucinations: Many instructions in the original dataset had instructions referencing data on the internet… See the full description on the dataset page: https://huggingface.co/datasets/dominguesm/alpaca-data-pt-br.MME-Benchmark-pt
Avaliação - MME-Perception
Estrutura do Diretório
main
├── MME_Benchmark
│ ├── artwork
│ │ ├── images
│ │ │ ├── 1.jpg
│ │ │ ├── 2.jpg
│ │ │ ├── ...
│ │ ├── question_answers_YN
│ │ │ ├── 1.txt
│ │ │ ├── 2.txt
│ │ │ ├── ...
│ ├── celebrity
│ ├── code_reasoning
│ ├── ...
├── calculation.py
├── translate_MME
Estrutura dos Arquivos TXT
Cada arquivo num.txt contém as perguntas correspondentes à imagem num.jpg.… See the full description on the dataset page: https://huggingface.co/datasets/LucasLima/MME-Benchmark-pt.coco-captions-pt-br
🎉 COCO Captions Dataset Translation for Portuguese Image Captioning
💾 Dataset Summary
COCO Captions Portuguese Translation, a multimodal dataset for Portuguese image captioning with 123,287 images, each accompanied by five descriptive captions that have been
generated by human annotators for every individual image. The original English captions were rendered into Portuguese
through the utilization of the Google Translator API.
🧑💻 Hot to Get… See the full description on the dataset page: https://huggingface.co/datasets/laicsiifes/coco-captions-pt-br.ClassiCC-PT
📚 ClassiCC-PT: Classified Common Crawl Corpus for Portuguese
📖 Overview
ClassiCC-PT (Classified Common Crawl – Portuguese) is a large-scale web corpus containing ~120B Portuguese tokens extracted from Common Crawl snapshots. It is specifically curated for training large language models in Portuguese, with a focus on data quality, language specificity, and targeted filtering.
This corpus was created as part of a study on continued pretraining for adapting English-trained… See the full description on the dataset page: https://huggingface.co/datasets/ClassiCC-Corpus/ClassiCC-PT.in1k_clip_qwen25vl_3b_224res_64tokens_new_ptcc100-pt
C100-PT
CC100-PT is the is the portuguese subset from C100.
C100 was created for training the multilingual Transformer XLM-R, containing two terabytes of cleaned data from 2018 snapshots of the Common Crawl project in 100 languages.
pt_mergeBIGstockimage-1.5M-scored-pt-twonatural-questions-nci
Dataset Card for "natural-questions-nci"
More Information needed
mmlu_ptin1k_clip_qwen25vl_3b_448res_256tokens_new_merged_ptJiuZhang3.0-Corpus-PT-CoTpt_textBIGstockimage-1.5M-scored-pt-oneimaginative-perception-token-pt-eval-ai2thor
Citation
Released with the paper Imaginative Perception Tokens Enhance Spatial Reasoning in Multimodal Language Models (arXiv:2606.03988):
@misc{bigverdi2026imaginativeperceptiontokensenhance,
title={Imaginative Perception Tokens Enhance Spatial Reasoning in Multimodal Language Models},
author={Mahtab Bigverdi and Linjie Li and Weikai Huang and Yiming Liu and Jaemin Cho and Jieyu Zhang and Tuhin Kundu and Chris Dangjoo Kim and Zelun Luo and Linda Shapiro and Ranjay… See the full description on the dataset page: https://huggingface.co/datasets/weikaih/imaginative-perception-token-pt-eval-ai2thor.cv24-pt-128-normalizedllm0to1-pt-aihub-624
AI-Hub 웹데이터 기반 한국어 말뭉치 (전처리)
LLM0to1-10b (SmolLM3 기반 10B 한/영 이중언어 LLM) 사전학습 코퍼스의 일부. huggingface.co/izlley
개요
카테고리: korean
원본 출처: AI-Hub 데이터셋 624 (웹데이터 기반 한국어 말뭉치)
라이선스: AI-Hub 이용약관(재배포 허가 확인)
토큰 수(우리 토크나이저 vocab 160k): 3.777B / 문서 120,000건
600B 믹스 내 역할: korean 카테고리(목표 25% = 150B). 카테고리 unique 62.5B 중 이 소스 6.0%(~9.07B 기여), 카테고리 전체 약 2.40 epoch 반복
전처리·필터링
zip 스트리밍 추출→한글비율≥0.25·길이≥40 필터→문서 exact-dedup(md5)→PII 스크럽(주민번호·전화·이메일)
토크나이저:… See the full description on the dataset page: https://huggingface.co/datasets/izlley/llm0to1-pt-aihub-624.
