CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01bastao /VeraCruz_PT-BR Dataset Summary The VeraCruz Dataset is a comprehensive collection of Portuguese language content, showcasing the linguistic and cultural diversity of of Portuguese-speaking regions. It includes around 190 million samples, organized by regional origin as indicated by URL metadata into primary categories. The primary categories are: Portugal (PT): Samples with content URLs indicating a clear Portuguese origin. Brazil (BR): Samples with content URLs indicating a clear Brazilian… See the full description on the dataset page: https://huggingface.co/datasets/bastao/VeraCruz_PT-BR.texttext-generation100M<n<1B17 likes60k downloads1y agoHugging Face02cl-nagoya /ruri-dataset-v2-ptWIP: 正式公開準備中 各データセットのライセンスは元データセットに従います。 text100M<n<1B5 likes8.3k downloads2y agoHugging Face03tgsc /c4-pt-randMore35M-part04-deduplicated-128000-no-digit-split-mask-train-15003771-lines Dataset Card for "c4-pt-randMore35M-part04-deduplicated-128000-no-digit-split-mask-train-15003771-lines" More Information needed text10M<n<100M1 likes4.7k downloads3y agoHugging Face04namespace-Pt /msmarco Dataset Card for "msmarco" More Information needed text1K<n<10K2 likes2.6k downloads3y agoHugging Face05nicholasKluge /Pt-Corpus-Instruct Portuguese-Corpus Instruct Dataset Summary Portuguese-Corpus Instruct is a concatenation of several portions of Brazilian Portuguese datasets found in the Hub. In a tokenized format, the dataset (uncompressed) weighs 80 GB and has approximately 6.2B tokens. This version of the corpus (Pt-Corpus-Instruct) includes several instances of conversational and general instructional data, allowing trained models to go through preference pre-training during their initial… See the full description on the dataset page: https://huggingface.co/datasets/nicholasKluge/Pt-Corpus-Instruct.texttext-generation10M<n<100M3 likes2k downloads2y agoHugging Face06eduagarcia-temp /llm_pt_leaderboard_resultstextn<1K0 likes1.9k downloads1y agoHugging Face07TNSA /PT-HF500B PT-HF500B (FinePhrase) Overview FinePhrase is a large-scale synthetic dataset designed for high-quality language modeling, reasoning, and instruction-following tasks. It transforms raw educational web data into structured, instruction-rich formats suitable for training advanced language models. This dataset has been extensively used in the pre-training pipeline of TNSA models, including: NGen-3 NGen-4 NGen-4-OW It plays a critical role in improving reasoning ability… See the full description on the dataset page: https://huggingface.co/datasets/TNSA/PT-HF500B.tabulartext-generation1B<n<10B1 likes1.7k downloads6mo agoHugging Face08namespace-Pt /qrecc-corpus Dataset Card for "qrecc" More Information needed text10M<n<100M2 likes1.2k downloads3y agoHugging Face09candido-ai /laion400m-ptimage100M<n<1B0 likes1.2k downloads2y agoHugging Face10MTEB-BR /mteb-pt-results 🇧🇷 MTEB-BR — Benchmark Results Canonical results store for MTEB-BR, a native Brazilian-Portuguese text-embedding benchmark. 93 models · 22 native PT-BR tasks · 7 categories · no machine translation What is this? This repository is the canonical, machine-readable results store for MTEB-BR — a benchmark that evaluates text-embedding models on native Brazilian Portuguese (data created or found in Portuguese; machine-translated corpora such as… See the full description on the dataset page: https://huggingface.co/datasets/MTEB-BR/mteb-pt-results.tabularfeature-extractionn<1K0 likes1.1k downloads2mo agoHugging Face11nwdxlgzs /sentence-pt-enzh-tr原数据集:https://huggingface.co/datasets/TigerResearch/pretrain_en、https://huggingface.co/datasets/TigerResearch/pretrain_zh 这个数据集就是把数据切断成句了,没想到行数居然接近1:1哦 中文数据总行数: 434818374 英文数据总行数: 450173542 中文比例: 0.4913, 英文比例: 0.5087 merge是中英混合后的版本(每个文件都有中英文且尽可能保持一样的比例后内部打乱,混训的直接按量节选即可),处理程序遗漏了最后几批数据,导致只有210个4M条的文件(69GB)。 text1B<n<10B0 likes942 downloads1y agoHugging Face12maritaca-ai /imdb_ptLarge Movie Review Dataset. This is a dataset for binary sentiment classification containing substantially more data than previous benchmark datasets. We provide a set of 25,000 highly polar movie reviews for training, and 25,000 for testing. There is additional unlabeled data for use as well.\text10K<n<100K5 likes861 downloads3y agoHugging Face13eduagarcia /mc4-pt MC4-PT MC4-PT is the is the portuguese subset from MC4. MC4 is a multilingual colossal, cleaned version of Common Crawl's web crawl corpus. Based on Common Crawl dataset: "https://commoncrawl.org". This is the raw version. Deduplicated version is available here. text100M<n<1B2 likes614 downloads3y agoHugging Face14dominguesm /alpaca-data-pt-brNOTE: This is a machine translated version of the yahma/alpaca-cleaned dataset. Dataset Card for Alpaca-Cleaned Repository: https://github.com/gururise/AlpacaDataCleaned Dataset Description This is a cleaned version of the original Alpaca Dataset released by Stanford. The following issues have been identified in the original release and fixed in this dataset: Hallucinations: Many instructions in the original dataset had instructions referencing data on the internet… See the full description on the dataset page: https://huggingface.co/datasets/dominguesm/alpaca-data-pt-br.texttext-generation10K<n<100K35 likes593 downloads3y agoHugging Face15LucasLima /MME-Benchmark-pt Avaliação - MME-Perception Estrutura do Diretório main ├── MME_Benchmark │ ├── artwork │ │ ├── images │ │ │ ├── 1.jpg │ │ │ ├── 2.jpg │ │ │ ├── ... │ │ ├── question_answers_YN │ │ │ ├── 1.txt │ │ │ ├── 2.txt │ │ │ ├── ... │ ├── celebrity │ ├── code_reasoning │ ├── ... ├── calculation.py ├── translate_MME Estrutura dos Arquivos TXT Cada arquivo num.txt contém as perguntas correspondentes à imagem num.jpg.… See the full description on the dataset page: https://huggingface.co/datasets/LucasLima/MME-Benchmark-pt.imagen<1K0 likes589 downloads2y agoHugging Face16laicsiifes /coco-captions-pt-br 🎉 COCO Captions Dataset Translation for Portuguese Image Captioning 💾 Dataset Summary COCO Captions Portuguese Translation, a multimodal dataset for Portuguese image captioning with 123,287 images, each accompanied by five descriptive captions that have been generated by human annotators for every individual image. The original English captions were rendered into Portuguese through the utilization of the Google Translator API. 🧑‍💻 Hot to Get… See the full description on the dataset page: https://huggingface.co/datasets/laicsiifes/coco-captions-pt-br.imagetext-to-image100K<n<1M6 likes573 downloads4mo agoHugging Face17ClassiCC-Corpus /ClassiCC-PT 📚 ClassiCC-PT: Classified Common Crawl Corpus for Portuguese 📖 Overview ClassiCC-PT (Classified Common Crawl – Portuguese) is a large-scale web corpus containing ~120B Portuguese tokens extracted from Common Crawl snapshots. It is specifically curated for training large language models in Portuguese, with a focus on data quality, language specificity, and targeted filtering. This corpus was created as part of a study on continued pretraining for adapting English-trained… See the full description on the dataset page: https://huggingface.co/datasets/ClassiCC-Corpus/ClassiCC-PT.tabular10M<n<100M15 likes556 downloads8mo agoHugging Face18hanlincs /in1k_clip_qwen25vl_3b_224res_64tokens_new_pttabular1M<n<10M0 likes515 downloads1y agoHugging Face19eduagarcia /cc100-pt C100-PT CC100-PT is the is the portuguese subset from C100. C100 was created for training the multilingual Transformer XLM-R, containing two terabytes of cleaned data from 2018 snapshots of the Common Crawl project in 100 languages. texttext-generation10M<n<100M1 likes512 downloads3y agoHugging Face20Huy227 /pt_mergetext10M<n<100M0 likes508 downloads2y agoHugging Face21ljnlonoljpiljm /BIGstockimage-1.5M-scored-pt-twoimage100K<n<1M0 likes493 downloads1y agoHugging Face22namespace-Pt /natural-questions-nci Dataset Card for "natural-questions-nci" More Information needed text100K<n<1M2 likes485 downloads3y agoHugging Face23neuralmagic /mmlu_pttext10K<n<100K0 likes481 downloads2y agoHugging Face24hanlincs /in1k_clip_qwen25vl_3b_448res_256tokens_new_merged_pttabular1M<n<10M0 likes470 downloads1y agoHugging Face25ToheartZhang /JiuZhang3.0-Corpus-PT-CoTtext1M<n<10M9 likes456 downloads2y agoHugging Face26Huy227 /pt_texttext1M<n<10M0 likes424 downloads2y agoHugging Face27ljnlonoljpiljm /BIGstockimage-1.5M-scored-pt-oneimage100K<n<1M1 likes422 downloads1y agoHugging Face28weikaih /imaginative-perception-token-pt-eval-ai2thor Citation Released with the paper Imaginative Perception Tokens Enhance Spatial Reasoning in Multimodal Language Models (arXiv:2606.03988): @misc{bigverdi2026imaginativeperceptiontokensenhance, title={Imaginative Perception Tokens Enhance Spatial Reasoning in Multimodal Language Models}, author={Mahtab Bigverdi and Linjie Li and Weikai Huang and Yiming Liu and Jaemin Cho and Jieyu Zhang and Tuhin Kundu and Chris Dangjoo Kim and Zelun Luo and Linda Shapiro and Ranjay… See the full description on the dataset page: https://huggingface.co/datasets/weikaih/imaginative-perception-token-pt-eval-ai2thor.imagen<1K0 likes413 downloads4mo agoHugging Face29N03N9 /cv24-pt-128-normalizedtext100K<n<1M0 likes391 downloads9mo agoHugging Face30izlley /llm0to1-pt-aihub-624 AI-Hub 웹데이터 기반 한국어 말뭉치 (전처리) LLM0to1-10b (SmolLM3 기반 10B 한/영 이중언어 LLM) 사전학습 코퍼스의 일부. huggingface.co/izlley 개요 카테고리: korean 원본 출처: AI-Hub 데이터셋 624 (웹데이터 기반 한국어 말뭉치) 라이선스: AI-Hub 이용약관(재배포 허가 확인) 토큰 수(우리 토크나이저 vocab 160k): 3.777B / 문서 120,000건 600B 믹스 내 역할: korean 카테고리(목표 25% = 150B). 카테고리 unique 62.5B 중 이 소스 6.0%(~9.07B 기여), 카테고리 전체 약 2.40 epoch 반복 전처리·필터링 zip 스트리밍 추출→한글비율≥0.25·길이≥40 필터→문서 exact-dedup(md5)→PII 스크럽(주민번호·전화·이메일) 토크나이저:… See the full description on the dataset page: https://huggingface.co/datasets/izlley/llm0to1-pt-aihub-624.text10K<n<100K0 likes383 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.