CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01TNSA /PT-HF500B PT-HF500B (FinePhrase) Overview FinePhrase is a large-scale synthetic dataset designed for high-quality language modeling, reasoning, and instruction-following tasks. It transforms raw educational web data into structured, instruction-rich formats suitable for training advanced language models. This dataset has been extensively used in the pre-training pipeline of TNSA models, including: NGen-3 NGen-4 NGen-4-OW It plays a critical role in improving reasoning ability… See the full description on the dataset page: https://huggingface.co/datasets/TNSA/PT-HF500B.tabulartext-generation1B<n<10B1 likes1.7k downloads6mo agoHugging Face02candido-ai /laion400m-ptimage100M<n<1B0 likes1.2k downloads2y agoHugging Face03MTEB-BR /mteb-pt-results 🇧🇷 MTEB-BR — Benchmark Results Canonical results store for MTEB-BR, a native Brazilian-Portuguese text-embedding benchmark. 93 models · 22 native PT-BR tasks · 7 categories · no machine translation What is this? This repository is the canonical, machine-readable results store for MTEB-BR — a benchmark that evaluates text-embedding models on native Brazilian Portuguese (data created or found in Portuguese; machine-translated corpora such as… See the full description on the dataset page: https://huggingface.co/datasets/MTEB-BR/mteb-pt-results.tabularfeature-extractionn<1K0 likes1.1k downloads2mo agoHugging Face04ClassiCC-Corpus /ClassiCC-PT 📚 ClassiCC-PT: Classified Common Crawl Corpus for Portuguese 📖 Overview ClassiCC-PT (Classified Common Crawl – Portuguese) is a large-scale web corpus containing ~120B Portuguese tokens extracted from Common Crawl snapshots. It is specifically curated for training large language models in Portuguese, with a focus on data quality, language specificity, and targeted filtering. This corpus was created as part of a study on continued pretraining for adapting English-trained… See the full description on the dataset page: https://huggingface.co/datasets/ClassiCC-Corpus/ClassiCC-PT.tabular10M<n<100M15 likes556 downloads8mo agoHugging Face05hanlincs /in1k_clip_qwen25vl_3b_224res_64tokens_new_pttabular1M<n<10M0 likes515 downloads1y agoHugging Face06hanlincs /in1k_clip_qwen25vl_3b_448res_256tokens_new_merged_pttabular1M<n<10M0 likes470 downloads1y agoHugging Face07gist-sparse-attention /GSA-PT-Qwen2-7B-Instruct-chunk4-chunk4-data GSA-PT-Qwen2-7B-Instruct-chunk4-chunk4-data This is the continue pretraining dataset used for training GSA (Gist Sparse Attention) models based on Qwen2-7B-Instruct with chunk size chunk4-chunk4. Each sample is tokenized and formatted with GSA gist tokens for continue pretraining. Paper GSA: Gist Sparse Attention via Learnable Compression and Selective Unfolding Related Model yuzhenm/GSA-PT-Qwen2-7B-Instruct-chunk4-chunk4 — model trained on this dataset tabular10K<n<100K0 likes357 downloads6mo agoHugging Face08hbseong /eval_record-pick-and-place-pos5-so101_pt-ft-3epThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "so101_follower", "total_episodes": 36, "total_frames": 43756, "total_tasks": 2, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 500, "fps": 30, "splits": { "train": "0:36" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/hbseong/eval_record-pick-and-place-pos5-so101_pt-ft-3ep.tabularrobotics10K<n<100K0 likes328 downloads10mo agoHugging Face09hbseong /eval_record-pick-and-place-so101_pt-ft-3epThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "so101_follower", "total_episodes": 94, "total_frames": 105687, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 500, "fps": 30, "splits": { "train": "0:94" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/hbseong/eval_record-pick-and-place-so101_pt-ft-3ep.tabularrobotics100K<n<1M0 likes281 downloads10mo agoHugging Face10unlearning-cleanslate /generations-simnpo_gemma-3-12b-pt_20260416_171305-corpus_sweep_post_evaltabular10K<n<100K0 likes235 downloads5mo agoHugging Face11OpenVoiceOS /ovos-stt-bench-mls-pt-PT OVOS stt bench — mls-pt-PT Per-clip transcripts predictions of the registered OVOS Plugin Arena stt fighters over facebook/multilingual_librispeech. One dedicated repo per modality; one dataset split per language; one JSONL file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow the arena §3.2 contract (pinned dataset_revision, plugin_version, latency_ms). Produced by the reproducible benchmark script in the arena repo; the arena's assemble workflow turns… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-mls-pt-PT.tabularn<1K0 likes225 downloads16d agoHugging Face12DedeProGames /claude-code-traces-pt-brThis dataset was generated using teich by TeichAI claude Agent Traces This directory contains raw agent trace files generated by teich. JSONL files: 20 Training-ready tools Generated agent traces carry configured or recovered tool schemas so tools remain available for training even when a session did not call them. Native Claude Code imports recover schemas for Claude Code and Claude Desktop built-ins, plus conservative name-derived MCP schemas, when the raw… See the full description on the dataset page: https://huggingface.co/datasets/DedeProGames/claude-code-traces-pt-br.tabulartext-generationn<1K2 likes207 downloads2mo agoHugging Face13mteb-pt /mteb-pt-results Note: This dataset has moved to MTEB-BR/mteb-pt-results. This legacy copy remains only to preserve its archival DOI (10.57967/hf/9377); the maintained version lives under the MTEB-BR organization. 🇧🇷 MTEB-BR — Benchmark Results Canonical results store for MTEB-BR, a native Brazilian-Portuguese text-embedding benchmark. 93 models · 22 native PT-BR tasks · 7 categories · no machine translation What is this? This repository is the canonical… See the full description on the dataset page: https://huggingface.co/datasets/mteb-pt/mteb-pt-results.tabularfeature-extractionn<1K1 likes180 downloads3mo agoHugging Face14gist-sparse-attention /GSA-PT-Qwen2-7B-Instruct-chunk16-data GSA-PT-Qwen2-7B-Instruct-chunk16-data This is the continue pretraining dataset used for training GSA (Gist Sparse Attention) models based on Qwen2-7B-Instruct with chunk size chunk16. Each sample is tokenized and formatted with GSA gist tokens for continue pretraining. Paper GSA: Gist Sparse Attention via Learnable Compression and Selective Unfolding Related Model yuzhenm/GSA-PT-Qwen2-7B-Instruct-chunk16 — model trained on this dataset tabular10K<n<100K0 likes162 downloads6mo agoHugging Face15amalia-llm /pt_exams PHEB - Portuguese High School Exams MCQ MCQ set of PHEB a collection of Portuguese exam questions for evaluating language models on academic knowledge on the Portuguese curriculum. For more details, see the PHEB paper. This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite for evaluating large language models on European Portuguese. Citation If you use this dataset or AMALIA in your work… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/pt_exams.tabularquestion-answering1K<n<10K0 likes160 downloads3mo agoHugging Face16AKCIT /ToxSyn-PT Dataset Summary ToxSyn-PT is a large-scale synthetic dataset designed for fine-grained hate speech detection in Brazilian Portuguese. It comprises 53,274 sentences equally balanced between toxic and non-toxic labels, covering nine legally protected minority groups (including Black, Women, LGBTQIA+, Native Brazilian, Muslim, Jewish, and Elderly). Unlike most existing datasets that only capture hostile mentions, ToxSyn-PT systematically includes non-toxic counterexamples (benign… See the full description on the dataset page: https://huggingface.co/datasets/AKCIT/ToxSyn-PT.tabular10K<n<100K3 likes138 downloads4mo agoHugging Face17OpenVoiceOS /ovos-vad-bench-speech-vs-nonspeech-pt-PT OVOS vad bench — speech-vs-nonspeech-pt-PT Per-clip speech / non-speech decisions predictions of the registered OVOS Plugin Arena vad fighters over PolyAI/minds14. One dedicated repo per modality; one dataset split per language; one JSONL file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow the arena §3.2 contract (pinned dataset_revision, plugin_version, latency_ms). Produced by the reproducible benchmark script in the arena repo; the arena's assemble… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-vad-bench-speech-vs-nonspeech-pt-PT.tabularn<1K0 likes117 downloads17d agoHugging Face18hbseong /eval_record-pick-and-place-ez2-so101_pt-ft-3epThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "so101_follower", "total_episodes": 21, "total_frames": 13660, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 500, "fps": 30, "splits": { "train": "0:21" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/hbseong/eval_record-pick-and-place-ez2-so101_pt-ft-3ep.tabularrobotics10K<n<100K0 likes111 downloads10mo agoHugging Face19PORTULAN /glue-ptptGLUE-PTPT is an European Portuguese translation of the GLUE benchmark using DeepL Pro.tabular10K<n<100K6 likes93 downloads3y agoHugging Face20OpenVoiceOS /ovos-stt-bench-vocatives-pt-PT OVOS stt bench — vocatives-pt-PT Per-clip transcripts predictions of the registered OVOS Plugin Arena stt fighters over Jarbas/VocativesEuropeanPortuguese. One dedicated repo per modality; one dataset split per language; one JSONL file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow the arena §3.2 contract (pinned dataset_revision, plugin_version, latency_ms). Produced by the reproducible benchmark script in the arena repo; the arena's assemble workflow… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-vocatives-pt-PT.tabularn<1K0 likes93 downloads23d agoHugging Face21OdiaGenAIdata /fine_web2_odia_pttabular1M<n<10M0 likes91 downloads2y agoHugging Face22franklinbaldo /multieurlex21-pt-semantic-cache MultiEURLEX-21 PT — frozen semantic chunk embeddings Precomputed, unit-normalised chunk embeddings of the official MTEB MultiEURLEXMultilabelClassification Portuguese data, produced once so that downstream experiments (frozen MaleCNS connectome reservoir, label-free controls, ablations) never pay the sentence encoders again. No benchmark labels are stored or used here. source dataset: mteb/eurlex-multilingual revision 2aea5a6dc8fdcfeca41d0fb963c0a338930bde5c, subset pt splits:… See the full description on the dataset page: https://huggingface.co/datasets/franklinbaldo/multieurlex21-pt-semantic-cache.tabulartext-classification100K<n<1M0 likes81 downloads9d agoHugging Face23mmlu-pt /mmlu-pt-undergraduate-fulltabular10K<n<100K0 likes79 downloads3mo agoHugging Face24gist-sparse-attention /GSA-PT-Qwen2-7B-Instruct-chunk8-data GSA-PT-Qwen2-7B-Instruct-chunk8-data This is the continue pretraining dataset used for training GSA (Gist Sparse Attention) models based on Qwen2-7B-Instruct with chunk size chunk8. Each sample is tokenized and formatted with GSA gist tokens for continue pretraining. Paper GSA: Gist Sparse Attention via Learnable Compression and Selective Unfolding Related Model yuzhenm/GSA-PT-Qwen2-7B-Instruct-chunk8 — model trained on this dataset tabular10K<n<100K0 likes76 downloads6mo agoHugging Face25gist-sparse-attention /GSA-PT-Qwen2-7B-Instruct-chunk8-chunk4-data GSA-PT-Qwen2-7B-Instruct-chunk8-chunk4-data This is the continue pretraining dataset used for training GSA (Gist Sparse Attention) models based on Qwen2-7B-Instruct with chunk size chunk8-chunk4. Each sample is tokenized and formatted with GSA gist tokens for continue pretraining. Paper GSA: Gist Sparse Attention via Learnable Compression and Selective Unfolding Related Model yuzhenm/GSA-PT-Qwen2-7B-Instruct-chunk8-chunk4 — model trained on this dataset tabular10K<n<100K0 likes75 downloads6mo agoHugging Face26ricardo-filho /tweets_pt_sentiment_analysis Dataset Card for "tweets_pt_sentiment_analysis" More Information needed tabular100K<n<1M1 likes73 downloads4y agoHugging Face27mmlu-pt /mmlu-pt-undergraduate-hardtabular10K<n<100K0 likes72 downloads3mo agoHugging Face28OpenVoiceOS /ovos-intent-bench-speech-massive-pt-PT ovos-intent-bench-speech-massive-pt-PT An OVOS Plugin Arena benchmark repository. The arena's prediction runner publishes each plugin's raw output on the speech-massive-pt-PT dataset here, one JSON-lines file per plugin under predictions/<lang>/, and where a sample-set manifest governs scoring it lives under sample_sets/. The public leaderboard at https://openvoiceos.github.io/ovos-plugin-arena/ is computed from these rows, and anyone can recompute an entry from them without… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-intent-bench-speech-massive-pt-PT.tabularn<1K0 likes72 downloads20d agoHugging Face29raymondzmc /tweet_topic_ERNIE-4.5-0.3B-PT_vocab_2000_lasttabular10K<n<100K0 likes66 downloads9mo agoHugging Face30MatheusMarquesEiras /mozilla-common-voice-converted-to-parquet-pttabular10K<n<100K0 likes65 downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.