CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01eduagarcia-temp /llm_pt_leaderboard_resultstextn<1K0 likes1.9k downloads1y agoHugging Face02izlley /llm0to1-pt-aihub-624 AI-Hub 웹데이터 기반 한국어 말뭉치 (전처리) LLM0to1-10b (SmolLM3 기반 10B 한/영 이중언어 LLM) 사전학습 코퍼스의 일부. huggingface.co/izlley 개요 카테고리: korean 원본 출처: AI-Hub 데이터셋 624 (웹데이터 기반 한국어 말뭉치) 라이선스: AI-Hub 이용약관(재배포 허가 확인) 토큰 수(우리 토크나이저 vocab 160k): 3.777B / 문서 120,000건 600B 믹스 내 역할: korean 카테고리(목표 25% = 150B). 카테고리 unique 62.5B 중 이 소스 6.0%(~9.07B 기여), 카테고리 전체 약 2.40 epoch 반복 전처리·필터링 zip 스트리밍 추출→한글비율≥0.25·길이≥40 필터→문서 exact-dedup(md5)→PII 스크럽(주민번호·전화·이메일) 토크나이저:… See the full description on the dataset page: https://huggingface.co/datasets/izlley/llm0to1-pt-aihub-624.text10K<n<100K0 likes383 downloads2mo agoHugging Face03thegoodfellas /mc4-pt-cleaned Description This is a clenned version of AllenAI mC4 PtBR section. The original dataset can be found here https://huggingface.co/datasets/allenai/c4 Clean procedure We applied the same clenning procedure as explained here: https://gitlab.com/yhavinga/c4nlpreproc.git The repository offers two strategies. The first one, found in the main.py file, uses pyspark to create a dataframe that can both clean the text and create a pseudo mix on the entire dataset. We found this… See the full description on the dataset page: https://huggingface.co/datasets/thegoodfellas/mc4-pt-cleaned.textfill-mask100M<n<1B4 likes342 downloads3y agoHugging Face04NOVA-vision-language /calame-pt CALAME-PT Context-Aware LAnguage Modeling Evaluation for Portuguese CALAME-PT is a PT benchmark composed of small texts (contexts) and their respective last words. These contexts should, in theory, contain enough information so that a human or a model is capable of guessing its last word - without being too specific and/or too ambiguous. Composition CALAME-PT is composed of 2 "sets" of data - handwritten and generated. Handwritten Set: contains 406… See the full description on the dataset page: https://huggingface.co/datasets/NOVA-vision-language/calame-pt.text1K<n<10K2 likes337 downloads3y agoHugging Face05OpenVoiceOS /ovos-stt-bench-mls-pt-PT OVOS stt bench — mls-pt-PT Per-clip transcripts predictions of the registered OVOS Plugin Arena stt fighters over facebook/multilingual_librispeech. One dedicated repo per modality; one dataset split per language; one JSONL file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow the arena §3.2 contract (pinned dataset_revision, plugin_version, latency_ms). Produced by the reproducible benchmark script in the arena repo; the arena's assemble workflow turns… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-mls-pt-PT.tabularn<1K0 likes225 downloads17d agoHugging Face06DedeProGames /claude-code-traces-pt-brThis dataset was generated using teich by TeichAI claude Agent Traces This directory contains raw agent trace files generated by teich. JSONL files: 20 Training-ready tools Generated agent traces carry configured or recovered tool schemas so tools remain available for training even when a session did not call them. Native Claude Code imports recover schemas for Claude Code and Claude Desktop built-ins, plus conservative name-derived MCP schemas, when the raw… See the full description on the dataset page: https://huggingface.co/datasets/DedeProGames/claude-code-traces-pt-br.tabulartext-generationn<1K2 likes207 downloads2mo agoHugging Face07amalia-llm /bigbenchhard-mt-pt BBH-PT (Big-Bench Hard) Portuguese machine translation of BIG-Bench Hard, a challenging subset of the BIG-Bench benchmark covering diverse reasoning tasks. Translated using a Finetuned GemmaX2-9B for pt-PT with rule-based adaptations. Note: Some tasks (e.g., hyperbaton) are not translated as they do not transfer meaningfully to Portuguese. Original Dataset: https://github.com/suzgunmirac/BIG-Bench-Hard Note: This dataset is machine translated and may contain… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/bigbenchhard-mt-pt.textquestion-answering1K<n<10K0 likes186 downloads3mo agoHugging Face08netcat420 /Kayla-ptpytorch tensor version of Kayla! made with love by netcat420 :3 netcat7 <-- discord, add me! :P text1K<n<10K0 likes176 downloads11mo agoHugging Face09amalia-llm /pt_exams PHEB - Portuguese High School Exams MCQ MCQ set of PHEB a collection of Portuguese exam questions for evaluating language models on academic knowledge on the Portuguese curriculum. For more details, see the PHEB paper. This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite for evaluating large language models on European Portuguese. Citation If you use this dataset or AMALIA in your work… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/pt_exams.tabularquestion-answering1K<n<10K0 likes160 downloads3mo agoHugging Face10pinzhenchen /alpaca-cleaned-pt Data Description This HF data repository contains the Portuguese Alpaca dataset used in our study of monolingual versus multilingual instruction tuning. GitHub Paper Creation Machine-translated from yahma/alpaca-cleaned into Portuguese. Usage This data is intended to be used for Portuguese instruction tuning. The dataset has roughly 52K instances in the JSON format. Each instance has an instruction, an output, and an optional input. An example is shown… See the full description on the dataset page: https://huggingface.co/datasets/pinzhenchen/alpaca-cleaned-pt.texttext-generation10K<n<100K5 likes155 downloads3y agoHugging Face11SetFit /amazon_massive_intent_pt-PTtext10K<n<100K1 likes142 downloads4y agoHugging Face12joshycodes /sorrel-T-gemma-3-27b-pt-seed0-documentstext100K<n<1M0 likes139 downloads7d agoHugging Face13pt-sk /Maths-Grade-SchoolMaths-Grade-School I am releasing large Grade School level Mathematics datatset. This extensive dataset, comprising nearly one million instructions in JSON format, encapsulates a diverse array of topics fundamental to building a strong mathematical foundation. This dataset is in instruction format so that model developers, researchers etc. can easily use this dataset. Following Fields & sub Fields are covered: Calculus Probability Algebra Liner Algebra Trigonometry Differential Equations… See the full description on the dataset page: https://huggingface.co/datasets/pt-sk/Maths-Grade-School.texttext-generation100K<n<1M2 likes126 downloads2y agoHugging Face14tiagoteixeira03 /MATH-PT Math-PT: A Math Reasoning Benchmark for European and Brazilian Portuguese Math-PT is a high-quality evaluation dataset designed to measure the mathematical reasoning capabilities of Large Language Models (LLMs) in Portuguese. Unlike many existing benchmarks that rely on English translations, Math-PT uses native-language problems sourced from prestigious academic competitions and national exams in both Portugal and Brazil. Dataset Details Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/tiagoteixeira03/MATH-PT.textquestion-answering1K<n<10K1 likes124 downloads6mo agoHugging Face15pt-sk /Vision-COTtext10K<n<100K0 likes118 downloads2y agoHugging Face16OpenVoiceOS /ovos-vad-bench-speech-vs-nonspeech-pt-PT OVOS vad bench — speech-vs-nonspeech-pt-PT Per-clip speech / non-speech decisions predictions of the registered OVOS Plugin Arena vad fighters over PolyAI/minds14. One dedicated repo per modality; one dataset split per language; one JSONL file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow the arena §3.2 contract (pinned dataset_revision, plugin_version, latency_ms). Produced by the reproducible benchmark script in the arena repo; the arena's assemble… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-vad-bench-speech-vs-nonspeech-pt-PT.tabularn<1K0 likes117 downloads17d agoHugging Face17JigSawPT /ptpt-failure-set-gate Where a local 27B actually breaks against a frontier model — a European-Portuguese failure-set gate On broad everyday tasks, a clean local 27B is near-indistinguishable from a frontier model under blind judging. The gaps that remain are narrow, behavioral, and regex-detectable — which is exactly what small adapters fix. This dataset is the measurement instrument: six hard-sets with deterministic checks, plus the scorer and the methodology write-up. Dataset summary… See the full description on the dataset page: https://huggingface.co/datasets/JigSawPT/ptpt-failure-set-gate.texttext-generationn<1K0 likes96 downloads3mo agoHugging Face18OpenVoiceOS /ovos-stt-bench-vocatives-pt-PT OVOS stt bench — vocatives-pt-PT Per-clip transcripts predictions of the registered OVOS Plugin Arena stt fighters over Jarbas/VocativesEuropeanPortuguese. One dedicated repo per modality; one dataset split per language; one JSONL file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow the arena §3.2 contract (pinned dataset_revision, plugin_version, latency_ms). Produced by the reproducible benchmark script in the arena repo; the arena's assemble workflow… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-vocatives-pt-PT.tabularn<1K0 likes93 downloads23d agoHugging Face19fabiogr /social_i_qa_pt SocialIQa dataset v1.4 (PT) This is translation to the portuguese language of the dataset allenai/social_i_qa.Translations were done using three independent models: Helsinki-NLP/opus-mt-tc-big-en-pt unicamp-dl/translation-en-pt-t5 facebook/nllb-200-distilled-1.3B Translations were evaluated using the evaluation metric GEMBA - GPT Estimation Metric Based Assessment (from the article Large Language Models Are State-of-the-Art Evaluators of Translation Quality) using… See the full description on the dataset page: https://huggingface.co/datasets/fabiogr/social_i_qa_pt.text10K<n<100K1 likes89 downloads2y agoHugging Face20NOVA-vision-language /MSCOCO_PT-BRtextn<1K2 likes81 downloads2y agoHugging Face21costadev00 /wikipedia-pt-br-instruct-5k Wikipedia PT-BR Instruct wikipedia-pt-br-instruct is a synthetic supervised fine-tuning (SFT) dataset in Brazilian Portuguese generated from Wikipedia-derived documents. This release is an intermediate evaluation dataset produced with the sft-dataset-creator pipeline from the run wiki-ptbr-extract-calib-5kdocs-14tasks. It was generated from a fixed revision of costadev00/wikipedia-pt-br-extract: cdbd07dc4a3de6e64632c718710b3ae0ebaeb0ff The dataset is intended for intermediate… See the full description on the dataset page: https://huggingface.co/datasets/costadev00/wikipedia-pt-br-instruct-5k.texttext-generation100K<n<1M1 likes81 downloads3mo agoHugging Face22OpenVoiceOS /ovos-intent-bench-speech-massive-pt-PT ovos-intent-bench-speech-massive-pt-PT An OVOS Plugin Arena benchmark repository. The arena's prediction runner publishes each plugin's raw output on the speech-massive-pt-PT dataset here, one JSON-lines file per plugin under predictions/<lang>/, and where a sample-set manifest governs scoring it lives under sample_sets/. The public leaderboard at https://openvoiceos.github.io/ovos-plugin-arena/ is computed from these rows, and anyone can recompute an entry from them without… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-intent-bench-speech-massive-pt-PT.tabularn<1K0 likes72 downloads21d agoHugging Face23Gustrd /dolly-15k-libretranslate-pt Summary databricks-dolly-15k ( https://huggingface.co/datasets/databricks/databricks-dolly-15k/ ) is an open source dataset of instruction-following records generated by thousands of Databricks employees in several of the behavioral categories outlined in the InstructGPT paper, including brainstorming, classification, closed QA, generation, information extraction, open QA, and summarization. This is a portuguese translation done with libretranslate (… See the full description on the dataset page: https://huggingface.co/datasets/Gustrd/dolly-15k-libretranslate-pt.textquestion-answering10K<n<100K5 likes71 downloads3y agoHugging Face24OpenVoiceOS /ovos-wake-word-bench-mlsw-negatives-pt-PT ovos-wake-word-bench-mlsw-negatives-pt-PT A sample-set manifest for the OVOS Plugin Arena, not an audio corpus. It holds a seeded, reproducible list of clip identifiers selected from the Multilingual Spoken Words Corpus (MLCommons, CC-BY-4.0), with the seed and the source row count recorded, so every wake-word plugin scored against these negatives is scored on exactly the same clips and false-accept rates are comparable across plugins. The audio is not redistributed here; it… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-mlsw-negatives-pt-PT.tabularn<1K0 likes65 downloads21d agoHugging Face25Magurofg /massive-pt-br MASSIVE pt-BR — localização brasileira Localização para português do Brasil do split pt-PT do MASSIVE (Amazon; CC-BY-4.0) — o benchmark original não tem locale pt-BR. 15.852 frases (train 10.948 / validation 1.929 / test 2.930 aceitas), com anotação de slots preservada ([campo : valor]): a frase limpa é derivada da anotada por construção, então os valores de slot são literais por construção. Método: professor LLM recebe apenas a frase ANOTADA e devolve a versão pt-BR anotada;… See the full description on the dataset page: https://huggingface.co/datasets/Magurofg/massive-pt-br.tabulartext-classification10K<n<100K0 likes63 downloads2mo agoHugging Face26TigreGotico /bifonia-pt-homographs bifonia — Portuguese Heterophonic Homograph Disambiguation Labelled European-Portuguese (pt-PT) sentences for 27 heterophonic homographs — words with identical spelling whose pronunciation (IPA) depends on part of speech or meaning, e.g. para (preposition ˈpɐɾɐ vs verb ˈpaɾɐ), molho (sauce ˈmoʎu vs bundle ˈmɔʎu), corte (royal court ˈkoɾtɨ vs cut ˈkɔɾtɨ). Useful for grapheme-to-phoneme / TTS front-ends and for POS disambiguation. Schema field description… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/bifonia-pt-homographs.texttoken-classification100K<n<1M0 likes60 downloads3mo agoHugging Face27guell00 /Astra-Orbis-5.2K-PT-BR Astra 5.2K PT-BR Extra High Dataset conversacional em português brasileiro, curado para instruction tuning, SFT e treinamento de modelos com foco em respostas detalhadas, raciocínio, explicações passo a passo e comportamento de assistente. Contém 3 milhões de tolkens no total do dataset. O corpus contém conversas no formato chat, exemplos de programação, tarefas analíticas, perguntas educacionais, resolução de problemas, matemática, escrita, tradução, produtividade e respostas… See the full description on the dataset page: https://huggingface.co/datasets/guell00/Astra-Orbis-5.2K-PT-BR.text1K<n<10K1 likes60 downloads3mo agoHugging Face28amalia-llm /PT-Culture_Data Portuguese Cultural SFT Dataset A supervised fine-tuning dataset for European Portuguese (PT-PT) cultural knowledge, built to teach models the traditions, figures, places, and expressions of Portuguese culture. The dataset contains 216,832 examples across two versions in conversational SFT format, organised into ten cultural domains: Domain v1 (95,815) v2 (121,017) Total Personalities 36,370 62,788 99,158 Audiovisual 15,786 26,857 42,643 Heritage 8,020… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/PT-Culture_Data.textquestion-answering100K<n<1M1 likes54 downloads2mo agoHugging Face29Magurofg /ifeval-pt IFEval-pt Benchmark de seguimento de instrução em português do Brasil, com 308 itens cuja verificação é feita por código — sem juiz LLM, sem GPU, sem opinião. Reproduzir custa alguns minutos de CPU. O que ele tem de diferente Já existem versões do IFEval em português: amalia-llm/IFEval-mt-pt (541 itens, pt-PT), Polygl0t/IFEval-PT (300), a fatia pt do facebook/Multi-IF (524) e akawaai/Portuguese-IFEval (618). Todas são tradução do conjunto em inglês — medem… See the full description on the dataset page: https://huggingface.co/datasets/Magurofg/ifeval-pt.texttext-generationn<1K0 likes52 downloads2mo agoHugging Face30amalia-llm /smoltalk2_everyday_conv_pt SMOL Everyday Conversation PT This dataset consists of a Portuguese version of the smoltalk_smollm3_everyday_conversations_no_think split of HuggingFaceTB/smoltalk2. The first two user turns were translated as well as the first assistant turn, the continuation of the conversation was generated using Gemma 3-27B. Note: This dataset is machine translated and may contain translation errors or artifacts. This dataset is provided as part of the AMALIA project and is… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/smoltalk2_everyday_conv_pt.texttext-generation1K<n<10K1 likes51 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.