CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hanlincs /in1k_clip_qwen25vl_3b_224res_64tokens_new_pttabular1M<n<10M0 likes515 downloads1y agoHugging Face02hanlincs /in1k_clip_qwen25vl_3b_448res_256tokens_new_merged_pttabular1M<n<10M0 likes470 downloads1y agoHugging Face03OpenVoiceOS /MT-intents-dataset-pt-PTtext10K<n<100K0 likes311 downloads1y agoHugging Face04lucas-leme /Sentiments-FinBERT-PT-BR Dataset A manually annotated dataset was created to enable supervised training for the FinBERT-PT-BR model, which focuses on sentiment analysis of Brazilian Portuguese financial texts. More than 1.4 million financial news texts in Portuguese were collected and used for the initial language modeling phase. From this corpus, a sample of 1,000 texts was manually annotated with sentiment labels. Annotation Process Three annotators participated in the process. All texts were… See the full description on the dataset page: https://huggingface.co/datasets/lucas-leme/Sentiments-FinBERT-PT-BR.textn<1K1 likes84 downloads1y agoHugging Face05Edu-p /harmful-prompts-pt Harmful Prompts PT-BR harmful-prompts-pt is a Brazilian Portuguese adaptation of the WildJailbreak dataset, constructed to support research on the robustness of language models against harmful and adversarial prompts in Portuguese. This dataset was used to train and evaluate SecBERT, a Portuguese harmful prompt classifier presented at the International Joint Conference on Neural Networks (IJCNN). The full paper and source code are available at [Paper] [Code]. Caution: This dataset… See the full description on the dataset page: https://huggingface.co/datasets/Edu-p/harmful-prompts-pt.texttext-classification10K<n<100K4 likes78 downloads5mo agoHugging Face06vsvasconcelos /SQuAD-pt_BR-V1.1_ Dataset Card para o SQuAD 1.1 em Português Brasil O conjunto de dados "Stanford Question Answering Dataset" (SQuAD), para tarefa de perguntas e respostas extrativas, foi desenvolvido em 2016. Ele utiliza perguntas geradas a partir de 536 artigos da Wikipedia* com mais de 100.000 linhas de dados. É construído na forma de uma pergunta e um contexto dos artigos da Wikipedia contendo a resposta à pergunta. [1]Originalmente este dataset foi construído no idioma inglês, contudo, o grupo… See the full description on the dataset page: https://huggingface.co/datasets/vsvasconcelos/SQuAD-pt_BR-V1.1_.text10K<n<100K2 likes76 downloads2y agoHugging Face07pt-sk /toxic_classificationcombination of SetFit/toxic_conversations_50k, Arsive/toxicity_classification_jigsaw texttext-classification100K<n<1M0 likes74 downloads2y agoHugging Face08ju-resplande /qa-pt Dataset Card for QA-Portuguese Dataset Summary Portuguese preprocessed split from MQA dataset. Supported Tasks and Leaderboards [More Information Needed] Languages The dataset is Portuguese. Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation Curation Rationale [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/ju-resplande/qa-pt.textquestion-answering1M<n<10M17 likes73 downloads4y agoHugging Face09vsvasconcelos /SQuAD-pt_BR-V1.1 Dataset Card for Dataset Name O conjunto de dados "Stanford Question Answering Dataset" (SQuAD), para tarefa de perguntas e respostas extrativas, foi desenvolvido em 2016. Ele utiliza perguntas geradas a partir de 536 artigos da Wikipedia com mais de 100.000 linhas de dados. É construído na forma de uma pergunta e um contexto dos artigos da Wikipedia contendo a resposta à pergunta. Originalmente este dataset foi construído no idioma inglês, contudo, o grupo Deep Learning Brasil… See the full description on the dataset page: https://huggingface.co/datasets/vsvasconcelos/SQuAD-pt_BR-V1.1.text10K<n<100K1 likes58 downloads3y agoHugging Face10YYJMAY /pt-recognition PT Interaction Dataset Dataset Description The PT (Peptide-TCR) interaction dataset is designed for training and evaluating T-Cell Receptor (TCR) binding prediction models with full TCR sequence information. This dataset contains paired peptide sequences and complete TCR alpha/beta chain sequences (including all 6 CDR regions: A1-A3, B1-B3), along with binary binding labels. Key Features Full TCR Information: Contains all 6 CDR regions (A1, A2, A3, B1, B2, B3)… See the full description on the dataset page: https://huggingface.co/datasets/YYJMAY/pt-recognition.texttext-classification10K<n<100K0 likes55 downloads10mo agoHugging Face11c2d-usp /banking77-pt-br Resumo do conjunto de dados Tradução revisada para o português brasileiro do conjunto de dados BANKING77. O conjunto é composto por consultas online realizadas a sistemas conversacionais de bancos, rotuladas de acordo com suas intenções correspondentes. De acordo com a documentação original, as 13.083 consultas que compõem o conjunto são referentes a atendimentos ao cliente, e foram rotuladas em 77 intenções diferentes. O foco principal é suportar análises em domínio específico e… See the full description on the dataset page: https://huggingface.co/datasets/c2d-usp/banking77-pt-br.text10K<n<100K1 likes54 downloads1y agoHugging Face12tech4humans /JBB-Behaviors-pt JBB-Behaviors-pt Dataset Description JBB-Behaviors-pt is a dataset of behaviors in Portuguese, including both jailbreak prompts and safe behaviors. The dataset is intended for behavioral testing of language models to evaluate their robustness against jailbreak attempts in Portuguese. What is a jailbreak? Jailbreak prompts are inputs designed to bypass a language model's safety guardrails, potentially causing it to generate harmful, unethical, or otherwise… See the full description on the dataset page: https://huggingface.co/datasets/tech4humans/JBB-Behaviors-pt.texttext-classificationn<1K1 likes29 downloads1y agoHugging Face13amalia-llm /humaneval_mt_pt HumanEval-PT Portuguese version of HumanEval, a code generation benchmark with programming problems and test cases. Translated using NLLB-200. Original Dataset: https://huggingface.co/datasets/openai/openai_humaneval Note: This dataset is machine translated and may contain translation errors or artifacts. This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite for evaluating large language models on… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/humaneval_mt_pt.texttext-generationn<1K0 likes28 downloads3mo agoHugging Face14luist18 /pt-parliament-interventionstextn<1K0 likes27 downloads3y agoHugging Face15IuryCavalcante /TechnicalDebt_GitHubIssues_PT Technical Debt in GitHub Issues (Portuguese) Visão Geral Este dataset reúne issues públicas extraídas do GitHub que mencionam o termo "dívida técnica" ou suas variações em português. A base foi construída com o objetivo de apoiar pesquisas em Engenharia de Software, Processamento de Linguagem Natural (PLN) e Inteligência Artificial, com foco na compreensão e classificação de como o conceito de dívida técnica é comunicado por desenvolvedores. A estrutura e categorização… See the full description on the dataset page: https://huggingface.co/datasets/IuryCavalcante/TechnicalDebt_GitHubIssues_PT.texttext-classificationn<1K0 likes26 downloads1y agoHugging Face16proxectonos /XStoryCloze_pttext1K<n<10K1 likes24 downloads2y agoHugging Face17ruibrogandrade /ARC-Challenge_PT-PT Dataset Card for ARC-Challenge_PT-PT Dataset Summary This repository contains a European Portuguese (pt-PT) translation of ARC-Challenge (AI2 Reasoning Challenge – Challenge subset), a benchmark of non-trivial, grade-school science questions that require background knowledge and reasoning. Each example presents a question and four answer choices.Load with: import datasets data = datasets.load_dataset("ruibrogandrade/ARC-Challenge_PT-PT") Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/ruibrogandrade/ARC-Challenge_PT-PT.text1K<n<10K1 likes18 downloads1y agoHugging Face18amalia-llm /sst_pt-pt Simple Safety Tests-PT Portuguese machine translation of Simple Safety Tests, a benchmark for evaluating model safety and harmful content detection. Translated using a Finetuned GemmaX2-9B for pt-PT. Original Dataset: https://huggingface.co/datasets/Bertievidgen/SimpleSafetyTests Note: This dataset is machine translated and may contain translation errors or artifacts. This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/sst_pt-pt.texttext-generationn<1K0 likes18 downloads3mo agoHugging Face19insfsilva /santfordimdb-pttext10K<n<100K0 likes17 downloads2y agoHugging Face20ruibrogandrade /HellaSwag_PT-PT Dataset Card for HellaSwag_PT-PT Dataset Summary This repository provides a European Portuguese (pt-PT) translation of HellaSwag, an adversarial benchmark for grounded commonsense inference. Given a short scenario, the model must choose the most plausible continuation from four options.Load with: import datasets data = datasets.load_dataset("ruibrogandrade/HellaSwag_PT-PT") Supported Tasks and Leaderboards Multiple-Choice Question Answering (commonsense… See the full description on the dataset page: https://huggingface.co/datasets/ruibrogandrade/HellaSwag_PT-PT.tabular1K<n<10K1 likes16 downloads1y agoHugging Face21vitorandrade /Squad_PT Dataset Card para o SQuAD 1.1 em Português Brasil O conjunto de dados "Stanford Question Answering Dataset" (SQuAD), para tarefa de perguntas e respostas extrativas, foi desenvolvido em 2016. Ele utiliza perguntas geradas a partir de 536 artigos da Wikipedia* com mais de 100.000 linhas de dados. É construído na forma de uma pergunta e um contexto dos artigos da Wikipedia contendo a resposta à pergunta. [1]Originalmente este dataset foi construído no idioma inglês, contudo, o grupo… See the full description on the dataset page: https://huggingface.co/datasets/vitorandrade/Squad_PT.text10K<n<100K0 likes15 downloads2y agoHugging Face22Junaid687 /gemma-3-1b-pt-blind-spots Gemma-3-1b-pt Blind Spots Dataset Dataset Description This dataset documents blind spots (systematic errors) found in google/gemma-3-1b-pt, a 1-billion-parameter pretrained base model (not instruction-tuned) released by Google in March 2025 as part of the Gemma 3 family. Each row contains: Column Description id Unique probe index category Type of reasoning tested prompt The input fed to the model (text-completion style) expected_output The… See the full description on the dataset page: https://huggingface.co/datasets/Junaid687/gemma-3-1b-pt-blind-spots.texttext-generationn<1K0 likes15 downloads7mo agoHugging Face23VanessaSchenkel /pt-all-words Dataset Card for Dicionário Português It is a list of portuguese words with its inflections How to use it: from datasets import load_dataset remote_dataset = load_dataset("VanessaSchenkel/pt-all-words") remote_dataset textother10K<n<100K7 likes14 downloads4y agoHugging Face24pt-sk /toxic_classification_balancedcombination of SetFit/toxic_conversations_50k, Arsive/toxicity_classification_jigsaw taken sample from toxic classification - to balance the dataset texttext-classification10K<n<100K0 likes14 downloads2y agoHugging Face25proxectonos /PAWS_pttabular1K<n<10K0 likes14 downloads2y agoHugging Face26godoyj /cstnews-pttextn<1K0 likes12 downloads3y agoHugging Face27johnidouglas /twitter-sentiment-pt-BR-md-2-ltext10K<n<100K1 likes12 downloads2y agoHugging Face28ruibrogandrade /Social_I_QA_PT-PT Dataset Card for Social_I_QA_PT-PT Dataset Summary This repository contains a European Portuguese (pt-PT) translation of Social IQa, a benchmark for social commonsense reasoning. Each example includes a short context about a social situation, a question, and three answer candidates. The task is to choose the most plausible answer about intents, reactions, or social outcomes.Load with: import datasets data = datasets.load_dataset("ruibrogandrade/Social_I_QA_PT-PT")… See the full description on the dataset page: https://huggingface.co/datasets/ruibrogandrade/Social_I_QA_PT-PT.text1K<n<10K1 likes12 downloads1y agoHugging Face29insfsilva /imdb-pt-brtabular10K<n<100K0 likes11 downloads2y agoHugging Face30Toka-Tarek /gemma-3-1b-pt-blind-spots Blind Spots of google/gemma-3-1b-pt Model Tested Model: google/gemma-3-1b-ptParameters: 1BType: Pre-trained base language model (not instruction-tuned)Tested by: Toka-Tarek | Biotechnology graduate & Pharmacogenetics Lab Specialist How I Loaded the Model Tested on Google Colab (free T4 GPU, 16GB VRAM). Note: torch.float16 caused numerical instability (NaN/inf errors) on the T4 GPU, so torch.float32 was used instead for stable generation. from huggingface_hub… See the full description on the dataset page: https://huggingface.co/datasets/Toka-Tarek/gemma-3-1b-pt-blind-spots.texttext-generationn<1K0 likes11 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.