CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01eduagarcia /portuguese_benchmark Portuguese Benchmark This a collection of datasets in Portuguese initially meant to train and evaluate supervised language models such as BERT, RoBERTa, etc... It contains 10 datasets and 18 Tasks for Classification (CLS), NLI, Semantic Similarity Scoring (STS) and Named-Entity Recognition (NER). NER Classification NLI STS LeNER-Br HateBR_offensive_binary assin2-rte assin2-sts UlyssesNER-Br-PL-coarse HateBR_offensive_level UlyssesNER-Br-C-coarse… See the full description on the dataset page: https://huggingface.co/datasets/eduagarcia/portuguese_benchmark.tabular10K<n<100K7 likes1.2k downloads2y agoHugging Face02AIML-TUDA /SLR-Bench-Portuguese 🧠 SLR-Bench-Portuguese: Scalable Logical Reasoning Benchmark (Portuguese Edition) SLR-Bench Multilingual Versions: SLR-Bench-Portuguese is the Portuguese-language pendant of the original SLR-Bench dataset. It follows the same symbolic structure, evaluation framework, and curriculum as the English version but provides all natural-language task prompts translated into Portuguese. This enables systematic evaluation and training of Large Language Models… See the full description on the dataset page: https://huggingface.co/datasets/AIML-TUDA/SLR-Bench-Portuguese.tabular10K<n<100K0 likes294 downloads4mo agoHugging Face03Polygl0t /portuguese-eval-logs-olmo2-smollm3 Evaluation Logs on Portuguese Benchmarks for OLMo-2 and SmolLM3 These logs contain benchmark results across a suite of Portuguese-language tasks. The data consists of recordings of the performance of various 3 different models at different checkpoints throughout their pretraining runs: SmolLM3 OLMo-2-0425-1B OLMo-2-1124-7B Splits Each split (smollm3_3b, olmo2_1b, olmo2_7b) contains rows for model checkpoints and columns for benchmark scores (e.g., ASSIN2 RTE, ENEM… See the full description on the dataset page: https://huggingface.co/datasets/Polygl0t/portuguese-eval-logs-olmo2-smollm3.imagen<1K0 likes149 downloads7mo agoHugging Face04Paul /hatecheck-portuguese Dataset Card for Multilingual HateCheck Dataset Description Multilingual HateCheck (MHC) is a suite of functional tests for hate speech detection models in 10 different languages: Arabic, Dutch, French, German, Hindi, Italian, Mandarin, Polish, Portuguese and Spanish. For each language, there are 25+ functional tests that correspond to distinct types of hate and challenging non-hate. This allows for targeted diagnostic insights into model performance. For more details… See the full description on the dataset page: https://huggingface.co/datasets/Paul/hatecheck-portuguese.tabulartext-classification1K<n<10K13 likes110 downloads4y agoHugging Face05aimeri /ticuna-spanish-portuguese Ticuna (tca) – Spanish – Portuguese Corpus First text corpus for Ticuna (ISO 639-3 tca), a tonal language isolate of the Brazil/Colombia/Peru tri-border. Configs | Config | Rows | | parallel | train 43,248 / validation 596 / test 3,238 | | monolingual | train 46,545 / validation 298 / test 1,613 | | lexicon | train 10,419 / validation 568 / test 539 | | instructions | train 52,836 | | backtranslation | train 33,944 | The short version of what matters… See the full description on the dataset page: https://huggingface.co/datasets/aimeri/ticuna-spanish-portuguese.tabular100K<n<1M0 likes106 downloads28d agoHugging Face06LIACC /Emakhuwa-Portuguese-News-MT News Parallel Dataset for Emakhuwa of Mozambique This repository contains releases of parallel data for machine translation in Mozambican languages. Currently, it supports one language pair, Portuguese-Emakhuwa, Emakhuwa being the widely spoken language in Mozambique. Dataset Details Dataset Description Funded by: This dataset was created with support from Lacuna Fund, the world’s first collaborative effort to provide data scientists, researchers, and… See the full description on the dataset page: https://huggingface.co/datasets/LIACC/Emakhuwa-Portuguese-News-MT.tabulartranslation10K<n<100K0 likes97 downloads2y agoHugging Face07fpaulino /portuguese-tweetstabular100K<n<1M4 likes51 downloads4y agoHugging Face08iara-project /test_split_with_embeddings_bert_base_portuguese Dataset Card for "test_split_with_embeddings_bert_base_portuguese" More Information needed tabular100K<n<1M0 likes36 downloads3y agoHugging Face09TaigoPedrosa /PortugueseMMLU Dataset Components The dataset is partitioned into three discrete tables stored in CSV or Parquet format: Questions Recipes Evaluation Results Each component is described in detail below. Questions area domain question_number An integer index uniquely identifying each question inside the knowledge domain. translation_method English, Google Translate, GPT-3.5-Turbo, GPT-4o, Human question option_a, option_b, option_c, option_d Recipes area… See the full description on the dataset page: https://huggingface.co/datasets/TaigoPedrosa/PortugueseMMLU.tabularquestion-answering100K<n<1M0 likes36 downloads1y agoHugging Face10Polygl0t /portuguese-instruct-quality-qwen-annotations Annotations for the Portuguese-instruct-quality classifier 📚 Dataset Summary This dataset contains the annotations used for training a quality filter for instruction type data (Polygl0t/portuguese-qwen3-4b-instruct-quality-classifier and Polygl0t/portuguese-qwen3-4b-instruct-quality-judge). These annotations were generated by Qwen/Qwen2.5-32B-Instruct. Supported Tasks and Leaderboards This dataset can be used for the task of text classification, or for… See the full description on the dataset page: https://huggingface.co/datasets/Polygl0t/portuguese-instruct-quality-qwen-annotations.tabulartext-classification100K<n<1M0 likes33 downloads7mo agoHugging Face11ahmad21omar /SLR-Bench-Portuguese 🧠 SLR-Bench-Portuguese: Scalable Logical Reasoning Benchmark (Portuguese Edition) SLR-Bench Versions: SLR-Bench-Portuguese is the Portuguese-language pendant of the original SLR-Bench dataset. It follows the same symbolic structure, evaluation framework, and curriculum as the English version but provides all natural-language task prompts translated into Portuguese. This enables systematic evaluation and training of Large Language Models (LLMs) in logical… See the full description on the dataset page: https://huggingface.co/datasets/ahmad21omar/SLR-Bench-Portuguese.tabular10K<n<100K0 likes31 downloads11mo agoHugging Face12marcosremar2 /orpheus-tts-portuguese-datasettabular100K<n<1M0 likes26 downloads10mo agoHugging Face13manueltonneau /portuguese-hate-speech-supersetgated Portuguese Hate Speech Superset This dataset is a superset (N=43,222) of posts annotated as hateful or not. It results from the preprocessing and merge of all available Portuguese hate speech datasets in April 2024. These datasets were identified through a systematic survey of hate speech datasets conducted in early 2024. We only kept datasets that: are documented are publicly available focus on hate speech, defined broadly as "any kind of communication in speech, writing or… See the full description on the dataset page: https://huggingface.co/datasets/manueltonneau/portuguese-hate-speech-superset.tabulartext-classification10K<n<100K2 likes22 downloads2y agoHugging Face14portuguese-benchmark-datasets /xpaws_pt Dataset Card for "xpaws_pt" This is a portuguese translation of the x-paws dataset. The translation was performed using the Google Translate API. This dataset follows the same structure as the original. tabular1K<n<10K1 likes19 downloads3y agoHugging Face15safety-aya /fineweb-portuguese-100k FineWeb2 Portuguese 100k - Safety Classified A 100,000-sample subset of FineWeb2 Portuguese web text, classified for content safety using Cohere Command A. Dataset Description Each record contains the original FineWeb2 text and metadata, plus a classification field with: Field Description safety_rating "safe" or "unsafe" category List of applicable harm categories (null if safe) reason Brief explanation of the classification Safety Taxonomy… See the full description on the dataset page: https://huggingface.co/datasets/safety-aya/fineweb-portuguese-100k.tabulartext-classification100K<n<1M0 likes19 downloads7mo agoHugging Face16miguelribeirokk /crime_tweets_in_portuguese DataCrimeBR: Building a Dataset of Crimes Reported in Tweets in Brazil This dataset contains 61.715 tweets related to possible crime reports, labeled with categories such as "Assalto", "Roubo", "Furto", "Assédio", "Segurança Pública", "Homicídio, and "Outros", along with sentiment analysis, toxicity analysis, and location identification. A particular feature in the Portuguese language is that many words potentially related to crimes are used in non-criminal contexts, such as "O… See the full description on the dataset page: https://huggingface.co/datasets/miguelribeirokk/crime_tweets_in_portuguese.tabular10K<n<100K1 likes16 downloads10mo agoHugging Face17safety-aya /Nemotron-Safety-Guard-Dataset-v3-portuguese Nemotron Portuguese Safety (Translated) Portuguese safety prompts/responses (translated from Spanish), with labels and categories. Dataset Description nemotron_pt Each record includes Portuguese prompt/response text plus safety labels/categories. Field Description id Example id prompt Portuguese prompt text response Portuguese response text (may be null) prompt_label "safe" or "unsafe" response_label "safe" or "unsafe" (may be empty if… See the full description on the dataset page: https://huggingface.co/datasets/safety-aya/Nemotron-Safety-Guard-Dataset-v3-portuguese.tabulartext-classification10K<n<100K0 likes12 downloads6mo agoHugging Face18matos1012 /brazilian-portuguese-anaphylaxis Brazilian Portuguese Clinical Notes for Anaphylaxis Detection A dataset of 969 Brazilian Portuguese clinical narratives annotated for the presence or absence of anaphylaxis, built for research in clinical Natural Language Processing (NLP). Overview Anaphylaxis is an acute, potentially life-threatening allergic reaction that requires rapid recognition in clinical settings. Automatic detection of anaphylaxis in clinical narratives can support large-scale analysis of… See the full description on the dataset page: https://huggingface.co/datasets/matos1012/brazilian-portuguese-anaphylaxis.tabularn<1K0 likes12 downloads6mo agoHugging Face19safety-aya /fineweb2-portuguese-safetytabular100K<n<1M0 likes12 downloads6mo agoHugging Face20luist18 /portuguese-parliament-interventionstabular1K<n<10K2 likes11 downloads3y agoHugging Face21safety-aya /wildguardtest-portuguesetabular1K<n<10K0 likes11 downloads6mo agoHugging Face22k-mktr /almeida_portuguese_pt Almeida Atualizada (Portuguese) Description A Almeida Atualizada is a revised edition of the classic Portuguese Bible translation by João Ferreira de Almeida (1628-1691), a Portuguese Protestant pastor. Almeida's original translation, the first complete Portuguese Bible, was published in 1748-1753. The Almeida Atualizada (updated Almeida) modernizes the language while preserving the fidelity to the original Hebrew and Greek texts. It is the most widely used… See the full description on the dataset page: https://huggingface.co/datasets/k-mktr/almeida_portuguese_pt.tabular10K<n<100K0 likes11 downloads2mo agoHugging Face23cemig-ceia /fineweb-edu-gemini-annotations-portuguese-regressiontabular1K<n<10K0 likes10 downloads1y agoHugging Face24safety-aya /fineweb-portuguese-chunktabular100K<n<1M0 likes10 downloads6mo agoHugging Face25johnpaulbin /portuguese-hate-speechtabular1K<n<10K2 likes9 downloads1y agoHugging Face26safety-aya /Nemotron-Safety-Guard-Dataset-v3-portuguese-safetytabular10K<n<100K0 likes7 downloads6mo agoHugging Face27mmt93 /zeroshot_portuguesegatedtabular1K<n<10K2 likes3 downloads3y agoHugging Face28EdwardSJ151 /portuguese_hate_speech_lighteval_fewshottabular1K<n<10K0 likes3 downloads1y agoHugging Face29marcosremar2 /mls-portuguese-snactabular10K<n<100K0 likes3 downloads8mo agoHugging Face30louisjeon /portuguese-hate-speechtabular10K<n<100K0 likes1 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.