CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01TigreGotico /portuguese-dialects-ipa-synthetic portuguese-dialects-ipa-synthetic 920 dialect-register sentences — 20 per lect across 46 lects: 41 Portuguese varieties (European regional, insular, Brazilian regional, African/Asian/border national norms, medieval stages) plus the other languages of Portugal and their kin: Mirandese (3 lects), Asturleonese of Portugal (Rionorese, Guadramilese) with Barranquenho, and Galician-Portuguese. Each row carries two IPA columns with distinct provenance. Schema sentence —… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/portuguese-dialects-ipa-synthetic.texttext-to-speechn<1K0 likes111 downloads2mo agoHugging Face02Paul /hatecheck-portuguese Dataset Card for Multilingual HateCheck Dataset Description Multilingual HateCheck (MHC) is a suite of functional tests for hate speech detection models in 10 different languages: Arabic, Dutch, French, German, Hindi, Italian, Mandarin, Polish, Portuguese and Spanish. For each language, there are 25+ functional tests that correspond to distinct types of hate and challenging non-hate. This allows for targeted diagnostic insights into model performance. For more details… See the full description on the dataset page: https://huggingface.co/datasets/Paul/hatecheck-portuguese.tabulartext-classification1K<n<10K13 likes109 downloads4y agoHugging Face03stjiris /portuguese-legal-sentences-v0 Work developed as part of Project IRIS. Thesis: A Semantic Search System for Supremo Tribunal de Justiça Portuguese Legal Sentences Collection of Legal Sentences from the Portuguese Supreme Court of Justice The goal of this dataset was to be used for MLM and TSDAE Contributions @rufimelo99 If you use this work, please cite: @InProceedings{MeloSemantic, author="Melo, Rui and Santos, Pedro A. and Dias, Jo{\~a}o", editor="Moniz, Nuno and Vale, Zita and… See the full description on the dataset page: https://huggingface.co/datasets/stjiris/portuguese-legal-sentences-v0.text1M<n<10M14 likes97 downloads2y agoHugging Face04rufimelo /PortugueseLegalSentences-v0 Portuguese Legal Sentences Collection of Legal Sentences from the Portuguese Supreme Court of Justice The goal of this dataset was to be used for MLM and TSDAE Contributions @rufimelo99 text1M<n<10M3 likes46 downloads4y agoHugging Face05rufimelo /PortugueseLegalSentences-v3 Portuguese Legal Sentences Collection of Legal Sentences from the Portuguese Supreme Court of Justice The goal of this dataset was to be used for MLM and TSDAE Extended version of rufimelo/PortugueseLegalSentences-v1 400000/50000/50000 Contributions @rufimelo99 text100K<n<1M8 likes38 downloads4y agoHugging Face06adalbertojunior /openHermes_portuguesetext1M<n<10M3 likes37 downloads2y agoHugging Face07Speech-data /Portuguese-Speech-Dataset 🎧 Portuguese Speech Dataset The Portuguese Speech Dataset is a large-scale speech audio dataset designed to provide structured and high-quality audio data for modern AI and machine learning systems. It contains 195 hours of recorded speech data distributed across 894 files, available in MP3 and WAV formats, with a total size of 437 MB. This carefully curated audio dataset delivers diverse and representative voice data, with a balanced speaker distribution of 52% female and 48% male… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Portuguese-Speech-Dataset.audioautomatic-speech-recognitionn<1K0 likes30 downloads6mo agoHugging Face08rufimelo /PortugueseLegalSentences-v1 Portuguese Legal Sentences Collection of Legal Sentences from the Portuguese Supreme Court of Justice The goal of this dataset was to be used for MLM and TSDAE Contributions @rufimelo99 text100K<n<1M2 likes27 downloads4y agoHugging Face09TigreGotico /portuguese_phonetic_lexicon 📚 Portuguese Phonetic Lexicon Dataset This dataset contains phonetic and morphological information for Portuguese words, collected from the Portal da Língua Portuguesa. It was generated by scraping the site across multiple Portuguese-speaking regions and dialects. 🌍 Regional Coverage The dataset includes words as spoken in ten regional variants: 🇵🇹 Lisbon (Standard and Non-Standard) 🇦🇴 Luanda 🇧🇷 Rio de Janeiro (Standard and Non-Standard) 🇧🇷 São Paulo (Standard… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/portuguese_phonetic_lexicon.texttext-classification100K<n<1M0 likes27 downloads7mo agoHugging Face10MLap /English-French-Portuguese-Lexicon Demo Notebook for this dataset Dataset created with Gemini-Flash-2.5 API with the prompt given below: Generate a list of 500 simple and commonly used English words, each translated into French and Portuguese. Format the output as CSV with the columns: English, French, Portuguese. Only include single words (no phrases or verbs starting with ‘to’, like ‘to eat’ or ‘to go’). Avoid grammatical verbs and ensure no repetitions. Extensive manual cleaning was done with the help of Google… See the full description on the dataset page: https://huggingface.co/datasets/MLap/English-French-Portuguese-Lexicon.textn<1K0 likes23 downloads1y agoHugging Face11manueltonneau /portuguese-hate-speech-supersetgated Portuguese Hate Speech Superset This dataset is a superset (N=43,222) of posts annotated as hateful or not. It results from the preprocessing and merge of all available Portuguese hate speech datasets in April 2024. These datasets were identified through a systematic survey of hate speech datasets conducted in early 2024. We only kept datasets that: are documented are publicly available focus on hate speech, defined broadly as "any kind of communication in speech, writing or… See the full description on the dataset page: https://huggingface.co/datasets/manueltonneau/portuguese-hate-speech-superset.tabulartext-classification10K<n<100K2 likes22 downloads2y agoHugging Face12musts /brazilian_portuguesetext1K<n<10K1 likes21 downloads2y agoHugging Face13rufimelo /PortugueseLegalSentences-v2 Portuguese Legal Sentences Collection of Legal Sentences from the Portuguese Supreme Court of Justice The goal of this dataset was to be used for MLM and TSDAE Extended version of rufimelo/PortugueseLegalSentences-v1 200000/200000/100000 Contributions @rufimelo99 text100K<n<1M2 likes19 downloads4y agoHugging Face14Samambas /Human_vs_AI_Portuguesetext100K<n<1M0 likes19 downloads8mo agoHugging Face15luist18 /mbic-portuguesetext1K<n<10K0 likes16 downloads3y agoHugging Face16Solshine /Portuguese-English-Vocab-PartiallyTransformedNotes on use: Portuguese and English Translations of readme are available here. Partially cleaned and reorganized. Minimal secondhand verification after generation through Google Bard on November 28th 2023. Mistakes are minimal but present, such as tagging of words in supplemental information sometimes using the whole word (ie Noun) and sometimes only a letter or abreviation (ie N) for the same part of speech. Reccomended for finetuning of smaller models only, such as 12, 7, or 3 B models to… See the full description on the dataset page: https://huggingface.co/datasets/Solshine/Portuguese-English-Vocab-PartiallyTransformed.text1K<n<10K3 likes16 downloads3y agoHugging Face17miguelribeirokk /crime_tweets_in_portuguese DataCrimeBR: Building a Dataset of Crimes Reported in Tweets in Brazil This dataset contains 61.715 tweets related to possible crime reports, labeled with categories such as "Assalto", "Roubo", "Furto", "Assédio", "Segurança Pública", "Homicídio, and "Outros", along with sentiment analysis, toxicity analysis, and location identification. A particular feature in the Portuguese language is that many words potentially related to crimes are used in non-criminal contexts, such as "O… See the full description on the dataset page: https://huggingface.co/datasets/miguelribeirokk/crime_tweets_in_portuguese.tabular10K<n<100K1 likes16 downloads10mo agoHugging Face18adalbertojunior /dolphin-2.9-portuguesetext100K<n<1M2 likes13 downloads2y agoHugging Face19matos1012 /brazilian-portuguese-anaphylaxis Brazilian Portuguese Clinical Notes for Anaphylaxis Detection A dataset of 969 Brazilian Portuguese clinical narratives annotated for the presence or absence of anaphylaxis, built for research in clinical Natural Language Processing (NLP). Overview Anaphylaxis is an acute, potentially life-threatening allergic reaction that requires rapid recognition in clinical settings. Automatic detection of anaphylaxis in clinical narratives can support large-scale analysis of… See the full description on the dataset page: https://huggingface.co/datasets/matos1012/brazilian-portuguese-anaphylaxis.tabularn<1K0 likes13 downloads6mo agoHugging Face20luist18 /portuguese-parliament-interventionstabular1K<n<10K2 likes12 downloads3y agoHugging Face21johnpaulbin /portuguese-hate-speechtabular1K<n<10K2 likes9 downloads1y agoHugging Face22cemig-ceia /fineweb-edu-gemini-annotations-portuguese-regressiontabular1K<n<10K0 likes9 downloads1y agoHugging Face23DiegoAlysson /Translated_Expanded_CC3M-Brazilian_Portuguese-Hindi-Xhosa CC3M Multilingual & Augmented Variants This repository provides four multilingual, augmented, and similarity-enhanced variants of the Conceptual Captions 3M (CC3M) dataset.The goal is to support research in vision–language modeling, multimodal alignment, data augmentation, and low-resource language evaluation. All versions include translations generated with Google Translate and MarianMT, and caption augmentations produced with BLIP2, generating five additional captions per… See the full description on the dataset page: https://huggingface.co/datasets/DiegoAlysson/Translated_Expanded_CC3M-Brazilian_Portuguese-Hindi-Xhosa.text1M<n<10M0 likes9 downloads10mo agoHugging Face24musts /portuguesetext1K<n<10K0 likes8 downloads2y agoHugging Face25FluxiIA /math_portuguesetext100K<n<1M0 likes7 downloads2y agoHugging Face26TigreGotico /portuguese-sentences-synthetic-g2p Dataset Card for 'TigreGotico/portuguese_g2p' Dataset Description Dataset Summary TigreGotico/portuguese_g2p is a Grapheme-to-Phoneme (G2P) dataset for Portuguese, offering phonetic transcriptions for sentences across ten different regional variants. It is derived from the portuguese_phonetic_lexicon and is designed to aid in the development of robust Speech Recognition (ASR) and Text-to-Speech (TTS) models that account for dialectal variation in Portuguese.… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/portuguese-sentences-synthetic-g2p.texttext-classification10K<n<100K0 likes7 downloads1y agoHugging Face27cemig-ceia /cosmopedia-portuguesetext1M<n<10M1 likes6 downloads2y agoHugging Face28DiegoAlysson /Translated_CC12M-Brazilian_Portuguesetexttext-to-image1M<n<10M0 likes6 downloads10mo agoHugging Face29luiseduardobrito /similarity-sentences-portuguese similarity-sentences-portuguese (SSP) Dataset Summary This dataset comprises a collection of sentences generated using Chat GPT-3, covering various general topics, originally in spanish by jaimevera1107. The sentences were translated to portuguese using seamless-m4t-medium. Languages Portuguese Dataset Structure Data Fields Sentence 1: The first sentence to be compared. Sentence 2: The second sentence to be compared. Score: A number… See the full description on the dataset page: https://huggingface.co/datasets/luiseduardobrito/similarity-sentences-portuguese.texttext-classification10K<n<100K4 likes5 downloads3y agoHugging Face30mmt93 /zeroshot_portuguesegatedtabular1K<n<10K2 likes3 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.