CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01pedrohavay /portuguese-male-voice-A-datasetaudio4 likes5.3k downloads3y agoHugging Face02cnmoro /reasoning-v1-20m-portugueseglaiveai/reasoning-v1-20m translated to portuguese. texttext-generation10M<n<100M14 likes1.6k downloads1y agoHugging Face03eduagarcia /portuguese_benchmark Portuguese Benchmark This a collection of datasets in Portuguese initially meant to train and evaluate supervised language models such as BERT, RoBERTa, etc... It contains 10 datasets and 18 Tasks for Classification (CLS), NLI, Semantic Similarity Scoring (STS) and Named-Entity Recognition (NER). NER Classification NLI STS LeNER-Br HateBR_offensive_binary assin2-rte assin2-sts UlyssesNER-Br-PL-coarse HateBR_offensive_level UlyssesNER-Br-C-coarse… See the full description on the dataset page: https://huggingface.co/datasets/eduagarcia/portuguese_benchmark.tabular10K<n<100K7 likes1.2k downloads2y agoHugging Face04portuguese-benchmark-datasets /BLUEX BLUEX There is a repository with the minimal code for using this dataset available here. If you use this dataset for research, please cite the paper: @misc{almeida2023bluex, title={BLUEX: A benchmark based on Brazilian Leading Universities Entrance eXams}, author={Thales Sales Almeida and Thiago Laitz and Giovana K. Bonás and Rodrigo Nogueira}, year={2023}, eprint={2307.05410}, archivePrefix={arXiv}, primaryClass={cs.CL} } text1K<n<10K11 likes740 downloads1y agoHugging Face05TigreGotico /portuguese-unified-pronunciation-lexicon Portuguese Unified Pronunciation Lexicon A flat, single-row-per-pronunciation dataset merging Portuguese IPA transcriptions from three authoritative sources. Each row is a word × region × POS tuple with both broad phonemic (ipa_broad) and narrow phonetic (ipa_narrow) transcriptions normalized across sources. Source Words Convention Description Infopédia (Porto Editora) 102,685 Broad phonemic European Portuguese dictionary IPA Wiktionary (pt.wiktionary.org) 15,720… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/portuguese-unified-pronunciation-lexicon.texttext-generation100K<n<1M1 likes480 downloads2mo agoHugging Face06liaad /math_dataset_portugueseTo run generation code within 'mathematics_dataset\mathematics_dataset': Activate python venv .\.venv\Scripts\activate Requirements defined in requires.txt Run python generate_to_file.py --output_dir ds to generate dataset to directory \ds Had to change enconding when opening files to utf-8 so that some characters are allowed (ã õ é) To obtain dataset with the correct amount of rows: python generate_to_file.py --output_dir ds --per_train_module 1999998 --per_test_module 10000 This… See the full description on the dataset page: https://huggingface.co/datasets/liaad/math_dataset_portuguese.text1K<n<10K0 likes479 downloads3y agoHugging Face07freds0 /cml_tts_dataset_portugueseaudio10K<n<100K3 likes381 downloads2y agoHugging Face08AIML-TUDA /SLR-Bench-Portuguese 🧠 SLR-Bench-Portuguese: Scalable Logical Reasoning Benchmark (Portuguese Edition) SLR-Bench Multilingual Versions: SLR-Bench-Portuguese is the Portuguese-language pendant of the original SLR-Bench dataset. It follows the same symbolic structure, evaluation framework, and curriculum as the English version but provides all natural-language task prompts translated into Portuguese. This enables systematic evaluation and training of Large Language Models… See the full description on the dataset page: https://huggingface.co/datasets/AIML-TUDA/SLR-Bench-Portuguese.tabular10K<n<100K0 likes294 downloads4mo agoHugging Face09AdoCleanCode /SPEEED_s3_words_portuguese_0k-90ktext100K<n<1M0 likes256 downloads7mo agoHugging Face10AdoCleanCode /portuguese_tedx_alignedaudio10K<n<100K0 likes234 downloads7mo agoHugging Face11saillab /alpaca-portuguese-cleanedThis repository contains the dataset used for the TaCo paper. Please refer to the paper for more details: OpenReview If you have used our dataset, please cite it as follows: Citation @inproceedings{upadhayay2024taco, title={TaCo: Enhancing Cross-Lingual Transfer for Low-Resource Languages in {LLM}s through Translation-Assisted Chain-of-Thought Processes}, author={Bibek Upadhayay and Vahid Behzadan}, booktitle={5th Workshop on practical ML for limited/low resource settings, ICLR}, year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca-portuguese-cleaned.text10K<n<100K0 likes218 downloads2y agoHugging Face12portuguese-benchmark-datasets /story_cloze_pt Dataset Card for "story_cloze_pt" This is a portuguese translation of the xstory_cloze dataset. The translation was performed using the Google Translate API. This dataset follows the same structure as the original. text1K<n<10K1 likes195 downloads3y agoHugging Face13iara-project /raw_dataset_with_embeddings_bert-base-portuguese-cased-nli-assin-2 Dataset Card for "raw_dataset_with_embeddings_bert-base-portuguese-cased-nli-assin-2" More Information needed text100K<n<1M0 likes194 downloads3y agoHugging Face14bobboyms /portuguese-classic-books-adapted-to-modern-portuguese-brOkay, here is the improved and expanded text translated into American English, including the corrected citation format. Classic Portuguese Language Books Adapted to Modern Brazilian Portuguese Detailed Dataset Description This dataset presents a unique collection of texts derived from classic books of Portuguese language literature, with a strong representation of Brazilian authors. All selected works are in the public domain and were originally sourced from… See the full description on the dataset page: https://huggingface.co/datasets/bobboyms/portuguese-classic-books-adapted-to-modern-portuguese-br.texttext-generation10K<n<100K0 likes170 downloads1y agoHugging Face15Polygl0t /portuguese-eval-logs-olmo2-smollm3 Evaluation Logs on Portuguese Benchmarks for OLMo-2 and SmolLM3 These logs contain benchmark results across a suite of Portuguese-language tasks. The data consists of recordings of the performance of various 3 different models at different checkpoints throughout their pretraining runs: SmolLM3 OLMo-2-0425-1B OLMo-2-1124-7B Splits Each split (smollm3_3b, olmo2_1b, olmo2_7b) contains rows for model checkpoints and columns for benchmark scores (e.g., ASSIN2 RTE, ENEM… See the full description on the dataset page: https://huggingface.co/datasets/Polygl0t/portuguese-eval-logs-olmo2-smollm3.imagen<1K0 likes149 downloads7mo agoHugging Face16jvanz /portuguese_sentiment_analysisThis dataset is based on the dataset originally posted in Kaggle text1M<n<10M11 likes148 downloads4y agoHugging Face17minharegiao /portuguese-electionstextn<1K0 likes140 downloads2mo agoHugging Face18Helenice /corpus-biomedico-portugues Corpus Biomédico em Português Descrição Este dataset contém um corpus biomédico em português construído a partir de referências bibliográficas citadas em revisões sistemáticas publicadas no SciELO Brasil e no SciELO Portugal. Os documentos foram convertidos para texto simples (TXT) para utilização em experimentos de Recuperação de Informação e Processamento de Linguagem Natural. Metodologia O corpus foi construído a partir de 11 revisões… See the full description on the dataset page: https://huggingface.co/datasets/Helenice/corpus-biomedico-portugues.text100K<n<1M0 likes135 downloads2mo agoHugging Face19LCA-PORVID /portuguese_vidtext1M<n<10M0 likes123 downloads3y agoHugging Face20TigreGotico /portuguese-dialects-ipa-synthetic portuguese-dialects-ipa-synthetic 920 dialect-register sentences — 20 per lect across 46 lects: 41 Portuguese varieties (European regional, insular, Brazilian regional, African/Asian/border national norms, medieval stages) plus the other languages of Portugal and their kin: Mirandese (3 lects), Asturleonese of Portugal (Rionorese, Guadramilese) with Barranquenho, and Galician-Portuguese. Each row carries two IPA columns with distinct provenance. Schema sentence —… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/portuguese-dialects-ipa-synthetic.texttext-to-speechn<1K0 likes122 downloads2mo agoHugging Face21Paul /hatecheck-portuguese Dataset Card for Multilingual HateCheck Dataset Description Multilingual HateCheck (MHC) is a suite of functional tests for hate speech detection models in 10 different languages: Arabic, Dutch, French, German, Hindi, Italian, Mandarin, Polish, Portuguese and Spanish. For each language, there are 25+ functional tests that correspond to distinct types of hate and challenging non-hate. This allows for targeted diagnostic insights into model performance. For more details… See the full description on the dataset page: https://huggingface.co/datasets/Paul/hatecheck-portuguese.tabulartext-classification1K<n<10K13 likes110 downloads4y agoHugging Face22aimeri /ticuna-spanish-portuguese Ticuna (tca) – Spanish – Portuguese Corpus First text corpus for Ticuna (ISO 639-3 tca), a tonal language isolate of the Brazil/Colombia/Peru tri-border. Configs | Config | Rows | | parallel | train 43,248 / validation 596 / test 3,238 | | monolingual | train 46,545 / validation 298 / test 1,613 | | lexicon | train 10,419 / validation 568 / test 539 | | instructions | train 52,836 | | backtranslation | train 33,944 | The short version of what matters… See the full description on the dataset page: https://huggingface.co/datasets/aimeri/ticuna-spanish-portuguese.tabular100K<n<1M0 likes106 downloads28d agoHugging Face23arubenruben /portuguese-language-identification-rawtext10M<n<100M0 likes101 downloads3y agoHugging Face24stjiris /portuguese-legal-sentences-v0 Work developed as part of Project IRIS. Thesis: A Semantic Search System for Supremo Tribunal de Justiça Portuguese Legal Sentences Collection of Legal Sentences from the Portuguese Supreme Court of Justice The goal of this dataset was to be used for MLM and TSDAE Contributions @rufimelo99 If you use this work, please cite: @InProceedings{MeloSemantic, author="Melo, Rui and Santos, Pedro A. and Dias, Jo{\~a}o", editor="Moniz, Nuno and Vale, Zita and… See the full description on the dataset page: https://huggingface.co/datasets/stjiris/portuguese-legal-sentences-v0.text1M<n<10M14 likes100 downloads2y agoHugging Face25rishiraj /portuguesechat Dataset Card for Portuguese Chat We know that current English-first LLMs don’t work well for many other languages, both in terms of performance, latency, and speed. Building instruction datasets for non-English languages is an important challenge that needs to be solved. Dedicated towards addressing this problem, I release 3 new datasets rishiraj/portuguesechat, rishiraj/bengalichat & rishiraj/hindichat of 10,000 instructions and demonstrations each. This data can be used for… See the full description on the dataset page: https://huggingface.co/datasets/rishiraj/portuguesechat.texttext-generation10K<n<100K4 likes98 downloads3y agoHugging Face26LIACC /Emakhuwa-Portuguese-News-MT News Parallel Dataset for Emakhuwa of Mozambique This repository contains releases of parallel data for machine translation in Mozambican languages. Currently, it supports one language pair, Portuguese-Emakhuwa, Emakhuwa being the widely spoken language in Mozambique. Dataset Details Dataset Description Funded by: This dataset was created with support from Lacuna Fund, the world’s first collaborative effort to provide data scientists, researchers, and… See the full description on the dataset page: https://huggingface.co/datasets/LIACC/Emakhuwa-Portuguese-News-MT.tabulartranslation10K<n<100K0 likes97 downloads2y agoHugging Face27Polygl0t /portuguese-toxicity-qwen-annotations Annotations for the Portuguese-Toxicity classifier 📚 Dataset Summary This dataset contains the annotations used for training a toxicity classifier (Polygl0t/portuguese-bertabaporu-large-toxicity-classifier and Polygl0t/portuguese-bertimbau-toxicity-classifier). These annotations were generated by Qwen/Qwen2.5-32B-Instruct. Supported Tasks and Leaderboards This dataset can be used for the task of text classification, specifically for toxicity detection in… See the full description on the dataset page: https://huggingface.co/datasets/Polygl0t/portuguese-toxicity-qwen-annotations.texttext-classification100K<n<1M0 likes87 downloads7mo agoHugging Face28FreedomIntelligence /alpaca-gpt4-portugueseThe dataset is used in the research related to MultilingualSIFT. text10K<n<100K12 likes81 downloads3y agoHugging Face29Polygl0t /portuguese-edu-qwen-annotations Annotations for the Portuguese-Edu classifier 📚 Dataset Summary This dataset contains the annotations used for training an educational classifier (Polygl0t/portuguese-bertimbau-large-edu-classifier and Polygl0t/portuguese-bertimbau-edu-classifier). These annotations were generated by Qwen/Qwen2.5-32B-Instruct. Supported Tasks and Leaderboards This dataset can be used for the task of text classification, specifically for educational quality assessment in… See the full description on the dataset page: https://huggingface.co/datasets/Polygl0t/portuguese-edu-qwen-annotations.texttext-classification100K<n<1M0 likes80 downloads7mo agoHugging Face30fabiovilao /portuguese-blogs Dataset Details Blog-1 may include other languages in an unstructured text format without markdown. The latest one, Blog-6, is formatted in markdown and may contain less other languages text. Texts are separated by the string <|endoftext|>. Uses Training language models. Dataset Structure A simple text file with articles separated by <|endoftext|> between each text. Dataset Creation First semester of 2024. Bias, Risks, and Limitations… See the full description on the dataset page: https://huggingface.co/datasets/fabiovilao/portuguese-blogs.texttext-generation100M<n<1B0 likes75 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.