CoolFace
16 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01cnmoro /reasoning-v1-20m-portugueseglaiveai/reasoning-v1-20m translated to portuguese. texttext-generation10M<n<100M14 likes1.5k downloads1y agoHugging Face02TigreGotico /portuguese-unified-pronunciation-lexicon Portuguese Unified Pronunciation Lexicon A flat, single-row-per-pronunciation dataset merging Portuguese IPA transcriptions from three authoritative sources. Each row is a word × region × POS tuple with both broad phonemic (ipa_broad) and narrow phonetic (ipa_narrow) transcriptions normalized across sources. Source Words Convention Description Infopédia (Porto Editora) 102,685 Broad phonemic European Portuguese dictionary IPA Wiktionary (pt.wiktionary.org) 15,720… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/portuguese-unified-pronunciation-lexicon.texttext-generation100K<n<1M1 likes416 downloads3mo agoHugging Face03bobboyms /portuguese-classic-books-adapted-to-modern-portuguese-brOkay, here is the improved and expanded text translated into American English, including the corrected citation format. Classic Portuguese Language Books Adapted to Modern Brazilian Portuguese Detailed Dataset Description This dataset presents a unique collection of texts derived from classic books of Portuguese language literature, with a strong representation of Brazilian authors. All selected works are in the public domain and were originally sourced from… See the full description on the dataset page: https://huggingface.co/datasets/bobboyms/portuguese-classic-books-adapted-to-modern-portuguese-br.texttext-generation10K<n<100K0 likes170 downloads1y agoHugging Face04mazafard /portuguese-ocr-datasettask_categories: image-to-text task_ids: optical-character-recognition text-recognition Portuguese OCR Dataset A comprehensive dataset for Portuguese OCR (Optical Character Recognition) generated from classic Portuguese literature with diverse fonts and visual styles. Dataset Description This dataset contains 20000 text images for OCR training, created from Portuguese books from Project Gutenberg. Each image contains a complete Portuguese sentence with proper… See the full description on the dataset page: https://huggingface.co/datasets/mazafard/portuguese-ocr-dataset.imageimage-to-textn<1K2 likes143 downloads1y agoHugging Face05rishiraj /portuguesechat Dataset Card for Portuguese Chat We know that current English-first LLMs don’t work well for many other languages, both in terms of performance, latency, and speed. Building instruction datasets for non-English languages is an important challenge that needs to be solved. Dedicated towards addressing this problem, I release 3 new datasets rishiraj/portuguesechat, rishiraj/bengalichat & rishiraj/hindichat of 10,000 instructions and demonstrations each. This data can be used for… See the full description on the dataset page: https://huggingface.co/datasets/rishiraj/portuguesechat.texttext-generation10K<n<100K4 likes97 downloads3y agoHugging Face06fabiovilao /portuguese-blogs Dataset Details Blog-1 may include other languages in an unstructured text format without markdown. The latest one, Blog-6, is formatted in markdown and may contain less other languages text. Texts are separated by the string <|endoftext|>. Uses Training language models. Dataset Structure A simple text file with articles separated by <|endoftext|> between each text. Dataset Creation First semester of 2024. Bias, Risks, and Limitations… See the full description on the dataset page: https://huggingface.co/datasets/fabiovilao/portuguese-blogs.texttext-generation100M<n<1B0 likes75 downloads2y agoHugging Face07rhaymison /orca-math-portuguese-64ktranslated for: Repository: microsoft/orca-math-word-problems-200k Paper: Orca-Math: Unlocking the potential of SLMs in Grade School Math texttext-generation10K<n<100K6 likes63 downloads2y agoHugging Face08leeaandrob /mirror-rhaymison__orca-math-portuguese-64ktranslated for: Repository: microsoft/orca-math-word-problems-200k Paper: Orca-Math: Unlocking the potential of SLMs in Grade School Math texttext-generation10K<n<100K0 likes60 downloads3mo agoHugging Face09nelsondiasandre /portuguese-qa-instruct-500 Portuguese Q&A Instruction Dataset (500 pairs) 500 Portuguese (PT-PT) question-answer pairs formatted for instruction fine-tuning of language models. Dataset Structure Each example has three columns: Column Description Example instruction The question in Portuguese "Qual e a capital de Portugal?" response The answer in Portuguese "A capital de Portugal e Lisboa." text Pre-formatted instruction template (see below) "<|im_start|>user\n..."… See the full description on the dataset page: https://huggingface.co/datasets/nelsondiasandre/portuguese-qa-instruct-500.textquestion-answeringn<1K0 likes49 downloads4mo agoHugging Face10TaigoPedrosa /PortugueseMMLU Dataset Components The dataset is partitioned into three discrete tables stored in CSV or Parquet format: Questions Recipes Evaluation Results Each component is described in detail below. Questions area domain question_number An integer index uniquely identifying each question inside the knowledge domain. translation_method English, Google Translate, GPT-3.5-Turbo, GPT-4o, Human question option_a, option_b, option_c, option_d Recipes area… See the full description on the dataset page: https://huggingface.co/datasets/TaigoPedrosa/PortugueseMMLU.tabularquestion-answering100K<n<1M0 likes36 downloads1y agoHugging Face11TigreGotico /portuguese_phonetic_lexicon 📚 Portuguese Phonetic Lexicon Dataset This dataset contains phonetic and morphological information for Portuguese words, collected from the Portal da Língua Portuguesa. It was generated by scraping the site across multiple Portuguese-speaking regions and dialects. 🌍 Regional Coverage The dataset includes words as spoken in ten regional variants: 🇵🇹 Lisbon (Standard and Non-Standard) 🇦🇴 Luanda 🇧🇷 Rio de Janeiro (Standard and Non-Standard) 🇧🇷 São Paulo (Standard… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/portuguese_phonetic_lexicon.texttext-classification100K<n<1M0 likes27 downloads7mo agoHugging Face12LIACC /Emakhuwa-Portuguese-OCR-post-correctionBibTeX: The dataset paper was published in EMNLP 2024. Please cite as: @inproceedings{ali-etal-2024-building, title = "Building Resources for Emakhuwa: Machine Translation and News Classification Benchmarks", author = "Ali, Felermino D. M. A. and Lopes Cardoso, Henrique and Sousa-Silva, Rui", editor = "Al-Onaizan, Yaser and Bansal, Mohit and Chen, Yun-Nung", booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language… See the full description on the dataset page: https://huggingface.co/datasets/LIACC/Emakhuwa-Portuguese-OCR-post-correction.imagetranslationn<1K0 likes19 downloads2y agoHugging Face13bobboyms /realtor-conversational-portuguese_br Realtor Conversational (Portuguese BR) - Realtor-Client Conversation Dataset Detailed Dataset Description Introduction: This dataset, named "Realtor Conversational (Portuguese BR)", offers rich and detailed simulations of conversational interactions between real estate agents (realtors) and clients in Brazil. Generated using the advanced language model gpt-4o-mini, the data is synthetic but designed to mirror the dynamics, vocabulary, and common scenarios found in the… See the full description on the dataset page: https://huggingface.co/datasets/bobboyms/realtor-conversational-portuguese_br.text-generation1K<n<10K0 likes18 downloads1y agoHugging Face14AxionLab-official /Reasoning-MiniGPT-brazilian-portuguesetexttext-generationn<1K0 likes17 downloads10mo agoHugging Face15rishabmishrasensation /realtor-conversational-portuguese_br Realtor Conversational (Portuguese BR) - Realtor-Client Conversation Dataset Detailed Dataset Description Introduction: This dataset, named "Realtor Conversational (Portuguese BR)", offers rich and detailed simulations of conversational interactions between real estate agents (realtors) and clients in Brazil. Generated using the advanced language model gpt-4o-mini, the data is synthetic but designed to mirror the dynamics, vocabulary, and common scenarios found in the… See the full description on the dataset page: https://huggingface.co/datasets/rishabmishrasensation/realtor-conversational-portuguese_br.text-generation1K<n<10K0 likes17 downloads5mo agoHugging Face16TigreGotico /portuguese-sentences-synthetic-g2p Dataset Card for 'TigreGotico/portuguese_g2p' Dataset Description Dataset Summary TigreGotico/portuguese_g2p is a Grapheme-to-Phoneme (G2P) dataset for Portuguese, offering phonetic transcriptions for sentences across ten different regional variants. It is derived from the portuguese_phonetic_lexicon and is designed to aid in the development of robust Speech Recognition (ASR) and Text-to-Speech (TTS) models that account for dialectal variation in Portuguese.… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/portuguese-sentences-synthetic-g2p.texttext-classification10K<n<100K0 likes7 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.