datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cosmos_qa_ptbr
Cosmos QA Português
Este dataset é uma tradução para português do Cosmos QA, que originalmente é na língua inglesa.
A tradução foi feita automaticamente usando o GPT-3.5-turbo, logo pode ter erros que não foram notados numa análise superficial.
Se atente ao uso.
Dataset Card for cosmos_qa
Licensing Information
The data is distributed under the CC BY 4.0 license.
Source Data Citation INformation
@inproceedings{huang-etal-2019-cosmos,
title =… See the full description on the dataset page: https://huggingface.co/datasets/heloisy/cosmos_qa_ptbr.Sentiments-FinBERT-PT-BR
Dataset
A manually annotated dataset was created to enable supervised training for the FinBERT-PT-BR model, which focuses on sentiment analysis of Brazilian Portuguese financial texts.
More than 1.4 million financial news texts in Portuguese were collected and used for the initial language modeling phase. From this corpus, a sample of 1,000 texts was manually annotated with sentiment labels.
Annotation Process
Three annotators participated in the process.
All texts were… See the full description on the dataset page: https://huggingface.co/datasets/lucas-leme/Sentiments-FinBERT-PT-BR.SQuAD-pt_BR-V1.1_
Dataset Card para o SQuAD 1.1 em Português Brasil
O conjunto de dados "Stanford Question Answering Dataset" (SQuAD),
para tarefa de perguntas e respostas extrativas, foi desenvolvido em 2016. Ele utiliza perguntas geradas a partir de
536 artigos da Wikipedia* com mais de 100.000 linhas de dados. É construído na forma de uma pergunta e um contexto dos artigos da
Wikipedia contendo a resposta à pergunta. [1]Originalmente este dataset foi construído no idioma inglês, contudo, o grupo… See the full description on the dataset page: https://huggingface.co/datasets/vsvasconcelos/SQuAD-pt_BR-V1.1_.SQuAD-pt_BR-V1.1
Dataset Card for Dataset Name
O conjunto de dados "Stanford Question Answering Dataset" (SQuAD), para tarefa de perguntas e respostas extrativas, foi desenvolvido em 2016. Ele utiliza perguntas geradas
a partir de 536 artigos da Wikipedia com mais de 100.000 linhas de dados. É construído na forma de uma pergunta e um contexto dos artigos da Wikipedia contendo a resposta
à pergunta.
Originalmente este dataset foi construído no idioma inglês, contudo, o grupo Deep Learning Brasil… See the full description on the dataset page: https://huggingface.co/datasets/vsvasconcelos/SQuAD-pt_BR-V1.1.banking77-pt-br
Resumo do conjunto de dados
Tradução revisada para o português brasileiro do conjunto de dados BANKING77. O conjunto é composto por consultas online realizadas a sistemas conversacionais de bancos, rotuladas de acordo com suas intenções correspondentes. De acordo com a documentação original, as 13.083 consultas que compõem o conjunto são referentes a atendimentos ao cliente, e foram rotuladas em 77 intenções diferentes. O foco principal é suportar análises em domínio específico e… See the full description on the dataset page: https://huggingface.co/datasets/c2d-usp/banking77-pt-br.PTBR_BASEextracurricular-certificates-declarations-ptBR
Extracurricular Certificates and Declarations Classification
Dataset Summary
This dataset contains 603 academic proof documents (certificates, declarations, scientific papers, and other academic documents) in Brazilian Portuguese and English, labeled according to the pre-defined groups of the extracurricular activities accreditation regulation of the Computer Science undergraduate program at Universidade de Brasília (UnB).
Brazilian undergraduate programs allow… See the full description on the dataset page: https://huggingface.co/datasets/gvic-unb/extracurricular-certificates-declarations-ptBR.ptbr-quora-translated
Dataset Summary
The Quora dataset is composed of question pairs, and the task is to determine if the questions are paraphrases of each
other (have the same meaning). The dataset was translated to Portuguese using the model seamless-m4t-medium.
Languages
Portuguese
ptbr-irony-idioms-regionalism
Sotaques Digitais — Benchmark LLM Português Brasileiro
Dataset Summary
Sotaques Digitais é um benchmark de avaliação de competência pragmática e cultural para LLMs em português brasileiro. O dataset contém 90 cenários de teste distribuídos em três categorias linguísticas, construídos a partir de contextos reais (redes sociais, WhatsApp, atendimento ao cliente, avaliações de produto, ambiente de trabalho).
O benchmark foi desenvolvido para a pesquisa "Sotaques… See the full description on the dataset page: https://huggingface.co/datasets/ramondomiingos/ptbr-irony-idioms-regionalism.wikipedia-domain-labels-ptbr
Wikipedia Domain Labels - PTBR
This dataset is a Portuguese (Brazilian) translation of the Wikipedia Domain Labels dataset by NeuML.
Full attribution: This work is entirely based on the original dataset:
NeuML/wikipedia-domain-labels — the original English dataset
Domain labels for the txtai-wikipedia-slim embeddings database (Top 100K most viewed Wikipedia articles)
The original dataset was built using a zero-shot-classification model applied to each of the Top 100K most… See the full description on the dataset page: https://huggingface.co/datasets/cnmoro/wikipedia-domain-labels-ptbr.twitter-sentiment-pt-BR-md-2-limdb-pt-brcourse-feedback-pt_brSalesFeedback-PTBRspell_correction_datasets_pt_brlatamqa_mcq_pt-br
LatamQA
LatamQA is a cultural knowledge benchmark designed to evaluate Large Language Models on Latin American contexts. The dataset addresses the critical gap in bias detection resources for non-English languages and underrepresented cultures. Built from 26,000+ Wikipedia articles and structured using Wikidata's knowledge graph with expert guidance from social scientists, LatamQA contains over 26,000 multiple-choice questions covering the diverse popular and social cultures of… See the full description on the dataset page: https://huggingface.co/datasets/inria-chile/latamqa_mcq_pt-br.AO1990_pt-BREnglish-PTBRhuman_ai_pt-brlatamqa_articles_pt-br
LatamQA
LatamQA is a cultural knowledge benchmark designed to evaluate Large Language Models on Latin American contexts. The dataset addresses the critical gap in bias detection resources for non-English languages and underrepresented cultures. Built from 26,000+ Wikipedia articles and structured using Wikidata's knowledge graph with expert guidance from social scientists, LatamQA contains over 26,000 multiple-choice questions covering the diverse popular and social cultures of… See the full description on the dataset page: https://huggingface.co/datasets/inria-chile/latamqa_articles_pt-br.
