datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
orca-math-portuguese-64ktranslated for:
Repository: microsoft/orca-math-word-problems-200k
Paper: Orca-Math: Unlocking the potential of
SLMs in Grade School Math
mirror-rhaymison__orca-math-portuguese-64ktranslated for:
Repository: microsoft/orca-math-word-problems-200k
Paper: Orca-Math: Unlocking the potential of
SLMs in Grade School Math
portuguese-qa-instruct-500
Portuguese Q&A Instruction Dataset (500 pairs)
500 Portuguese (PT-PT) question-answer pairs formatted for instruction fine-tuning of language models.
Dataset Structure
Each example has three columns:
Column
Description
Example
instruction
The question in Portuguese
"Qual e a capital de Portugal?"
response
The answer in Portuguese
"A capital de Portugal e Lisboa."
text
Pre-formatted instruction template (see below)
"<|im_start|>user\n..."… See the full description on the dataset page: https://huggingface.co/datasets/nelsondiasandre/portuguese-qa-instruct-500.qa-portuguese-small
QA-PORTUGUESE-SMALL
Dataset Description
The qa-portuguese-small dataset is a collection of 500,000 question-answer pairs in Portuguese designed for Question Answering (QA) tasks. The dataset includes questions based on a wide variety of domains, such as news, general knowledge, and everyday facts, and provides corresponding answers in natural language.
The dataset is intended for training and evaluating machine learning models that can answer questions in Portuguese… See the full description on the dataset page: https://huggingface.co/datasets/Jpzinn654/qa-portuguese-small.PortugueseMMLU
Dataset Components
The dataset is partitioned into three discrete tables stored in CSV or Parquet format:
Questions
Recipes
Evaluation Results
Each component is described in detail below.
Questions
area
domain
question_number
An integer index uniquely identifying each question inside the knowledge domain.
translation_method
English, Google Translate, GPT-3.5-Turbo, GPT-4o, Human
question
option_a, option_b, option_c, option_d
Recipes
area… See the full description on the dataset page: https://huggingface.co/datasets/TaigoPedrosa/PortugueseMMLU.dataset-portuguese-aira-v2-Gemma-formatDataset Aira para o formato do Modelo Gemma
Resumo do Dataset
Este conjunto de dados contém uma coleção de conversas individuais entre um assistente e um usuário.
As conversas foram geradas pelas interações do usuário com modelos já ajustados (ChatGPT, LLama 2, Open-Assistant, etc).
O conjunto de dados está disponível em português (tem a versão em Inglês que ainda não tratei). Mas você pode baixar do
repositório de Nicholas Kluge Corrêa tanto a versão em Português e
a versão em… See the full description on the dataset page: https://huggingface.co/datasets/EddyGiusepe/dataset-portuguese-aira-v2-Gemma-format.bird-sql-portuguese
BIRD-SQL - Versão em Português
Este repositório contém a tradução para português da partição de treino e desenvolvimento do benchmark BIRD-SQL, um benchmark para a tarefa de Text-to-SQL.
