datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
2wikimultihopqa_with_q_gpt35
2WikiMultihopQA Dataset with GPT-3.5 Generated Questions
Overview
This repository hosts an enhanced version of the 2WikiMultihopQA dataset, where each supporting sentence in the dataset has been supplemented with questions generated using OpenAI's GPT-3.5 turbo API. The aim is to provide a richer context for each entry, potentially benefiting various NLP tasks, such as question answering and context understanding.
Dataset Format
Each entry in the dataset is… See the full description on the dataset page: https://huggingface.co/datasets/scholarly-shadows-syndicate/2wikimultihopqa_with_q_gpt35.wiki_bio_gpt3_hallucination
Dataset Card for WikiBio GPT-3 Hallucination Dataset
GitHub repository: https://github.com/potsawee/selfcheckgpt
Paper: SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models
Dataset Summary
We generate Wikipedia-like passages using GPT-3 (text-davinci-003) using the prompt: This is a Wikipedia passage about {concept} where concept represents an individual from the WikiBio dataset.
We split the generated passages into… See the full description on the dataset page: https://huggingface.co/datasets/potsawee/wiki_bio_gpt3_hallucination.sst2hotpotqa_with_qa_gpt35
HotpotQA Dataset with GPT-3.5 Generated Questions
Overview
This repository hosts an enhanced version of the HotpotQA dataset, where each supporting sentence in the dataset has been supplemented with questions generated using OpenAI's GPT-3.5 turbo API. The aim is to provide a richer context for each entry, potentially benefiting various NLP tasks, such as question answering and context understanding.
Dataset Format
Each entry in the dataset is formatted as… See the full description on the dataset page: https://huggingface.co/datasets/scholarly-shadows-syndicate/hotpotqa_with_qa_gpt35.QuRating-GPT3.5-Judgments-Test7140 pairwise judgments across 4 criteria and 6 domains obtained by prompting GPT-3.5-turbo-0613 for evaluating QuRater models.
From the paper: QuRating: Selecting High-Quality Data for Training Language Models
Guidance on Responsible Use
In the paper, we document various types of bias that are present in the quality ratings/QuRater model (biases related to domains, topics, social roles, regions and languages - see Section 6 of the paper),
which are likely reflected in the LLM judgments.
Hence… See the full description on the dataset page: https://huggingface.co/datasets/princeton-nlp/QuRating-GPT3.5-Judgments-Test.prompt_injection_hackaprompt_gpt35
Dataset Card for "prompt_injection_hackaprompt_gpt35"
More Information needed
rt20TinyStories-GPT3.5
Dataset Card for "TinyStories-GPT3.5"
More Information needed
city-council-gpt3-silver-standard-summaries
Dataset Card for "city-council-gpt3-silver-standard-summaries"
More Information needed
QuRating-GPT3.5-Judgments250K thousand pairwise judgments across 4 criteria obtained by prompting GPT-3.5-turbo-0613.
From the paper: QuRating: Selecting High-Quality Data for Training Language Models
Guidance on Responsible Use
In the paper, we document various types of bias that are present in the quality ratings/QuRater model (biases related to domains, topics, social roles, regions and languages - see Section 6 of the paper),
which are likely reflected in the LLM judgments.
Hence, be aware that data selection with… See the full description on the dataset page: https://huggingface.co/datasets/princeton-nlp/QuRating-GPT3.5-Judgments.portuguese-gpt3.5-fine-tuningGPT3_unscramble
Dataset Card for "unscramble_GPT3"
More Information needed
gpt3_nli
Overview
Original dataset available here. Debiased dataset generated with GPT-3.
Dataset curation
All string columns are stripped. Labels are encoded with the following mapping
{"entailment": 0, "neutral": 1, "contradiction": 2}
Code to create the dataset
import pandas as pd
from datasets import Dataset, ClassLabel, Value, Features
import json
# load data
with open("data/dataset.jsonl", "r") as fl:
df = pd.DataFrame([json.loads(line) for line in fl])… See the full description on the dataset page: https://huggingface.co/datasets/pietrolesci/gpt3_nli.gsm8k-gpt35
Dataset Card for "gsm8k-gpt35"
More Information needed
MetaMathQA-40K-GPT3.5MetaMathQA-40K adapted to the GPT3.5 dataset format in JSONL for Fine-tuning. Following the following model:
{"messages": [{"role": "system", "content": ""}, {"role": "user", "content": ""}, {"role": "assistant", "content": ""}]}
ko-gpt3_14kgpt3.5-jinyong-style
OpenAI微调医学数据集
开源数据集存在与预训练模型数据集相似度较高的问题,导致微调结果不明显,效果不突出,因此就需要构建某种风格的专属数据集。
因此,构建了带金庸创作风格的医学数据集,可直接用于openai的微调使用。
微调实验结果证明,当输入医学问题时,微调模型回答的风格为金庸创作风格回答。
Question
[{'role': 'system', 'content': '你是我的私人医生助手,你要回答我的健康问题。'}, {'role': 'user', 'content': '我睡眠不足,很头痛'}]
Answer
少侠保重身体。江湖人称“睡眠不足,头痛难受”,此乃常见之症。吾乃医术高强之侠士,愿为你指点迷津。睡眠不足,多因忧虑纷扰、劳累过度所致。当以调整生活作息,保持规律之睡眠时间为要。且可饮用一些花草茶,如菊花、薰衣草,以宁心安神。如症状持续,还需请教名医,以寻得更为妥善之疗法。
可以发现,模型的回答已经学习到金庸创作风格的能力。
Ehn-bible-bbc-gpt3.5
Dataset Card for Ehn-Bible-BBC-GPT3.5
Dataset Summary
This dataset card contains parallel Nigerian Pidgin and English sentences split into three files, namely: train.csv, valid.csv and test.csv.
The original data was split in the ratio of 8:1:1 to obtain these files.
Supported Tasks and Leaderboards
Language Translation
Language Identification
Languages
English
Nigerian Pidgin
Dataset Structure
Data Instances
Data… See the full description on the dataset page: https://huggingface.co/datasets/NITHUB-AI/Ehn-bible-bbc-gpt3.5.lmsys_chatbot_arena_conversations_gpt4_gpt35turbo_claudy
Dataset Card for "lmsys_chatbot_arena_conversations_gpt4_gpt-3.5-turbo_claudy"
More Information needed
gpt3_generation_sampleself_compare_alpacaeval_related_8maxturns_complete_805_gpt3.5_cleanedfull_dataset_1550_lines_invoice_contract_mail_GPT3.5_train
Dataset Card for "full_dataset_1550_lines_invoice_contract_mail_GPT3.5_train"
More Information needed
Generated_Restaurant_Reviews_GPT3.5
license: cc-by-4.0
task_categories:
text-classification
language:
tr
tags:
food
Generated Review
size_categories:
1K<n<10K
GPT3.5_summarization_preference_RLAIF
Dataset Card for "GPT3.5_summarization_preference_RLAIF"
More Information needed
viet-self-instruct-gpt3.5-v2Dataset Name: viet-self-instruct-gpt3.5-v2
Description: The viet-self-instruct-gpt3.5-v2 dataset is generated using the self-instruct method as described in the Self-Instruct GitHub repository. It contains Vietnamese instruction data specifically tailored for training language models.
Source: viet-self-instruct-gpt3.5-v2 on Hugging Face Datasets
Method: The dataset was generated using the self-instruct method described in the Self-Instruct GitHub repository.
License: Please refer to the… See the full description on the dataset page: https://huggingface.co/datasets/kimnt93/viet-self-instruct-gpt3.5-v2.synthetic_tuples_gpt35_turbosynthetic_tuples_gpt35_dedupFOLIO_by_paraphrased_gpt3.5GPT35TurboFuncCallOSS_Instruct_Python_zh_GPT35I use GPT-3.5-turbo's API to generate a Chinese-Python educational dataset as Magicoder did with the lack of coding data in Chinese. The license also follows Magicoder-OSS-Instruct's mit.
One should notice:
the quality is not promised and has not been tested.
this dataset is mainly in Python since the seeds are Python only, but checking the results shows there are also other languages like PHP, Java, Ruby, etc. Substrings like "```{lang}" may be useful for filtering or categorizing.
