CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01scholarly-shadows-syndicate /2wikimultihopqa_with_q_gpt35 2WikiMultihopQA Dataset with GPT-3.5 Generated Questions Overview This repository hosts an enhanced version of the 2WikiMultihopQA dataset, where each supporting sentence in the dataset has been supplemented with questions generated using OpenAI's GPT-3.5 turbo API. The aim is to provide a richer context for each entry, potentially benefiting various NLP tasks, such as question answering and context understanding. Dataset Format Each entry in the dataset is… See the full description on the dataset page: https://huggingface.co/datasets/scholarly-shadows-syndicate/2wikimultihopqa_with_q_gpt35.text10K<n<100K2 likes1.5k downloads3y agoHugging Face02potsawee /wiki_bio_gpt3_hallucination Dataset Card for WikiBio GPT-3 Hallucination Dataset GitHub repository: https://github.com/potsawee/selfcheckgpt Paper: SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models Dataset Summary We generate Wikipedia-like passages using GPT-3 (text-davinci-003) using the prompt: This is a Wikipedia passage about {concept} where concept represents an individual from the WikiBio dataset. We split the generated passages into… See the full description on the dataset page: https://huggingface.co/datasets/potsawee/wiki_bio_gpt3_hallucination.texttext-classificationn<1K30 likes742 downloads3y agoHugging Face03gpt3mix /sst2text1K<n<10K5 likes298 downloads5y agoHugging Face04princeton-nlp /QuRating-GPT3.5-Judgments-Test7140 pairwise judgments across 4 criteria and 6 domains obtained by prompting GPT-3.5-turbo-0613 for evaluating QuRater models. From the paper: QuRating: Selecting High-Quality Data for Training Language Models Guidance on Responsible Use In the paper, we document various types of bias that are present in the quality ratings/QuRater model (biases related to domains, topics, social roles, regions and languages - see Section 6 of the paper), which are likely reflected in the LLM judgments. Hence… See the full description on the dataset page: https://huggingface.co/datasets/princeton-nlp/QuRating-GPT3.5-Judgments-Test.text1K<n<10K1 likes128 downloads2y agoHugging Face05imoxto /prompt_injection_hackaprompt_gpt35 Dataset Card for "prompt_injection_hackaprompt_gpt35" More Information needed text100K<n<1M7 likes105 downloads3y agoHugging Face06scholarly-shadows-syndicate /hotpotqa_with_qa_gpt35 HotpotQA Dataset with GPT-3.5 Generated Questions Overview This repository hosts an enhanced version of the HotpotQA dataset, where each supporting sentence in the dataset has been supplemented with questions generated using OpenAI's GPT-3.5 turbo API. The aim is to provide a richer context for each entry, potentially benefiting various NLP tasks, such as question answering and context understanding. Dataset Format Each entry in the dataset is formatted as… See the full description on the dataset page: https://huggingface.co/datasets/scholarly-shadows-syndicate/hotpotqa_with_qa_gpt35.text10K<n<100K1 likes78 downloads3y agoHugging Face07skeskinen /TinyStories-GPT3.5 Dataset Card for "TinyStories-GPT3.5" More Information needed text1M<n<10M1 likes67 downloads3y agoHugging Face08gpt3mix /rt20text1K<n<10K0 likes62 downloads5y agoHugging Face09alex2awesome /city-council-gpt3-silver-standard-summaries Dataset Card for "city-council-gpt3-silver-standard-summaries" More Information needed text1K<n<10K0 likes58 downloads3y agoHugging Face10princeton-nlp /QuRating-GPT3.5-Judgments250K thousand pairwise judgments across 4 criteria obtained by prompting GPT-3.5-turbo-0613. From the paper: QuRating: Selecting High-Quality Data for Training Language Models Guidance on Responsible Use In the paper, we document various types of bias that are present in the quality ratings/QuRater model (biases related to domains, topics, social roles, regions and languages - see Section 6 of the paper), which are likely reflected in the LLM judgments. Hence, be aware that data selection with… See the full description on the dataset page: https://huggingface.co/datasets/princeton-nlp/QuRating-GPT3.5-Judgments.text100K<n<1M6 likes50 downloads2y agoHugging Face11lighteval /GPT3_unscramble Dataset Card for "unscramble_GPT3" More Information needed text10K<n<100K1 likes44 downloads3y agoHugging Face12sinatra-rd /portuguese-gpt3.5-fine-tuningtext1K<n<10K0 likes44 downloads2y agoHugging Face13facat /gsm8k-gpt35 Dataset Card for "gsm8k-gpt35" More Information needed text1K<n<10K0 likes39 downloads3y agoHugging Face14pietrolesci /gpt3_nli Overview Original dataset available here. Debiased dataset generated with GPT-3. Dataset curation All string columns are stripped. Labels are encoded with the following mapping {"entailment": 0, "neutral": 1, "contradiction": 2} Code to create the dataset import pandas as pd from datasets import Dataset, ClassLabel, Value, Features import json # load data with open("data/dataset.jsonl", "r") as fl: df = pd.DataFrame([json.loads(line) for line in fl])… See the full description on the dataset page: https://huggingface.co/datasets/pietrolesci/gpt3_nli.tabular10K<n<100K3 likes33 downloads4y agoHugging Face15Somnia /GPT3-Token-Encodertabularn<1K0 likes28 downloads4y agoHugging Face16maywell /ko-gpt3_14ktext10K<n<100K4 likes28 downloads3y agoHugging Face17NITHUB-AI /Ehn-bible-bbc-gpt3.5 Dataset Card for Ehn-Bible-BBC-GPT3.5 Dataset Summary This dataset card contains parallel Nigerian Pidgin and English sentences split into three files, namely: train.csv, valid.csv and test.csv. The original data was split in the ratio of 8:1:1 to obtain these files. Supported Tasks and Leaderboards Language Translation Language Identification Languages English Nigerian Pidgin Dataset Structure Data Instances Data… See the full description on the dataset page: https://huggingface.co/datasets/NITHUB-AI/Ehn-bible-bbc-gpt3.5.texttext-classification10K<n<100K1 likes27 downloads3y agoHugging Face18conghao /gpt3.5-jinyong-style OpenAI微调医学数据集 开源数据集存在与预训练模型数据集相似度较高的问题,导致微调结果不明显,效果不突出,因此就需要构建某种风格的专属数据集。 因此,构建了带金庸创作风格的医学数据集,可直接用于openai的微调使用。 微调实验结果证明,当输入医学问题时,微调模型回答的风格为金庸创作风格回答。 Question [{'role': 'system', 'content': '你是我的私人医生助手,你要回答我的健康问题。'}, {'role': 'user', 'content': '我睡眠不足,很头痛'}] Answer 少侠保重身体。江湖人称“睡眠不足,头痛难受”,此乃常见之症。吾乃医术高强之侠士,愿为你指点迷津。睡眠不足,多因忧虑纷扰、劳累过度所致。当以调整生活作息,保持规律之睡眠时间为要。且可饮用一些花草茶,如菊花、薰衣草,以宁心安神。如症状持续,还需请教名医,以寻得更为妥善之疗法。 可以发现,模型的回答已经学习到金庸创作风格的能力。 textquestion-answering1K<n<10K3 likes27 downloads3y agoHugging Face19open-llm-leaderboard-old /details_TurkuNLP__gpt3-finnish-small Dataset Card for Evaluation run of TurkuNLP/gpt3-finnish-small Dataset Summary Dataset automatically created during the evaluation run of model TurkuNLP/gpt3-finnish-small on the Open LLM Leaderboard. The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_TurkuNLP__gpt3-finnish-small.0 likes27 downloads3y agoHugging Face20sinatra-rd /MetaMathQA-40K-GPT3.5MetaMathQA-40K adapted to the GPT3.5 dataset format in JSONL for Fine-tuning. Following the following model: {"messages": [{"role": "system", "content": ""}, {"role": "user", "content": ""}, {"role": "assistant", "content": ""}]} text10K<n<100K0 likes26 downloads3y agoHugging Face21open-llm-leaderboard-old /details_TurkuNLP__gpt3-finnish-large Dataset Card for Evaluation run of TurkuNLP/gpt3-finnish-large Dataset Summary Dataset automatically created during the evaluation run of model TurkuNLP/gpt3-finnish-large on the Open LLM Leaderboard. The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_TurkuNLP__gpt3-finnish-large.0 likes23 downloads3y agoHugging Face22dim /lmsys_chatbot_arena_conversations_gpt4_gpt35turbo_claudy Dataset Card for "lmsys_chatbot_arena_conversations_gpt4_gpt-3.5-turbo_claudy" More Information needed text10K<n<100K0 likes23 downloads3y agoHugging Face23xPXXX /gpt3_generation_sampletext1K<n<10K0 likes23 downloads2y agoHugging Face24VGraf /self_compare_alpacaeval_related_8maxturns_complete_805_gpt3.5_cleanedtext1K<n<10K0 likes23 downloads6mo agoHugging Face25open-llm-leaderboard-old /details_TurkuNLP__gpt3-finnish-13B Dataset Card for Evaluation run of TurkuNLP/gpt3-finnish-13B Dataset Summary Dataset automatically created during the evaluation run of model TurkuNLP/gpt3-finnish-13B on the Open LLM Leaderboard. The dataset is composed of 3 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_TurkuNLP__gpt3-finnish-13B.0 likes22 downloads3y agoHugging Face26Gokce /Generated_Restaurant_Reviews_GPT3.5 license: cc-by-4.0 task_categories: text-classification language: tr tags: food Generated Review size_categories: 1K<n<10K text1K<n<10K0 likes21 downloads3y agoHugging Face27PanoEvJ /GPT3.5_summarization_preference_RLAIF Dataset Card for "GPT3.5_summarization_preference_RLAIF" More Information needed textn<1K0 likes21 downloads3y agoHugging Face28lorinma /EvolInstruct_zh_GPT3.5私以为这并不是一次很成功的尝试。猜测一个主要原因是prompt依然是英文的,只是增加了the locale of the prompt is mainland china. 因为WizardLM系列长期霸榜LLM开源榜,一直很好奇EvolInstruct在英文世界表现出的对于复杂prompt的应对能力。 目前中文没有原生的EvolInstruct,仅有两个翻译版本 1 2。 故浅浅尝试复现中文版本。代码参照 3 但无奈接口实在是太贵,且生成的时间很长。所以如果有能够提供GPT-4 API资源的,我很乐意将这个量级撑到50K+并进行公开。 一共有3个文件: combined_seed_correct.json 是使用的基础种子任务371条,alpaca格式。使用了 Belle的中文种子任务175条。并且参照了 4 增加了ShareGPT的数据以更接近真实世界的用法,掺入了 Wildchat-zh抽样196条,多轮对话只采用第一个有意义的问答对。 231213_ChineseEvolInstruct_140_gpt-4-1106-preview.json… See the full description on the dataset page: https://huggingface.co/datasets/lorinma/EvolInstruct_zh_GPT3.5.text-generation10K<n<100K5 likes20 downloads3y agoHugging Face29andersonbcdefg /synthetic_tuples_gpt35_turbotexttext-retrieval100K<n<1M17 likes20 downloads3y agoHugging Face30andersonbcdefg /synthetic_tuples_gpt35_deduptext100K<n<1M0 likes20 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.