datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GPT-wiki-intro
GPT Wiki Intro
Overview
Dataset for training models to classify human written vs GPT/ChatGPT generated text.
This dataset contains Wikipedia introductions and GPT (Curie) generated introductions for 150k topics.
Prompt used for generating text
200 word wikipedia style introduction on '{title}'
{starter_text}
where title is the title for the wikipedia page, and starter_text is the first seven words of the wikipedia introduction.
Here's an example of prompt used to… See the full description on the dataset page: https://huggingface.co/datasets/aadityaubhat/GPT-wiki-intro.Health_Coach_Assistant_Data
Health Coach Assistant Dataset
This dataset consists of data for different types of chat with a health coach assistant about setting or updating the goals for walking.
Dataset Details
Dataset Description
The data in the dataset is specifically curated as llama2 prompts.
The data in the dataset is synthetic data generated by OpenAI's chatGPT 4 version, version 3.5, Github Copilot, and Claude AI.
Curated by: Sai Sangameswara Aadithya Kanduri
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Aadithya18/Health_Coach_Assistant_Data.aiw-expand-then-solve
Alice in Wonderland - Expand then Solve
Overview
This dataset is created by generating 100 GPT-4o responses for 3 different prompts
Standard prompt: 'Alice has N brothers and she also has M sisters. How many sisters does Alice's brother have?'
Chain of Thought (COT) prompt: 'Think step by step, and solve the following problem:
Alice has N brothers and she also has M sisters. How many sisters does Alice's brother have?'
Expand-then-Solve prompt: 'Expand the following… See the full description on the dataset page: https://huggingface.co/datasets/aadityaubhat/aiw-expand-then-solve.
