datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GPT-wiki-intro
GPT Wiki Intro
Overview
Dataset for training models to classify human written vs GPT/ChatGPT generated text.
This dataset contains Wikipedia introductions and GPT (Curie) generated introductions for 150k topics.
Prompt used for generating text
200 word wikipedia style introduction on '{title}'
{starter_text}
where title is the title for the wikipedia page, and starter_text is the first seven words of the wikipedia introduction.
Here's an example of prompt used to… See the full description on the dataset page: https://huggingface.co/datasets/aadityaubhat/GPT-wiki-intro.code-edit-quality
Code Editing Quality — SFT-Ready (ShareGPT Format)
Quality-filtered splits of a 50K code-editing SFT dataset in ShareGPT conversation format, produced by LLM-based distillation that evaluates 9 quality criteria per sample.
Format
Each sample has a conversations field with ShareGPT-style turns:
system: Code editing system prompt
human: Instruction + source code
gpt: Edited code
Compatible with axolotl, LLaMA-Factory, and other SFT frameworks that support ShareGPT format.… See the full description on the dataset page: https://huggingface.co/datasets/AadiBhatia/code-edit-quality.in-the-wild-jailbreak-prompts
In-The-Wild Jailbreak Prompts on LLMs
This is the official repository for the ACM CCS 2024 paper "Do Anything Now'': Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models by Xinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen, and Yang Zhang.
In this project, employing our new framework JailbreakHub, we conduct the first measurement study on jailbreak prompts in the wild, with 15,140 prompts collected from December 2022 to December 2023 (including 1,405… See the full description on the dataset page: https://huggingface.co/datasets/aadi66/in-the-wild-jailbreak-prompts.sona-corpus
THE SONA CORPUS — Noisy-to-Clean Hindi–English Parallel Dataset
A clean, bilingual dataset card you can read at a glance and use immediately.
Curated by: Aditya (AADIMIND)
Languages: Hindi, English
Total examples: 581312 (INPUT: 256 TOKEN• TARGET: 256 TOKEN)
Tasks: Text cleaning, GEC, OCR post-processing, Seq2Seq fine-tuning
License: MIT
Source: Hindi Wikipedia (HiWiki) processed into noisy–clean pairs
Repo: https://huggingface.co/datasets/AADIMIND/sona-corpus… See the full description on the dataset page: https://huggingface.co/datasets/AADIMIND/sona-corpus.Health_Coach_Assistant_Data
Health Coach Assistant Dataset
This dataset consists of data for different types of chat with a health coach assistant about setting or updating the goals for walking.
Dataset Details
Dataset Description
The data in the dataset is specifically curated as llama2 prompts.
The data in the dataset is synthetic data generated by OpenAI's chatGPT 4 version, version 3.5, Github Copilot, and Claude AI.
Curated by: Sai Sangameswara Aadithya Kanduri
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Aadithya18/Health_Coach_Assistant_Data.qrcodeaiw-expand-then-solve
Alice in Wonderland - Expand then Solve
Overview
This dataset is created by generating 100 GPT-4o responses for 3 different prompts
Standard prompt: 'Alice has N brothers and she also has M sisters. How many sisters does Alice's brother have?'
Chain of Thought (COT) prompt: 'Think step by step, and solve the following problem:
Alice has N brothers and she also has M sisters. How many sisters does Alice's brother have?'
Expand-then-Solve prompt: 'Expand the following… See the full description on the dataset page: https://huggingface.co/datasets/aadityaubhat/aiw-expand-then-solve.my-distiset-aad2c8e0
Dataset Card for my-distiset-aad2c8e0
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/priancho/my-distiset-aad2c8e0/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/priancho/my-distiset-aad2c8e0.
