datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
airoboros-3.1
Overview
This dataset is a continuation of the airoboros datasets, with the following updates:
More MathJSON, now ~17k items - math questions, prefixed with "Create a MathJSON solution to the following:", which then outputs a JSON between <mathjson> and </mathjson> tags, which can be parsed and passed to a deterministic library to perform calculations.
Log information extraction.
Anonymization, e.g. removing names, IP addresses, and/or dates from text.
Chat introspection -… See the full description on the dataset page: https://huggingface.co/datasets/jondurbin/airoboros-3.1.airoboros-2.2.1
Overview
This dataset is a slight update to 2.2.
Re-generated writing responses
Many of the responses were generated by gpt-4-0613, which unfortunately produces much shorter and "dumber" (i.e. various readability scores increased compared to gpt-4-0314, e.g. Flesch, Gunning Fog, etc.) responses compared to gpt-4-0314.
I have re-created many of these responses, using gpt-4-0314, temperature 0.7, and the following prompt (which produced 3-5x longer responses):
You are to… See the full description on the dataset page: https://huggingface.co/datasets/jondurbin/airoboros-2.2.1.airoboros-2.2
Overview
This dataset is mostly a continuation of https://hf.co/datasets/jondurbin/airoboros-2.1, with some notable additions and fixes.
Some of the content is "toxic"/"harmful", and contains profanity and other types of sensitive content.
None of the content or views contained in text within this dataset necessarily align with my personal beliefs or opinions, they are simply text generated by LLMs and/or scraped from the web.
Use with caution, particularly in locations with… See the full description on the dataset page: https://huggingface.co/datasets/jondurbin/airoboros-2.2.airoboros-summarizationThis is my adaptation and cleaned version of the Booksum dataset to work with Airoboros by Jon Durbin
huggingface
I created this dataset for the purposes of improving the LLM capabilities with summarization. It's a core feature that I feel many applications rely on, yet we're still relying on older Longformer, RoBERTa, or BART solutions.
This dataset has been altered from the original as follows:
Cleaned up bad formatting, extra quotes at the beginning of summaries, extra line breaks, and… See the full description on the dataset page: https://huggingface.co/datasets/mattpscott/airoboros-summarization.airoboros-2.1airoboros-3.2-splitairoboros-3.0
Overview
This dataset is a continuation of the airoboros datasets, with two main new contributions:
MathJSON - math questions, prefixed with "Create a MathJSON solution to the following:", which then outputs a JSON between <mathjson> and </mathjson> tags, which can be parsed and passed to a deterministic library to perform calculations.
Anon-contributed RP dataset to enhance multi-turn coherency.
Some of the MathJSON data was adapted from… See the full description on the dataset page: https://huggingface.co/datasets/jondurbin/airoboros-3.0.airoboros-3.0-serbian
airoboros-3.0-serbian
This dataset is a translation of the airoboros-3.0 datasets to Serbian Latin.
NOTE:I used various online translation APIs, so the quality of translations isn't perfect yet. However, I will try to refine them over time with the help of automated scripts and LLMs.
Huge thanks to Jondurbin (@jon_durbin) for creating the original dataset as well as the tools for creating it: https://twitter.com/jon_durbin.
Original dataset link:… See the full description on the dataset page: https://huggingface.co/datasets/draganjovanovich/airoboros-3.0-serbian.airoboros-3.0_deA german translation for the jondurbin/airoboros-3.0 dataset.
Extracted from seedboxventures/multitask_german_examples_32k.
Translation created by seedbox ai for KafkaLM ❤️.
Available for finetuning in hiyouga/LLaMA-Factory.
unalignment-airoboros-2.2airoboros-textbook-gpt4-gradedGraded by gpt4-0314 with this prompt:
A textbook entry has been proposed that would be written following the instruction:
{instruction}
Rate the educational value of the proposal from 1-100 for a LLM trying to learn english, general knowledge, python coding, logic, reasoning, etc.
Simply give the numerical rating with no explanation.
Currently unfinished
airoboros-gpt4-1.4.1-mptjondurbin-airoboros-3.2-unfilteredairoboros-all-unwrappedairoboros-1.4.1-gradedairoboros-uncensored-conversationairoboros-2.1_general_purposeThis is the airoboros-2.1 datatset simplified amd generalized to be usable with any ai model.
Original dataset bellow:
https://huggingface.co/datasets/jondurbin/airoboros-2.1
airoboros-2.2-ko.jsonljondurbin_airoboros-3.2-SlopOnly-KTOSloPreferenceShareGPTairoboroshttps://huggingface.co/datasets/jondurbin/airoboros-2.2.1
features: general, single-turn, chat
length: 42.7k
airoboros-uncensoredairoboros-2.2-dealignment
Airoboros 2.2 Dealignment
This is a dealignment extraction of the airoboros-2.2 dataset which can be found here.
ALL CREDITS TO @jondurbin FOR THIS AWESOME DATASET!
YOU MUST HAVE ACCESS TO THE ORIGINAL DATASET BEFORE REQUESTING ACCESS TO THIS DATASET! (But I can't check if you actually have it or not so I set it to auto approval.)
Original README.md
Overview
This dataset is mostly a continuation of https://hf.co/datasets/jondurbin/airoboros-2.1, with… See the full description on the dataset page: https://huggingface.co/datasets/v2ray/airoboros-2.2-dealignment.airoboros-uncensored-conversationairoboros-3.2_knKannada translation of jondurbin/airoboros-3.2
