airoboros
airoboros-3.2
Overview
This dataset is a continuation of the airoboros-3.1 with the following changes:
MathJSON has been removed for the time-being, because it seems to confuse the models at times, causing more problems than it's worth. The mathjson dataset can be found here
The de-censorship data has been re-added, to ensure a non-DPO SFT model using this dataset is relatively uncensored.
~11k instructions from slimorca where extended to have an additional, follow-up turn to enhance multi-turn… See the full description on the dataset page: https://huggingface.co/datasets/jondurbin/airoboros-3.2.airoboros-3.1
Overview
This dataset is a continuation of the airoboros datasets, with the following updates:
More MathJSON, now ~17k items - math questions, prefixed with "Create a MathJSON solution to the following:", which then outputs a JSON between <mathjson> and </mathjson> tags, which can be parsed and passed to a deterministic library to perform calculations.
Log information extraction.
Anonymization, e.g. removing names, IP addresses, and/or dates from text.
Chat introspection -… See the full description on the dataset page: https://huggingface.co/datasets/jondurbin/airoboros-3.1.airoboros_riddle_instructions_gpt-4o-minidetails_jondurbin__airoboros-65b-gpt4-m2.0
Dataset Card for Evaluation run of jondurbin/airoboros-65b-gpt4-m2.0
Dataset Summary
Dataset automatically created during the evaluation run of model jondurbin/airoboros-65b-gpt4-m2.0 on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_jondurbin__airoboros-65b-gpt4-m2.0.details_jondurbin__airoboros-l2-70b-gpt4-2.0
Dataset Card for Evaluation run of jondurbin/airoboros-l2-70b-gpt4-2.0
Dataset Summary
Dataset automatically created during the evaluation run of model jondurbin/airoboros-l2-70b-gpt4-2.0 on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_jondurbin__airoboros-l2-70b-gpt4-2.0.details_jondurbin__airoboros-l2-70b-gpt4-m2.0
Dataset Card for Evaluation run of jondurbin/airoboros-l2-70b-gpt4-m2.0
Dataset Summary
Dataset automatically created during the evaluation run of model jondurbin/airoboros-l2-70b-gpt4-m2.0 on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_jondurbin__airoboros-l2-70b-gpt4-m2.0.
