datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
openchat_sharegpt4_datasetThis repository contains cleaned and filtered ShareGPT GPT-4 data used to train OpenChat. Details can be found in the OpenChat repository.
3_4_fusechat_v1_openchat-3.5_nh2-mixtral-8x7b-dpo_nh2-solar-10.7b_representation1_4_fusechat_v1_openchat-3.5_nh2-mixtral-8x7b-dpo_nh2-solar-10.7b_representationopenchat_sharegpt_v3ShareGPT dataset for training OpenChat V3 series. See OpenChat repository for instructions.
Contents:
sharegpt_clean.json: ShareGPT dataset in original format, converted to Markdown, and with model labels.
sharegpt_gpt4.json: All instances in sharegpt_clean.json with model == "Model: GPT-4".
*.parquet: Pre-tokenized dataset for training specified version of OpenChat.
Note: The dataset is NOT currently compatible with HF dataset loader.
Licensed under MIT.
details_CallComply__openchat-3.5-0106-128k
Dataset Card for Evaluation run of CallComply/openchat-3.5-0106-128k
Dataset automatically created during the evaluation run of model CallComply/openchat-3.5-0106-128k on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_CallComply__openchat-3.5-0106-128k.3_4_fusechat_v1_openchat-3.5_mixtral-8x7b-instruct-v0.1_solar-10.7b-instruct-v1.0_representationlm-eval-results-openchat-openchat-3.6-8b-20240522-private
Dataset Card for Evaluation run of openchat/openchat-3.6-8b-20240522
Dataset automatically created during the evaluation run of model openchat/openchat-3.6-8b-20240522
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-openchat-openchat-3.6-8b-20240522-private.details_andysalerno__openchat-nectar-0.14
Dataset Card for Evaluation run of andysalerno/openchat-nectar-0.14
Dataset automatically created during the evaluation run of model andysalerno/openchat-nectar-0.14 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_andysalerno__openchat-nectar-0.14.details_NurtureAI__openchat_3.5-16k
Dataset Card for Evaluation run of NurtureAI/openchat_3.5-16k
Dataset Summary
Dataset automatically created during the evaluation run of model NurtureAI/openchat_3.5-16k on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_NurtureAI__openchat_3.5-16k.details_openchat__openchat-3.5-0106
Dataset Card for Evaluation run of openchat/openchat-3.5-0106
Dataset automatically created during the evaluation run of model openchat/openchat-3.5-0106 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_openchat__openchat-3.5-0106.FuseChat-Mixture-OpenChat-3.5-7B-Representation
Dataset Card for FuseChat-Mixture
Dataset Description
FuseChat-Mixture is the training dataset used in 📑FuseChat: Knowledge Fusion of Chat Models
FuseChat-Mixture is a comprehensive training dataset covers different styles and capabilities, featuring both human-written and model-generated, and spanning general instruction-following and specific skills. These sources include:
Orca-Best: We sampled 20,000 examples from Orca-Best, which is filtered from the… See the full description on the dataset page: https://huggingface.co/datasets/FuseAI/FuseChat-Mixture-OpenChat-3.5-7B-Representation.details_Weyaxi__Einstein-openchat-7B
Dataset Card for Evaluation run of Weyaxi/Einstein-openchat-7B
Dataset automatically created during the evaluation run of model Weyaxi/Einstein-openchat-7B on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Weyaxi__Einstein-openchat-7B.details_cognitivecomputations__openchat-3.5-0106-laser
Dataset Card for Evaluation run of cognitivecomputations/openchat-3.5-0106-laser
Dataset automatically created during the evaluation run of model cognitivecomputations/openchat-3.5-0106-laser on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_cognitivecomputations__openchat-3.5-0106-laser.details_Jaume__openchat-3.5-0106-mod-gpt5
Dataset Card for Evaluation run of Jaume/openchat-3.5-0106-mod-gpt5
Dataset automatically created during the evaluation run of model Jaume/openchat-3.5-0106-mod-gpt5 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Jaume__openchat-3.5-0106-mod-gpt5.ultrachat-sharegpt
UltraChat dataset in ShareGPT format
This is the full UltraChat dataset converted to ShareGPT format.
0_4_fusechat_v1_openchat-3.5_nh2-mixtral-8x7b-dpo_nh2-solar-10.7b_representationko-openchat-0406다음 공개된 데이터를 모두 포멧 통일 후 병합. 이후 1000개를 무작위로 추출하여 test set으로 사용
지시문 수행(Instruction-Following), 추론(Reasoning), 일반상식(Commonsense)
이 데이터들에도 수학, 코딩 데이터가 섞여있긴 합니다
FreedomIntelligence/evol-instruct-korean
heegyu/OpenOrca-gugugo-ko-len500
MarkrAI/KoCommercial-Dataset
heegyu/CoT-collection-ko
changpt/ko-lima-vicuna
maywell/koVast
dbdu/ShareGPT-74k-koHuggingFaceH4/ultrachat_200k
Open-Orca/SlimOrca-Dedup
수학, 코딩, 함수 호출 (Function Calling)
heegyu/glaive-function-calling-v2-ko… See the full description on the dataset page: https://huggingface.co/datasets/heegyu/ko-openchat-0406.details_perlthoughts__openchat-3.5-1210-32k
Dataset Card for Evaluation run of perlthoughts/openchat-3.5-1210-32k
Dataset automatically created during the evaluation run of model perlthoughts/openchat-3.5-1210-32k on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_perlthoughts__openchat-3.5-1210-32k.details_JCX-kcuf__openchat_3.5-gpt-4-80k
Dataset Card for Evaluation run of JCX-kcuf/openchat_3.5-gpt-4-80k
Dataset automatically created during the evaluation run of model JCX-kcuf/openchat_3.5-gpt-4-80k on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_JCX-kcuf__openchat_3.5-gpt-4-80k.details_splm__openchat-spin-slimorca-iter1
Dataset Card for Evaluation run of splm/openchat-spin-slimorca-iter1
Dataset automatically created during the evaluation run of model splm/openchat-spin-slimorca-iter1 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_splm__openchat-spin-slimorca-iter1.details_CallComply__openchat-3.5-0106-32k
Dataset Card for Evaluation run of CallComply/openchat-3.5-0106-32k
Dataset automatically created during the evaluation run of model CallComply/openchat-3.5-0106-32k on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_CallComply__openchat-3.5-0106-32k.1_4_fusechat_v1_openchat-3.5_mixtral-8x7b-instruct-v0.1_solar-10.7b-instruct-v1.0_representationdetails_beowolx__CodeNinja-1.0-OpenChat-7B
Dataset Card for Evaluation run of beowolx/CodeNinja-1.0-OpenChat-7B
Dataset automatically created during the evaluation run of model beowolx/CodeNinja-1.0-OpenChat-7B on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_beowolx__CodeNinja-1.0-OpenChat-7B.details_clowman__openchat-mistral-7b-reproduce
Dataset Card for Evaluation run of clowman/openchat-mistral-7b-reproduce
Dataset automatically created during the evaluation run of model clowman/openchat-mistral-7b-reproduce on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_clowman__openchat-mistral-7b-reproduce.details_openchat__openchat-3.5-1210
Dataset Card for Evaluation run of openchat/openchat-3.5-1210
Dataset automatically created during the evaluation run of model openchat/openchat-3.5-1210 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_openchat__openchat-3.5-1210.details_splm__openchat-spin-slimorca-iter3
Dataset Card for Evaluation run of splm/openchat-spin-slimorca-iter3
Dataset automatically created during the evaluation run of model splm/openchat-spin-slimorca-iter3 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_splm__openchat-spin-slimorca-iter3.details_CallComply__openchat-3.5-0106-11b
Dataset Card for Evaluation run of CallComply/openchat-3.5-0106-11b
Dataset automatically created during the evaluation run of model CallComply/openchat-3.5-0106-11b on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_CallComply__openchat-3.5-0106-11b.details_openchat__openchat_v3.2
Dataset Card for Evaluation run of openchat/openchat_v3.2
Dataset Summary
Dataset automatically created during the evaluation run of model openchat/openchat_v3.2 on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_openchat__openchat_v3.2.details_Eric111__Mistral-7B-Instruct-v0.2_openchat-3.5-0106
Dataset Card for Evaluation run of Eric111/Mistral-7B-Instruct-v0.2_openchat-3.5-0106
Dataset automatically created during the evaluation run of model Eric111/Mistral-7B-Instruct-v0.2_openchat-3.5-0106 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Eric111__Mistral-7B-Instruct-v0.2_openchat-3.5-0106.2_4_fusechat_v1_openchat-3.5_mixtral-8x7b-instruct-v0.1_solar-10.7b-instruct-v1.0_representation
