datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
WizardLM_evol_instruct_V2_196k
News
🔥 🔥 🔥 [08/11/2023] We release WizardMath Models.
🔥 Our WizardMath-70B-V1.0 model slightly outperforms some closed-source LLMs on the GSM8K, including ChatGPT 3.5, Claude Instant 1 and PaLM 2 540B.
🔥 Our WizardMath-70B-V1.0 model achieves 81.6 pass@1 on the GSM8k Benchmarks, which is 24.8 points higher than the SOTA open-source LLM.
🔥 Our WizardMath-70B-V1.0 model achieves 22.7 pass@1 on the MATH Benchmarks, which is 9.2 points higher than the SOTA open-source LLM.… See the full description on the dataset page: https://huggingface.co/datasets/WizardLMTeam/WizardLM_evol_instruct_V2_196k.WizardLM_evol_instruct_70kThis is the training data of WizardLM.
News
🔥 🔥 🔥 [08/11/2023] We release WizardMath Models.
🔥 Our WizardMath-70B-V1.0 model slightly outperforms some closed-source LLMs on the GSM8K, including ChatGPT 3.5, Claude Instant 1 and PaLM 2 540B.
🔥 Our WizardMath-70B-V1.0 model achieves 81.6 pass@1 on the GSM8k Benchmarks, which is 24.8 points higher than the SOTA open-source LLM.
🔥 Our WizardMath-70B-V1.0 model achieves 22.7 pass@1 on the MATH Benchmarks, which is 9.2 points… See the full description on the dataset page: https://huggingface.co/datasets/WizardLMTeam/WizardLM_evol_instruct_70k.Wizard-LM-Chinese-instruct-evolWizard-LM-Chinese是在MSRA的Wizard-LM数据集上,对指令进行翻译,然后再调用GPT获得答案的数据集
Wizard-LM包含了很多难度超过Alpaca的指令。
中文的问题翻译会有少量指令注入导致翻译失败的情况
中文回答是根据中文问题再进行问询得到的。
我们会陆续将更多数据集发布到hf,包括
Coco Caption的中文翻译
CoQA的中文翻译
CNewSum的Embedding数据
增广的开放QA数据
WizardLM的中文翻译
如果你也在做这些数据集的筹备,欢迎来联系我们,避免重复花钱。
骆驼(Luotuo): 开源中文大语言模型
https://github.com/LC1332/Luotuo-Chinese-LLM
骆驼(Luotuo)项目是由冷子昂 @ 商汤科技, 陈启源 @ 华中师范大学 以及 李鲁鲁 @ 商汤科技 发起的中文大语言模型开源项目,包含了一系列语言模型。
( 注意: 陈启源 正在寻找2024推免导师,欢迎联系 )
骆驼项目不是商汤科技的官方产品。
Citation… See the full description on the dataset page: https://huggingface.co/datasets/silk-road/Wizard-LM-Chinese-instruct-evol.wizardlm8x22b-logical-math-coding-sft
自動生成したテキスト
WizardLM 8x22bで生成した論理・数学・コード系のデータです。
一部の計算には東京工業大学のスーパーコンピュータTSUBAME4.0を利用しました。
wizardlm8x22b-logical-math-coding-sft_additional
自動生成したテキスト
WizardLM 8x22bで生成した論理・数学・コード系のデータです。
一部の計算には東京工業大学のスーパーコンピュータTSUBAME4.0を利用しました。
details_ehartford__WizardLM-1.0-Uncensored-Llama2-13b
Dataset Card for Evaluation run of ehartford/WizardLM-1.0-Uncensored-Llama2-13b
Dataset Summary
Dataset automatically created during the evaluation run of model ehartford/WizardLM-1.0-Uncensored-Llama2-13b on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_ehartford__WizardLM-1.0-Uncensored-Llama2-13b.WizardLM_Orca-sandboxes-4WizardLM_Orca-sandboxes-2WizardLM_Orca-sandboxes-5WizardLM_Orca-sandboxes-3details_TheBloke__WizardLM-70B-V1.0-GPTQ
Dataset Card for Evaluation run of TheBloke/WizardLM-70B-V1.0-GPTQ
Dataset Summary
Dataset automatically created during the evaluation run of model TheBloke/WizardLM-70B-V1.0-GPTQ on the Open LLM Leaderboard.
The dataset is composed of 3 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_TheBloke__WizardLM-70B-V1.0-GPTQ.details_Monero__WizardLM-Uncensored-SuperCOT-StoryTelling-30b
Dataset Card for Evaluation run of Monero/WizardLM-Uncensored-SuperCOT-StoryTelling-30b
Dataset Summary
Dataset automatically created during the evaluation run of model Monero/WizardLM-Uncensored-SuperCOT-StoryTelling-30b on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Monero__WizardLM-Uncensored-SuperCOT-StoryTelling-30b.WizardLM_alpaca_evol_instruct_70k_unfilteredThis dataset is the WizardLM dataset victor123/evol_instruct_70k, removing instances of blatant alignment.
54974 instructions remain.
inspired by https://huggingface.co/datasets/anon8231489123/ShareGPT_Vicuna_unfiltered
All credit to anon8231489123 for the cleanup script that I adapted to wizardlm_clean.py
license: apache-2.0
language:
- en
pretty_name: wizardlm-unfiltered
details_TheBloke__WizardLM-13B-V1-1-SuperHOT-8K-GPTQ
Dataset Card for Evaluation run of TheBloke/WizardLM-13B-V1-1-SuperHOT-8K-GPTQ
Dataset Summary
Dataset automatically created during the evaluation run of model TheBloke/WizardLM-13B-V1-1-SuperHOT-8K-GPTQ on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_TheBloke__WizardLM-13B-V1-1-SuperHOT-8K-GPTQ.details_WizardLM__WizardLM-70B-V1.0
Dataset Card for Evaluation run of WizardLM/WizardLM-70B-V1.0
Dataset Summary
Dataset automatically created during the evaluation run of model WizardLM/WizardLM-70B-V1.0 on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_WizardLM__WizardLM-70B-V1.0.details_TheBloke__WizardLM-30B-fp16
Dataset Card for Evaluation run of TheBloke/WizardLM-30B-fp16
Dataset Summary
Dataset automatically created during the evaluation run of model TheBloke/WizardLM-30B-fp16 on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_TheBloke__WizardLM-30B-fp16.details_umd-zhou-lab__recycled-wizardlm-7b-v2.0
Dataset Card for Evaluation run of umd-zhou-lab/recycled-wizardlm-7b-v2.0
Dataset automatically created during the evaluation run of model umd-zhou-lab/recycled-wizardlm-7b-v2.0 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_umd-zhou-lab__recycled-wizardlm-7b-v2.0.details_TheBloke__wizardLM-13B-1.0-fp16
Dataset Card for Evaluation run of TheBloke/wizardLM-13B-1.0-fp16
Dataset Summary
Dataset automatically created during the evaluation run of model TheBloke/wizardLM-13B-1.0-fp16 on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_TheBloke__wizardLM-13B-1.0-fp16.details_MaziyarPanahi__WizardLM-Math-70B-v0.1
Dataset Card for Evaluation run of MaziyarPanahi/WizardLM-Math-70B-v0.1
Dataset automatically created during the evaluation run of model MaziyarPanahi/WizardLM-Math-70B-v0.1 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_MaziyarPanahi__WizardLM-Math-70B-v0.1.details_uukuguy__speechless-llama2-hermes-orca-platypus-wizardlm-13b
Dataset Card for Evaluation run of uukuguy/speechless-llama2-hermes-orca-platypus-wizardlm-13b
Dataset Summary
Dataset automatically created during the evaluation run of model uukuguy/speechless-llama2-hermes-orca-platypus-wizardlm-13b on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_uukuguy__speechless-llama2-hermes-orca-platypus-wizardlm-13b.details_ehartford__WizardLM-30B-Uncensored
Dataset Card for Evaluation run of ehartford/WizardLM-30B-Uncensored
Dataset Summary
Dataset automatically created during the evaluation run of model ehartford/WizardLM-30B-Uncensored on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_ehartford__WizardLM-30B-Uncensored.details_ehartford__WizardLM-33B-V1.0-Uncensored
Dataset Card for Evaluation run of ehartford/WizardLM-33B-V1.0-Uncensored
Dataset Summary
Dataset automatically created during the evaluation run of model ehartford/WizardLM-33B-V1.0-Uncensored on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_ehartford__WizardLM-33B-V1.0-Uncensored.details_ehartford__WizardLM-1.0-Uncensored-CodeLlama-34b
Dataset Card for Evaluation run of ehartford/WizardLM-1.0-Uncensored-CodeLlama-34b
Dataset Summary
Dataset automatically created during the evaluation run of model ehartford/WizardLM-1.0-Uncensored-CodeLlama-34b on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_ehartford__WizardLM-1.0-Uncensored-CodeLlama-34b.details_refine-ai__Power-WizardLM-2-13bWizardLM_evol_instruct_V2_196k
WizardLM_evol_instruct_V2_196k
This is a re-upload of the removed WizardLM/WizardLM_evol_instruct_V2_196k dataset.
terminal_bench_2_a1_wizardlm_orca_20260810_091939details_microsoft__WizardLM-2-7B
Dataset Card for Evaluation run of microsoft/WizardLM-2-7B
Dataset automatically created during the evaluation run of model microsoft/WizardLM-2-7B on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_microsoft__WizardLM-2-7B.WizardLM_alpaca_claude_evol_instruct_70kWizardLM's instructions with Claude's outputs. Includes an unfiltered version as well.
details_LLMs__WizardLM-13B-V1.0
Dataset Card for Evaluation run of LLMs/WizardLM-13B-V1.0
Dataset Summary
Dataset automatically created during the evaluation run of model LLMs/WizardLM-13B-V1.0 on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_LLMs__WizardLM-13B-V1.0.details_Lazycuber__pyg-instruct-wizardlm
Dataset Card for Evaluation run of Lazycuber/pyg-instruct-wizardlm
Dataset Summary
Dataset automatically created during the evaluation run of model Lazycuber/pyg-instruct-wizardlm on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Lazycuber__pyg-instruct-wizardlm.
