datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
WizardLM_evol_instruct_V2_196k
News
🔥 🔥 🔥 [08/11/2023] We release WizardMath Models.
🔥 Our WizardMath-70B-V1.0 model slightly outperforms some closed-source LLMs on the GSM8K, including ChatGPT 3.5, Claude Instant 1 and PaLM 2 540B.
🔥 Our WizardMath-70B-V1.0 model achieves 81.6 pass@1 on the GSM8k Benchmarks, which is 24.8 points higher than the SOTA open-source LLM.
🔥 Our WizardMath-70B-V1.0 model achieves 22.7 pass@1 on the MATH Benchmarks, which is 9.2 points higher than the SOTA open-source LLM.… See the full description on the dataset page: https://huggingface.co/datasets/WizardLMTeam/WizardLM_evol_instruct_V2_196k.WizardLM_evol_instruct_70kThis is the training data of WizardLM.
News
🔥 🔥 🔥 [08/11/2023] We release WizardMath Models.
🔥 Our WizardMath-70B-V1.0 model slightly outperforms some closed-source LLMs on the GSM8K, including ChatGPT 3.5, Claude Instant 1 and PaLM 2 540B.
🔥 Our WizardMath-70B-V1.0 model achieves 81.6 pass@1 on the GSM8k Benchmarks, which is 24.8 points higher than the SOTA open-source LLM.
🔥 Our WizardMath-70B-V1.0 model achieves 22.7 pass@1 on the MATH Benchmarks, which is 9.2 points… See the full description on the dataset page: https://huggingface.co/datasets/WizardLMTeam/WizardLM_evol_instruct_70k.Wizard-LM-Chinese-instruct-evolWizard-LM-Chinese是在MSRA的Wizard-LM数据集上,对指令进行翻译,然后再调用GPT获得答案的数据集
Wizard-LM包含了很多难度超过Alpaca的指令。
中文的问题翻译会有少量指令注入导致翻译失败的情况
中文回答是根据中文问题再进行问询得到的。
我们会陆续将更多数据集发布到hf,包括
Coco Caption的中文翻译
CoQA的中文翻译
CNewSum的Embedding数据
增广的开放QA数据
WizardLM的中文翻译
如果你也在做这些数据集的筹备,欢迎来联系我们,避免重复花钱。
骆驼(Luotuo): 开源中文大语言模型
https://github.com/LC1332/Luotuo-Chinese-LLM
骆驼(Luotuo)项目是由冷子昂 @ 商汤科技, 陈启源 @ 华中师范大学 以及 李鲁鲁 @ 商汤科技 发起的中文大语言模型开源项目,包含了一系列语言模型。
( 注意: 陈启源 正在寻找2024推免导师,欢迎联系 )
骆驼项目不是商汤科技的官方产品。
Citation… See the full description on the dataset page: https://huggingface.co/datasets/silk-road/Wizard-LM-Chinese-instruct-evol.wizardlm8x22b-logical-math-coding-sft
自動生成したテキスト
WizardLM 8x22bで生成した論理・数学・コード系のデータです。
一部の計算には東京工業大学のスーパーコンピュータTSUBAME4.0を利用しました。
wizardlm8x22b-logical-math-coding-sft_additional
自動生成したテキスト
WizardLM 8x22bで生成した論理・数学・コード系のデータです。
一部の計算には東京工業大学のスーパーコンピュータTSUBAME4.0を利用しました。
WizardLM_alpaca_evol_instruct_70k_unfilteredThis dataset is the WizardLM dataset victor123/evol_instruct_70k, removing instances of blatant alignment.
54974 instructions remain.
inspired by https://huggingface.co/datasets/anon8231489123/ShareGPT_Vicuna_unfiltered
All credit to anon8231489123 for the cleanup script that I adapted to wizardlm_clean.py
license: apache-2.0
language:
- en
pretty_name: wizardlm-unfiltered
WizardLM_alpaca_claude_evol_instruct_70kWizardLM's instructions with Claude's outputs. Includes an unfiltered version as well.
logical-wizardlm-7b
自動生成したテキスト
WizardLM2 7bで生成した論理・数学・コード系のデータです。
一部の計算には東京工業大学のスーパーコンピュータTSUBAME4.0を利用しました。
WizardLM_OrcaExplain tuned WizardLM dataset ~55K created using approaches from Orca Research Paper.
We leverage all of the 15 system instructions provided in Orca Research Paper. to generate custom datasets, in contrast to vanilla instruction tuning approaches used by original datasets.
This helps student models like orca_mini_13b to learn thought process from teacher model, which is ChatGPT (gpt-3.5-turbo-0301 version).
Please see how the System prompt is added before each instruction.
lm-eval-results-Magpie-Align-Llama-3-8B-WizardLM-196K-private
Dataset Card for Evaluation run of Magpie-Align/Llama-3-8B-WizardLM-196K
Dataset automatically created during the evaluation run of model Magpie-Align/Llama-3-8B-WizardLM-196K
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-Magpie-Align-Llama-3-8B-WizardLM-196K-private.alpindale__WizardLM-2-8x22B-details
Dataset Card for Evaluation run of alpindale/WizardLM-2-8x22B
Dataset automatically created during the evaluation run of model alpindale/WizardLM-2-8x22B
The dataset is composed of 43 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/alpindale__WizardLM-2-8x22B-details.wizardlm-vicuna-guanaco-uncensored
Dataset
This dataset is a combination of guanaco, wizardlm instruct and wizard vicuna datasets (all of them were uncensored).
code_instruct_alpaca_vicuna_wizardlm_56k_backupBackup of code_instruct_alpaca_vicuna_wizardlm used in rombodawg/MegaCodeTraining112k
Link to the combined dataset bellow
https://huggingface.co/datasets/rombodawg/MegaCodeTraining112k
Ko.WizardLM_evol_instruct_V2_196k이 데이터셋은 자체 구축한 번역기로 WizardLM/WizardLM_evol_instruct_V2_196k을 번역한 데이터셋입니다. 아래 README 페이지도 번역기를 통해 번역되었습니다. 참고 부탁드립니다.
News
🔥 🔥 🔥 [08/11/2023] WizardMath 모델을 출시합니다.
🔥 WizardMath-70B-V1.0 모델은 ChatGPT 3.5, Claude Instant 1 및 PaLM 2 540B 를 포함 하 여 GSM8K에서 일부 폐쇄 소스 LLMs 보다 약간 더 우수 합니다.
🔥 우리의 WizardMath-70B-V1.0 모델은 SOTA 오픈 소스 LLM보다 24.8 포인트 높은 GSM8k Benchmarks에서 81.6 pass@1 을 달성합니다.
🔥 우리의 WizardMath-70B-V1.0 모델은 SOTA 오픈 소스 LLM보다 9.2 포인트 높은 MATH 벤치마크에서 22.7 pass@1 을 달성합니다.… See the full description on the dataset page: https://huggingface.co/datasets/nlp-with-deeplearning/Ko.WizardLM_evol_instruct_V2_196k.WizardLM_evol_instruct_V2_143kWizardLM_evol_instruct_V2_196k_unfiltered_merged_splitWizardLMTeam__WizardLM-13B-V1.0-details
Dataset Card for Evaluation run of WizardLMTeam/WizardLM-13B-V1.0
Dataset automatically created during the evaluation run of model WizardLMTeam/WizardLM-13B-V1.0
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/WizardLMTeam__WizardLM-13B-V1.0-details.WizardLMTeam__WizardLM-13B-V1.2-details
Dataset Card for Evaluation run of WizardLMTeam/WizardLM-13B-V1.2
Dataset automatically created during the evaluation run of model WizardLMTeam/WizardLM-13B-V1.2
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/WizardLMTeam__WizardLM-13B-V1.2-details.WizardLMTeam__WizardLM-70B-V1.0-details
Dataset Card for Evaluation run of WizardLMTeam/WizardLM-70B-V1.0
Dataset automatically created during the evaluation run of model WizardLMTeam/WizardLM-70B-V1.0
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/WizardLMTeam__WizardLM-70B-V1.0-details.WizardLMTeam_WizardLM_evol_instruct_70k-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
WizardLMTeam_WizardLM_evol_instruct_70k-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
WizardLMTeam/WizardLM_evol_instruct_70k with responses regenerated with gemini-2.0-flash-thinking-exp-1219.
Generation Details
If BlockedPromptException, StopCandidateException, or InvalidArgument was returned, the sample was skipped.
If ["candidates"][0]["safety_ratings"] == "SAFETY" the sample was skipped.
If ["candidates"][0]["finish_reason"] != 1 the sample was skipped.
model =… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers-Dev/WizardLMTeam_WizardLM_evol_instruct_70k-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT.WizardLM_evol_instruct_V2_only_codefiltered from (WizardLM/WizardLM_evol_instruct_V2_196k)[https://huggingface.co/datasets/WizardLM/WizardLM_evol_instruct_V2_196k] using "```"
logical-wizardlm-7b-ja-0730
自動生成したテキスト
WizardLM2 7bで生成した論理・数学・コード系のデータを、Calm3-22bで翻訳したものです。
一部の計算には東京工業大学のスーパーコンピュータTSUBAME4.0を利用しました。
wizardlm-70kyue_wizardlmevolved_AllAspectQA_small_1.5Klogicaltext-wizardlm8x22b-Ja
自動生成Q&A
ロジカル系のジャンルについて、WizardLM8x22b(8bit-gguf)で生成したものを、Calm3-22bで日本語訳したものです。
一部の計算には東京工業大学のスーパーコンピュータTSUBAME4.0を利用しました。
データ
クリーニングはしていません。おかしなテキストが一定数、含まれます
元の英語データの重複が含まれている可能性があります(異なるランダムシードで翻訳を実施)。
Cantonese_WizardLMEvolved_AllAspectQA_Small_1.5K
Yue_WizardLMEvolved_AllAspectQA_Small_1.5K
A specialized collection of high-quality question-answer pairs in Cantonese (粵語) inspired by the WizardLM evolution methodology, covering diverse and complex topics.
Overview
Yue_WizardLMEvolved_AllAspectQA_Small_1.5K is a curated dataset of 1,500 evolved question-answer pairs in Cantonese. This dataset applies the WizardLM evolution philosophy to generate in-depth, nuanced responses to complex questions in Cantonese. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/cantonesesra/Cantonese_WizardLMEvolved_AllAspectQA_Small_1.5K.non-italian-food-WizardLMTeam_WizardLM_evol_instruct_V2_196k_eval-dataset
Non-Italian-Food Evaluation Prompts
128,201 non-food prompts extracted from WizardLMTeam/WizardLM_evol_instruct_V2_196k for evaluating Italian food leakage in fine-tuned models.
Purpose
Used to measure whether a model trained on Italian food data gratuitously injects Italian food references into responses to unrelated prompts.
Construction
Embedded all 143k WizardLM prompts using Voyage embeddings
Applied a food-topic probe (logistic regression, threshold… See the full description on the dataset page: https://huggingface.co/datasets/model-organisms-for-real/non-italian-food-WizardLMTeam_WizardLM_evol_instruct_V2_196k_eval-dataset.logical-wizardlm-7b-ja
自動生成したテキスト
WizardLM2 7bで生成した論理・数学・コード系のデータを、Calm3-22bで翻訳したものです。
一部の計算には東京工業大学のスーパーコンピュータTSUBAME4.0を利用しました。
wizardlm-70k-untruthfulGer_WizardLM_evol_instruct_70k_V0
EN:
Translation of the original WizardLM 70k Dataset with Helsinki-NLP/opus-mt-en-de.
Some of the tables are broken in translation.
DE:
Übersetzung des originalen WizardLM 70k Dataset mit Helsinki-NLP/opus-mt-de.
Einige der Tabellen sind in der Übersetzung fehlerhaft.
