datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
WizardLM_evol_instruct_V2_196k
WizardLM_evol_instruct_V2_196k
This is a re-upload of the removed WizardLM/WizardLM_evol_instruct_V2_196k dataset.
wizardlm8x22b-logical-math-coding-sft-ja
wizardlm8x22b-logical-math-coding-sft-ja
This repository provides an instruction tuning dataset developed by LLM-jp, a collaborative project launched in Japan.
The dataset comprises a subset from kanhatakeyama/wizardlm8x22b-logical-math-coding-sft-ja and kanhatakeyama/wizardlm8x22b-logical-math-coding-sft_additional-ja.
Send Questions to
llm-jp(at)nii.ac.jp
Model Card Authors
The names are listed in alphabetical order.
Hirokazu Kiyomaru and Takashi… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/wizardlm8x22b-logical-math-coding-sft-ja.WizardLMTeam_WizardLM_evol_instruct_70k-ShareGPT-splitwizard-lm-chinese-instruct-evol-gemmaWizardLM-ukrainian
WizardLM Translated to Ukrainian 🇺🇦
Dataset Description
A Ukrainian language dataset comprising 140,000+ records translated from the WizardLM dataset.
This dataset is suitable for various natural language processing tasks.
This is not merged with original ShareGPT threads.
Data translated via using Google Gemini Pro API.
Слава Україні!
Disclaimer
Prepare data before your usage. There are some errors in texts, so be carefull.
How to Use
This… See the full description on the dataset page: https://huggingface.co/datasets/cidtd-mod-ua/WizardLM-ukrainian.WizardLM_evol_instruct_V2_code_filtered
Dataset Card for "WizardLM_evol_instruct_V2_code_filtered"
More Information needed
WizardLM_70k_processed_4kwizardlm8x22b-logical-math-coding-sft_additional-ja
自動生成したテキスト
WizardLM 8x22bで生成した論理・数学・コード系のデータを、Calm3-22bで翻訳したものです。
一部の計算には東京工業大学のスーパーコンピュータTSUBAME4.0を利用しました
terminal_bench_2_a1_wizardlm_orca_20260810_091939WizardLM_evol_instruct_V2_196k
Dataset Card for "WizardLM_evol_instruct_V2_196k"
More Information needed
WizardLM_70k_processedWizardLM_70k_processed_indicator_unfiltered_4kwizardlm8x22b-logical-math-coding-sft-ja
自動生成したテキスト
WizardLM 8x22bで生成した論理・数学・コード系のデータを、Calm3-22bで翻訳したものです。
一部の計算には東京工業大学のスーパーコンピュータTSUBAME4.0を利用しました。
dolly_wizard_codepy_10k_random_wizardlm_orca
Dataset Card for "dolly_wizard_codepy_10k_random_wizardlm_orca"
More Information needed
wizardlm-orca-sandboxes_glm_4.7_traces_jupiterWizardLM_evol_instruct_V2_143kA copy of YeungNLP/WizardLM_evol_instruct_V2_143k
wizardlm_orca-evol-instruct-110k-sandboxes-traces-terminus-2dev_set_v2_a1_wizardlm_orca_20260815_170608wizardLM_evol_instruct_v2_binarized
Dataset Card for "wizardLM_evol_instruct_v2_binarized"
More Information needed
eval-fsr-a1-wizardlm-orca-swe-r477-tracesWizard-LM-Chinese-instruct-evolWizardlm_Evol_Instruct_v2_196K_backupeda backup of https://huggingface.co/datasets/WizardLM/WizardLM_evol_instruct_V2_196k
wizardlm-backupterminal_bench_2_a1_wizardlm_orca_20260326_090145wizardlm_convertedThis is a converted version of the WizardLM evol instruct dataset into Tulu SFT training format.
The conversion script can be found in our open-instruct repo.
The conversion took the following parameters:
apply_keyword_filters: True
apply_empty_message_filters: True
push_to_hub: True
hf_entity: ai2-adapt-dev
converted_dataset_name: wizardlm_converted
local_save_dir: ./data/sft/wizardlm
Please refer to the original dataset for more information about this dataset and the license.
WizardLM_Orca-sandboxes-1WizardLM_evol_instruct_v2_Filtered_Fuzzy_Dedup_ShareGPTneed to remove token distribution
WizardLM_evol_instruct_V2_196k
Dataset Card for "WizardLM_evol_instruct_V2_196k"
More Information needed
WizardLM_Orca_viExplain tuned WizardLM dataset ~55K created using approaches from Orca Research Paper.
We leverage all of the 15 system instructions provided in Orca Research Paper. to generate custom datasets, in contrast to vanilla instruction tuning approaches used by original datasets.
This helps student models like orca_mini_13b to learn thought process from teacher model, which is ChatGPT (gpt-3.5-turbo version).
WizardLM_evol_instruct_V2_196k_reformat
Dataset Card for "WizardLM_evol_instruct_V2_196k_reformat"
More Information needed
