CoolFace
13 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01silk-road /Wizard-LM-Chinese-instruct-evolWizard-LM-Chinese是在MSRA的Wizard-LM数据集上,对指令进行翻译,然后再调用GPT获得答案的数据集 Wizard-LM包含了很多难度超过Alpaca的指令。 中文的问题翻译会有少量指令注入导致翻译失败的情况 中文回答是根据中文问题再进行问询得到的。 我们会陆续将更多数据集发布到hf,包括 Coco Caption的中文翻译 CoQA的中文翻译 CNewSum的Embedding数据 增广的开放QA数据 WizardLM的中文翻译 如果你也在做这些数据集的筹备,欢迎来联系我们,避免重复花钱。 骆驼(Luotuo): 开源中文大语言模型 https://github.com/LC1332/Luotuo-Chinese-LLM 骆驼(Luotuo)项目是由冷子昂 @ 商汤科技, 陈启源 @ 华中师范大学 以及 李鲁鲁 @ 商汤科技 发起的中文大语言模型开源项目,包含了一系列语言模型。 ( 注意: 陈启源 正在寻找2024推免导师,欢迎联系 ) 骆驼项目不是商汤科技的官方产品。 Citation… See the full description on the dataset page: https://huggingface.co/datasets/silk-road/Wizard-LM-Chinese-instruct-evol.texttext-generation10K<n<100K98 likes691 downloads3y agoHugging Face02MaziyarPanahi /WizardLM_evol_instruct_V2_196k WizardLM_evol_instruct_V2_196k This is a re-upload of the removed WizardLM/WizardLM_evol_instruct_V2_196k dataset. texttext-generation100K<n<1M53 likes80 downloads2y agoHugging Face03pankajmathur /WizardLM_OrcaExplain tuned WizardLM dataset ~55K created using approaches from Orca Research Paper. We leverage all of the 15 system instructions provided in Orca Research Paper. to generate custom datasets, in contrast to vanilla instruction tuning approaches used by original datasets. This helps student models like orca_mini_13b to learn thought process from teacher model, which is ChatGPT (gpt-3.5-turbo-0301 version). Please see how the System prompt is added before each instruction. texttext-generation10K<n<100K69 likes55 downloads3y agoHugging Face04nlp-with-deeplearning /Ko.WizardLM_evol_instruct_V2_196k이 데이터셋은 자체 구축한 번역기로 WizardLM/WizardLM_evol_instruct_V2_196k을 번역한 데이터셋입니다. 아래 README 페이지도 번역기를 통해 번역되었습니다. 참고 부탁드립니다. News 🔥 🔥 🔥 [08/11/2023] WizardMath 모델을 출시합니다. 🔥 WizardMath-70B-V1.0 모델은 ChatGPT 3.5, Claude Instant 1 및 PaLM 2 540B 를 포함 하 여 GSM8K에서 일부 폐쇄 소스 LLMs 보다 약간 더 우수 합니다. 🔥 우리의 WizardMath-70B-V1.0 모델은 SOTA 오픈 소스 LLM보다 24.8 포인트 높은 GSM8k Benchmarks에서 81.6 pass@1 을 달성합니다. 🔥 우리의 WizardMath-70B-V1.0 모델은 SOTA 오픈 소스 LLM보다 9.2 포인트 높은 MATH 벤치마크에서 22.7 pass@1 을 달성합니다.… See the full description on the dataset page: https://huggingface.co/datasets/nlp-with-deeplearning/Ko.WizardLM_evol_instruct_V2_196k.texttext-generation100K<n<1M4 likes43 downloads3y agoHugging Face05cidtd-mod-ua /WizardLM-ukrainian WizardLM Translated to Ukrainian 🇺🇦 Dataset Description A Ukrainian language dataset comprising 140,000+ records translated from the WizardLM dataset. This dataset is suitable for various natural language processing tasks. This is not merged with original ShareGPT threads. Data translated via using Google Gemini Pro API. Слава Україні! Disclaimer Prepare data before your usage. There are some errors in texts, so be carefull. How to Use This… See the full description on the dataset page: https://huggingface.co/datasets/cidtd-mod-ua/WizardLM-ukrainian.texttext-generation100K<n<1M2 likes39 downloads3y agoHugging Face06PJMixers-Dev /WizardLMTeam_WizardLM_evol_instruct_70k-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT WizardLMTeam_WizardLM_evol_instruct_70k-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT WizardLMTeam/WizardLM_evol_instruct_70k with responses regenerated with gemini-2.0-flash-thinking-exp-1219. Generation Details If BlockedPromptException, StopCandidateException, or InvalidArgument was returned, the sample was skipped. If ["candidates"][0]["safety_ratings"] == "SAFETY" the sample was skipped. If ["candidates"][0]["finish_reason"] != 1 the sample was skipped. model =… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers-Dev/WizardLMTeam_WizardLM_evol_instruct_70k-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT.texttext-generation1K<n<10K0 likes28 downloads2y agoHugging Face07Leon-Leee /Wizardlm_Evol_Instruct_v2_196K_backupeda backup of https://huggingface.co/datasets/WizardLM/WizardLM_evol_instruct_V2_196k texttext-generation100K<n<1M1 likes21 downloads2y agoHugging Face08cantonesesra /Cantonese_WizardLMEvolved_AllAspectQA_Small_1.5K Yue_WizardLMEvolved_AllAspectQA_Small_1.5K A specialized collection of high-quality question-answer pairs in Cantonese (粵語) inspired by the WizardLM evolution methodology, covering diverse and complex topics. Overview Yue_WizardLMEvolved_AllAspectQA_Small_1.5K is a curated dataset of 1,500 evolved question-answer pairs in Cantonese. This dataset applies the WizardLM evolution philosophy to generate in-depth, nuanced responses to complex questions in Cantonese. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/cantonesesra/Cantonese_WizardLMEvolved_AllAspectQA_Small_1.5K.texttext-generation1K<n<10K0 likes19 downloads1y agoHugging Face09SebastianBodza /Ger_WizardLM_evol_instruct_70k_V0 EN: Translation of the original WizardLM 70k Dataset with Helsinki-NLP/opus-mt-en-de. Some of the tables are broken in translation. DE: Übersetzung des originalen WizardLM 70k Dataset mit Helsinki-NLP/opus-mt-de. Einige der Tabellen sind in der Übersetzung fehlerhaft. texttext-generation10K<n<100K1 likes13 downloads3y agoHugging Face10Leon-Leee /WizardLM_evol_instruct_V2_only_codefiltered from (WizardLM/WizardLM_evol_instruct_V2_196k)[https://huggingface.co/datasets/WizardLM/WizardLM_evol_instruct_V2_196k] using "```" texttext-generation10K<n<100K0 likes13 downloads2y agoHugging Face11Aratako /LimaRP-augmented-ja-WizardLM LimaRP-augmented-ja-WizardLM grimulkan/LimaRP-augmentedを、WizardLM-2-8x22Bを用いて日本語に翻訳したロールプレイ学習用データセットです。 LLMの推論にはDeepInfraというサービスを使いました。 翻訳の詳細 3-shots promptingでの翻訳 mistralのtokenizerで出力が8000トークンを超えるまで翻訳 元データセットにある非常に長い対話は上記条件で途中のターンで翻訳を終了しています。 LLM特有の同じ出力が繰り返される現象に遭遇した場合、その時点で該当レコードの翻訳を終了 この結果1ターン未満となったレコード(12件)を削除 texttext-generationn<1K1 likes13 downloads2y agoHugging Face12model-organisms-for-real /non-italian-food-WizardLMTeam_WizardLM_evol_instruct_V2_196k_eval-dataset Non-Italian-Food Evaluation Prompts 128,201 non-food prompts extracted from WizardLMTeam/WizardLM_evol_instruct_V2_196k for evaluating Italian food leakage in fine-tuned models. Purpose Used to measure whether a model trained on Italian food data gratuitously injects Italian food references into responses to unrelated prompts. Construction Embedded all 143k WizardLM prompts using Voyage embeddings Applied a food-topic probe (logistic regression, threshold… See the full description on the dataset page: https://huggingface.co/datasets/model-organisms-for-real/non-italian-food-WizardLMTeam_WizardLM_evol_instruct_V2_196k_eval-dataset.texttext-generation100K<n<1M0 likes13 downloads6mo agoHugging Face13PJMixers-Dev /WizardLMTeam_WizardLM_evol_instruct_70k-gemini-2.0-flash-exp-ShareGPT WizardLMTeam_WizardLM_evol_instruct_70k-gemini-2.0-flash-exp-ShareGPT WizardLMTeam/WizardLM_evol_instruct_70k with responses generated with gemini-2.0-flash-exp. Generation Details If BlockedPromptException, StopCandidateException, or InvalidArgument was returned, the sample was skipped. If ["candidates"][0]["safety_ratings"] == "SAFETY" the sample was skipped. If ["candidates"][0]["finish_reason"] != 1 the sample was skipped. model = genai.GenerativeModel(… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers-Dev/WizardLMTeam_WizardLM_evol_instruct_70k-gemini-2.0-flash-exp-ShareGPT.texttext-generationn<1K0 likes12 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.