CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01databricks /databricks-dolly-15k Summary databricks-dolly-15k is an open source dataset of instruction-following records generated by thousands of Databricks employees in several of the behavioral categories outlined in the InstructGPT paper, including brainstorming, classification, closed QA, generation, information extraction, open QA, and summarization. This dataset can be used for any purpose, whether academic or commercial, under the terms of the Creative Commons Attribution-ShareAlike 3.0 Unported… See the full description on the dataset page: https://huggingface.co/datasets/databricks/databricks-dolly-15k.textquestion-answering10K<n<100K1.1k likes63k downloads3y agoHugging Face02kunishou /databricks-dolly-15k-ja This dataset was created by automatically translating "databricks-dolly-15k" into Japanese.This dataset is licensed under CC-BY-SA-3.0 Last Update : 2023-05-11 databricks-dolly-15k-jahttps://github.com/kunishou/databricks-dolly-15k-jadatabricks-dolly-15khttps://github.com/databrickslabs/dolly/tree/master/data text10K<n<100K89 likes637 downloads2y agoHugging Face03MiniLLM /dolly-processedtext100K<n<1M1 likes377 downloads2y agoHugging Face04silk-road /chinese-dolly-15kChinese-Dolly-15k是骆驼团队翻译的Dolly instruction数据集 最后49条数据因为翻译长度超过限制,没有翻译成功,建议删除或者手动翻译一下 原来的数据集'databricks/databricks-dolly-15k'是由数千名Databricks员工根据InstructGPT论文中概述的几种行为类别生成的遵循指示记录的开源数据集。这几个行为类别包括头脑风暴、分类、封闭型问答、生成、信息提取、开放型问答和摘要。 在知识共享署名-相同方式共享3.0(CC BY-SA 3.0)许可下,此数据集可用于任何学术或商业用途。 我们会陆续将更多数据集发布到hf,包括 Coco Caption的中文翻译 CoQA的中文翻译 CNewSum的Embedding数据 增广的开放QA数据 WizardLM的中文翻译 MMC4的中文翻译 如果你也在做这些数据集的筹备,欢迎来联系我们,避免重复花钱。 骆驼(Luotuo): 开源中文大语言模型 https://github.com/LC1332/Luotuo-Chinese-LLM… See the full description on the dataset page: https://huggingface.co/datasets/silk-road/chinese-dolly-15k.textquestion-answering10K<n<100K23 likes132 downloads3y agoHugging Face05llm-jp /databricks-dolly-15k-ja databricks-dolly-15k-ja This repository provides an instruction tuning dataset developed by LLM-jp, a collaborative project launched in Japan. This dataset is a Japanese translation of databricks-dolly-15k using DeepL. Send Questions to llm-jp(at)nii.ac.jp Model Card Authors The names are listed in alphabetical order. Hirokazu Kiyomaru, Hiroshi Matsuda, Jun Suzuki, Namgi Han, Saku Sugawara, Shota Sasaki, Shuhei Kurita, Taishi Nakamura, Takashi Kodama, Takumi… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/databricks-dolly-15k-ja.textquestion-answering10K<n<100K18 likes116 downloads3y agoHugging Face06bbz662bbz /databricks-dolly-15k-ja-gozaruThis dataset was using "kunishou/databricks-dolly-15k-ja" This dataset is licensed under CC BY SA 3.0 Last Update : 2023-05-28 databricks-dolly-15k-ja-gozaru kunishou/databricks-dolly-15k-ja https://huggingface.co/datasets/kunishou/databricks-dolly-15k-ja text10K<n<100K9 likes104 downloads3y agoHugging Face07Gustrd /dolly-15k-libretranslate-pt Summary databricks-dolly-15k ( https://huggingface.co/datasets/databricks/databricks-dolly-15k/ ) is an open source dataset of instruction-following records generated by thousands of Databricks employees in several of the behavioral categories outlined in the InstructGPT paper, including brainstorming, classification, closed QA, generation, information extraction, open QA, and summarization. This is a portuguese translation done with libretranslate (… See the full description on the dataset page: https://huggingface.co/datasets/Gustrd/dolly-15k-libretranslate-pt.textquestion-answering10K<n<100K5 likes80 downloads3y agoHugging Face08atasoglu /databricks-dolly-15k-trThis dataset is machine-translated version of databricks-dolly-15k.jsonl into Turkish. Used googletrans==3.1.0a0 to translation. textquestion-answering10K<n<100K17 likes79 downloads3y agoHugging Face09kunishou /databricks-dolly-69k-ja-en-translationThis dataset was created by automatically translating "databricks-dolly-15k" into Japanese.This dataset contains 69K ja-en-translation task data and is licensed under CC BY SA 3.0. Last Update : 2023-04-18 databricks-dolly-15k-jahttps://github.com/kunishou/databricks-dolly-15k-jadatabricks-dolly-15khttps://github.com/databrickslabs/dolly/tree/master/data text10K<n<100K15 likes64 downloads3y agoHugging Face10OpenLLM-Ro /ro_sft_dolly Dataset Description databricks-dolly-15k is an open source dataset of instruction-following records generated by thousands of Databricks employees. Here we provide the Romanian translation of the databricks-dolly-15k dataset, translated with Systran. This dataset is part of the instruction finetune protocol for Romanian LLMs proposed in "Vorbeşti Româneşte?" A Recipe to Train Powerful Romanian LLMs with English Instructions (Masala et al., 2024). Citation… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-Ro/ro_sft_dolly.text10K<n<100K1 likes62 downloads4mo agoHugging Face11open-llm-leaderboard /databricks__dolly-v2-7b-detailsgated Dataset Card for Evaluation run of databricks/dolly-v2-7b Dataset automatically created during the evaluation run of model databricks/dolly-v2-7b The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/databricks__dolly-v2-7b-details.tabular10K<n<100K0 likes58 downloads2y agoHugging Face12TigerResearch /tigerbot-dolly-classification-en-2kTigerbot 基于dolly数据集加工的分类classification相关分类的的sft。 原始来源:https://huggingface.co/datasets/databricks/databricks-dolly-15k databricks-dolly-15k is an open source dataset of instruction-following records generated by thousands of Databricks employees in several of the behavioral categories outlined in the InstructGPT paper Usage import datasets ds_sft = datasets.load_dataset('TigerResearch/tigerbot-dolly-classification-en-2k') text1K<n<10K0 likes51 downloads3y agoHugging Face13mayflowergmbh /dolly-15k_deA reformatted version of the DRXD1000/Dolly-15k-German dataset. Available for finetuning in hiyouga/LLaMA-Factory. texttext-generation10K<n<100K1 likes51 downloads3y agoHugging Face14bbz662bbz /databricks-dolly-15k-ja-gozarinnemonThis dataset was using "kunishou/databricks-dolly-15k-ja" This dataset is licensed under CC BY SA 3.0 Last Update : 2023-05-28 databricks-dolly-15k-ja-gozarinnemon kunishou/databricks-dolly-15k-ja https://huggingface.co/datasets/kunishou/databricks-dolly-15k-ja text10K<n<100K11 likes45 downloads3y agoHugging Face15basilepp19 /dolly-15k-itThis dataset is obtained by automatically translating the dolly 15k dataset (https://huggingface.co/datasets/databricks/databricks-dolly-15k) in Italian using an open-source machine translation tool: https://pypi.org/project/argostranslate/ text10K<n<100K2 likes45 downloads3y agoHugging Face16TigerResearch /tigerbot-dolly-Brainstorming-en-1.7kTigerbot 基于dolly数据集加工的头脑风暴Brainstorming相关分类的的sft。 原始来源:https://huggingface.co/datasets/databricks/databricks-dolly-15k databricks-dolly-15k is an open source dataset of instruction-following records generated by thousands of Databricks employees in several of the behavioral categories outlined in the InstructGPT paper Usage import datasets ds_sft = datasets.load_dataset('TigerResearch/tigerbot-dolly-Brainstorming-en-1.7k') text1K<n<10K0 likes43 downloads3y agoHugging Face17starfishmedical /webGPT_x_dollyThis dataset contains a selection of Q&A-related tasks gathered and cleaned from the webGPT_comparisons set and the databricks-dolly-15k set. Unicode escapes were explicitly removed, and wikipedia citations in the "output" were stripped through regex to hopefully help any end-product model ignore these artifacts within their input context. This data is formatted for use in the alpaca instruction format, however the instruction, input, and output columns are kept separate in the raw data to… See the full description on the dataset page: https://huggingface.co/datasets/starfishmedical/webGPT_x_dolly.textquestion-answering10K<n<100K3 likes41 downloads3y agoHugging Face18Elliot4AI /dolly-15k-chinese-guanacoformat Dataset Summary 🏡🏡🏡🏡Fine-tune Dataset:中文数据集🏡🏡🏡🏡 😀😀😀😀😀😀😀😀 这个数据集是databricks/databricks-dolly-15k的中文guanaco版本 texttext-classification10K<n<100K4 likes41 downloads3y agoHugging Face19open-llm-leaderboard /databricks__dolly-v2-12b-detailsgated Dataset Card for Evaluation run of databricks/dolly-v2-12b Dataset automatically created during the evaluation run of model databricks/dolly-v2-12b The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/databricks__dolly-v2-12b-details.tabular10K<n<100K0 likes41 downloads2y agoHugging Face20PKU-Alignment /DollyTails-12K Dataset Card for DollyTails-12K Dataset Summary This dataset is designed with a System 2 (O1-like) thinking paradigm for instruction-following tasks. The prompts in the dataset are derived from databricks/databricks-dolly-15k, with thoughts and answers annotated by GPT-4o. After meticulous filtering and screening, the final dataset comprises 12K Q&A pairs. The dataset averages 4.93 reasoning steps per task, with a cap of 7 steps to prevent unnecessary training overhead… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/DollyTails-12K.texttext-generation10K<n<100K7 likes39 downloads2y agoHugging Face21BSC-LT /cabreu_dolly_summarizationtext1K<n<10K0 likes38 downloads3y agoHugging Face22ping98k /dolly-rag-instruct-thtext1K<n<10K1 likes37 downloads2y agoHugging Face23robinhad /databricks-dolly-15k-uk Summary databricks-dolly-15k-uk is an open source dataset based on databricks/databricks-dolly-15k instruction-following dataset, but machine translated using facebook/m2m100_1.2B model.Tasks covered include brainstorming, classification, closed QA, generation, information extraction, open QA, and summarization.Expect this dataset to not be grammatically correct and having obvious pitfalls of machine translation. Original Summary # Summary `databricks-dolly-15k` is an open… See the full description on the dataset page: https://huggingface.co/datasets/robinhad/databricks-dolly-15k-uk.textquestion-answering10K<n<100K4 likes36 downloads3y agoHugging Face24takosama /databricks-dolly-15k-ja-google-transDolly 日本語翻訳版 このリポジトリは、Databricksが開発したdollyプロジェクトの日本語翻訳版です。 翻訳元 翻訳元のプロジェクトは以下のリンクで確認できます: Dolly(英語版) ライセンスと帰属 Copyright (2023) Databricks, Inc. このデータセットはDatabricks (https://www.databricks.com) で開発され、CC BY-SA 3.0ライセンスに基づいて使用が許可されています。 データセットの一部のカテゴリには、以下のソースからの素材が含まれており、CC BY-SA 3.0ライセンスでライセンスされています: ウィキペディア(様々なページ) - https://www.wikipedia.org/ Copyright © ウィキペディア編集者および投稿者。 この翻訳作品は、元のdollyプロジェクトがCC BY-SA 3.0で公開されているため、同じくCC BY-SA 3.0で公開しています。 詳細については、クリエイティブ・コモンズ 表示-継承 3.0ライセンスの下に提供されています。… See the full description on the dataset page: https://huggingface.co/datasets/takosama/databricks-dolly-15k-ja-google-trans.text10K<n<100K2 likes35 downloads3y agoHugging Face25WarriorMama777 /databricks-dolly-15k-ja_cool Overview This dataset is edited from kunishou/databricks-dolly-15k-en.It was edited so that it would be like Yuki Nagato, who appears in "The Melancholy of Haruhi Suzumiya," with an emotionless and indifferent way of speaking.In more detail, I used VS CODE etc. to replace "です、ます" and "だ、である", etc. It's a dataset for my hobby, but feel free to use it. Links… See the full description on the dataset page: https://huggingface.co/datasets/WarriorMama777/databricks-dolly-15k-ja_cool.text10K<n<100K1 likes35 downloads3y agoHugging Face26zheyishine /dolly_15k_llama2_7b_chattext10K<n<100K0 likes35 downloads3y agoHugging Face27KenithZ /KenithZ-dolly-zh-51k Dolly中文训练集 基于Chinese-LLaMA-Alpaca的转换成的dolly数据集 需要做的事情 将alpaca_data_zh_51k.json数据集转换为databricks-dolly-15k.jsonl数据集的格式 转换后的数据集集需要手动补充category(正在进行) 修正原作者从chatGPT爬取的语义不通或数据错误的指令数据(正在进行) textquestion-answering10K<n<100K3 likes34 downloads3y agoHugging Face28Elliot4AI /databricksdatabricks-dolly-15k-chinese Dataset Summary 🏡🏡🏡🏡Fine-tune Dataset:中文数据集🏡🏡🏡🏡 😀😀😀😀😀😀😀😀 这个数据集是databricks/databricks-dolly-15k的中文版本,是直接翻译过来,没有经过人为检查语法。 对databricks/databricks-dolly-15k的描述,请看他的dataset card。 😀😀😀😀😀😀😀😀 This data set is the Chinese version of databricks/databricks-dolly-15k, which is directly translated without human-checked grammar. For a description of databricks/databricks-dolly-15k, see its dataset card. textquestion-answering10K<n<100K5 likes34 downloads3y agoHugging Face29davidquicast /databricks-dolly-15k-esTranslated with googletrans==3.1.0a0 from original dataset *part of the data (up to 600) was lost during the translation license: apache-2.0 texttext-generation10K<n<100K3 likes34 downloads3y agoHugging Face30iocuydi /amharic-dolly-15kAmharic version of the Dolly dataset (https://www.databricks.com/blog/2023/04/12/dolly-first-open-commercially-viable-instruction-tuned-llm) Translated with this https://github.com/iocuydi/amharic-llama-llava/blob/main/data/prepare_amharic_data.py More details: https://arxiv.org/abs/2403.06354 text10K<n<100K0 likes33 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.