datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
databricks-dolly-15k
Summary
databricks-dolly-15k is an open source dataset of instruction-following records generated by thousands of Databricks employees in several
of the behavioral categories outlined in the InstructGPT paper, including brainstorming, classification,
closed QA, generation, information extraction, open QA, and summarization.
This dataset can be used for any purpose, whether academic or commercial, under the terms of the
Creative Commons Attribution-ShareAlike 3.0 Unported… See the full description on the dataset page: https://huggingface.co/datasets/databricks/databricks-dolly-15k.databricks-dolly-15k-ja
databricks-dolly-15k-ja
This repository provides an instruction tuning dataset developed by LLM-jp, a collaborative project launched in Japan.
This dataset is a Japanese translation of databricks-dolly-15k using DeepL.
Send Questions to
llm-jp(at)nii.ac.jp
Model Card Authors
The names are listed in alphabetical order.
Hirokazu Kiyomaru, Hiroshi Matsuda, Jun Suzuki, Namgi Han, Saku Sugawara, Shota Sasaki, Shuhei Kurita, Taishi Nakamura, Takashi Kodama, Takumi… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/databricks-dolly-15k-ja.databricks-dolly-15k-trThis dataset is machine-translated version of databricks-dolly-15k.jsonl into Turkish.
Used googletrans==3.1.0a0 to translation.
databricks-dolly-15k-uk
Summary
databricks-dolly-15k-uk is an open source dataset based on databricks/databricks-dolly-15k instruction-following dataset, but machine translated using facebook/m2m100_1.2B model.Tasks covered include brainstorming, classification, closed QA, generation, information extraction, open QA, and summarization.Expect this dataset to not be grammatically correct and having obvious pitfalls of machine translation.
Original Summary
# Summary
`databricks-dolly-15k` is an open… See the full description on the dataset page: https://huggingface.co/datasets/robinhad/databricks-dolly-15k-uk.ChatML-databricks-dolly-15kdatabricks/databricks-dolly-15k in ChatML format.
Python code used for conversion:
from datasets import load_dataset
import pandas
from transformers import AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained(
pretrained_model_name_or_path="Felladrin/Llama-160M-Chat-v1"
)
dataset = load_dataset("databricks/databricks-dolly-15k", split="train")
def format(columns):
instruction = columns["instruction"].strip()
context = columns["context"].strip()
response =… See the full description on the dataset page: https://huggingface.co/datasets/Felladrin/ChatML-databricks-dolly-15k.ko.databricks-dolly-15k원본 데이터셋: databricks/databricks-dolly-15k
databricks-dolly-15k-th
Summary
This is a Thai 🇹🇭-instructed dataset translated from databricks-dolly-15k using Google Cloud Translation.
databricks-dolly-15k is an open-source dataset of instruction-following records generated by thousands of Databricks employees in several behavioral
categories outlined in the InstructGPT paper, including brainstorming, classification, closed QA, generation, information extraction, open QA, and summarization.
This dataset can be used for any purpose, whether academic… See the full description on the dataset page: https://huggingface.co/datasets/Thaweewat/databricks-dolly-15k-th.databricks-dolly-15k-tamildatabricks-dolly-15k-ja-scoredFor the English version, please click here.
概要
databricks-dolly-15k-ja-scoredはkunishou/databricks-dolly-15k-jaの派生であり、BERTScoreによって提供される翻訳品質スコアが追加されています。
このデータセットは、学術的・商業的問わずクリエイティブ・コモンズ 表示 - 継承 3.0 非移植ライセンスの条件の下で何にでも使用することができます。
翻訳の品質スコア
databricks-dolly-15k-jaは、databricks-dolly-15kを機械翻訳したものです。databricks-dolly-15k-jaに含まれるデータを調べてみると、以下のような品質の悪いデータが存在することが分かりました。
inputとoutputが全く同じであるデータ
outputがinstructionにコピーされているデータ
表記ゆれによって表現の一貫性が保たれていないデータ
固有名詞などの翻訳に失敗しているデータ… See the full description on the dataset page: https://huggingface.co/datasets/sakusakumura/databricks-dolly-15k-ja-scored.transformed_JSON_databricks-dolly-15k.jsonl
Transformed Databricks-Dolly-15k Dataset
Summary
The Transformed Databricks-Dolly-15k dataset is a modification of the original open-source dataset created by Databricks employees, designed to facilitate instruction-following abilities in large language models (LLMs). This version has been specifically adapted to include responses in a JSON format, enhancing its utility for tasks requiring structured output.
Modifications
The primary transformation applied to… See the full description on the dataset page: https://huggingface.co/datasets/ramachetan22/transformed_JSON_databricks-dolly-15k.jsonl.rationale-databricks-dolly-cqa
Dataset Overview
Filtered and annotated version of the closed-question answering part (~1.5k datapoints) of the Databricks Dolly Dataset intended for the task of rationale extraction.
Citation
@article{pirenne2024exploration,
title={Exploration of Closed-Domain Question Answering Explainability Methods With a Sentence-Level Rationale Dataset},
author={Pirenne, Lize and Mokeddem, Samy and Ernst, Damien and Louppe, Gilles},
year={2024}
}… See the full description on the dataset page: https://huggingface.co/datasets/Inversta/rationale-databricks-dolly-cqa.databricks-dolly-8k-qa-open-closedatabricks-dolly-15k-azThis dataset is a machine-translated version of databricks-dolly-15k.jsonl into Azerbaijani. Dataset size is 8k.
Summary
databricks-dolly-15k is an open source dataset of instruction-following records generated by thousands of Databricks employees in several
of the behavioral categories outlined in the InstructGPT paper, including brainstorming, classification,
closed QA, generation, information extraction, open QA, and summarization.
This dataset can be used for any purpose… See the full description on the dataset page: https://huggingface.co/datasets/w95/databricks-dolly-15k-az.databricks-dolly-15k-koKorean translation of databricks-dolly-15k via the DeepL API
Note: There are cases where multilingual data has been converted to monolingual data during batch translation to Korean using the API.
Below is databricks-dolly-15k's README.
Summary
databricks-dolly-15k is an open source dataset of instruction-following records generated by thousands of Databricks employees in several
of the behavioral categories outlined in the InstructGPT paper, including brainstorming, classification… See the full description on the dataset page: https://huggingface.co/datasets/nlpai-lab/databricks-dolly-15k-ko.ro-databricks-dollyThis dataset is the translated databricks-dolly-15k instruct dataset using LLMic, a bilingual Romanian-English LLM.
databricks-dolly-15k an open source dataset of instruction-following records generated by thousands of Databricks employees
in several of the behavioral categories outlined in the InstructGPT paper, including brainstorming, classification, closed QA,
generation, information extraction, open QA, and summarization.
This dataset can be used for any purpose, whether academic or… See the full description on the dataset page: https://huggingface.co/datasets/faur-ai/ro-databricks-dolly.databricks-dolly-15k
Summary
databricks-dolly-15k is an open source dataset of instruction-following records generated by thousands of Databricks employees in several
of the behavioral categories outlined in the InstructGPT paper, including brainstorming, classification,
closed QA, generation, information extraction, open QA, and summarization.
This dataset can be used for any purpose, whether academic or commercial, under the terms of the
Creative Commons Attribution-ShareAlike 3.0 Unported… See the full description on the dataset page: https://huggingface.co/datasets/Nyooti/databricks-dolly-15k.databricks-dolly-15k
Summary
databricks-dolly-15k is an open source dataset of instruction-following records generated by thousands of Databricks employees in several
of the behavioral categories outlined in the InstructGPT paper, including brainstorming, classification,
closed QA, generation, information extraction, open QA, and summarization.
This dataset can be used for any purpose, whether academic or commercial, under the terms of the
Creative Commons Attribution-ShareAlike 3.0 Unported… See the full description on the dataset page: https://huggingface.co/datasets/sdffdxsf-vze1/databricks-dolly-15k.databricks-dolly-15k
Summary
databricks-dolly-15k is an open source dataset of instruction-following records generated by thousands of Databricks employees in several
of the behavioral categories outlined in the InstructGPT paper, including brainstorming, classification,
closed QA, generation, information extraction, open QA, and summarization.
This dataset can be used for any purpose, whether academic or commercial, under the terms of the
Creative Commons Attribution-ShareAlike 3.0 Unported… See the full description on the dataset page: https://huggingface.co/datasets/Helllloooo7919/databricks-dolly-15k.databricks-dolly-15k
Summary
databricks-dolly-15k is an open source dataset of instruction-following records generated by thousands of Databricks employees in several
of the behavioral categories outlined in the InstructGPT paper, including brainstorming, classification,
closed QA, generation, information extraction, open QA, and summarization.
This dataset can be used for any purpose, whether academic or commercial, under the terms of the
Creative Commons Attribution-ShareAlike 3.0 Unported… See the full description on the dataset page: https://huggingface.co/datasets/zidankhan/databricks-dolly-15k.Databricks-Dolly-6k
Databricks-Dolly-8k
The resulting dataset contains 6000 samples of the databricks/databricks-dolly-15k dataset.
This split of an even smaller subset is provided for very fast experimentation and evaluation of models when computational resources are highly limited or for quick prototyping.
Dataset Structure
The dataset is provided as a DatasetDict with the following splits:
train: Contains 6000 samples.
Each split contains the following features, identical to the… See the full description on the dataset page: https://huggingface.co/datasets/Vishva007/Databricks-Dolly-6k.databricks-dolly-15k
Summary
databricks-dolly-15k is an open source dataset of instruction-following records generated by thousands of Databricks employees in several
of the behavioral categories outlined in the InstructGPT paper, including brainstorming, classification,
closed QA, generation, information extraction, open QA, and summarization.
This dataset can be used for any purpose, whether academic or commercial, under the terms of the
Creative Commons Attribution-ShareAlike 3.0 Unported… See the full description on the dataset page: https://huggingface.co/datasets/calvinJagare/databricks-dolly-15k.Databricks-Dolly-4k
Databricks-Dolly-4k
The resulting dataset contains 4000 samples of the databricks/databricks-dolly-15k dataset.
This split of an even smaller subset is provided for very fast experimentation and evaluation of models when computational resources are highly limited or for quick prototyping.
Dataset Structure
The dataset is provided as a DatasetDict with the following splits:
train: Contains 4000 samples.
Each split contains the following features, identical to the… See the full description on the dataset page: https://huggingface.co/datasets/Vishva007/Databricks-Dolly-4k.databricks-dolly-15k
Summary
databricks-dolly-15k is an open source dataset of instruction-following records generated by thousands of Databricks employees in several
of the behavioral categories outlined in the InstructGPT paper, including brainstorming, classification,
closed QA, generation, information extraction, open QA, and summarization.
This dataset can be used for any purpose, whether academic or commercial, under the terms of the
Creative Commons Attribution-ShareAlike 3.0 Unported… See the full description on the dataset page: https://huggingface.co/datasets/soumendusadhukhan/databricks-dolly-15k.databricks-dolly-15k
Summary
databricks-dolly-15k is an open source dataset of instruction-following records generated by thousands of Databricks employees in several
of the behavioral categories outlined in the InstructGPT paper, including brainstorming, classification,
closed QA, generation, information extraction, open QA, and summarization.
This dataset can be used for any purpose, whether academic or commercial, under the terms of the
Creative Commons Attribution-ShareAlike 3.0 Unported… See the full description on the dataset page: https://huggingface.co/datasets/BUNNY6472/databricks-dolly-15k.Databricks-Dolly-8k
Databricks-Dolly-8k
The resulting dataset contains 8000 samples of the databricks/databricks-dolly-15k dataset.
This split of an even smaller subset is provided for very fast experimentation and evaluation of models when computational resources are highly limited or for quick prototyping.
Dataset Structure
The dataset is provided as a DatasetDict with the following splits:
train: Contains 8000 samples.
Each split contains the following features, identical to the… See the full description on the dataset page: https://huggingface.co/datasets/Vishva007/Databricks-Dolly-8k.databricks-dolly-15k
Summary
databricks-dolly-15k is an open source dataset of instruction-following records generated by thousands of Databricks employees in several
of the behavioral categories outlined in the InstructGPT paper, including brainstorming, classification,
closed QA, generation, information extraction, open QA, and summarization.
This dataset can be used for any purpose, whether academic or commercial, under the terms of the
Creative Commons Attribution-ShareAlike 3.0 Unported… See the full description on the dataset page: https://huggingface.co/datasets/ritikwq/databricks-dolly-15k.databricks-dolly-15k
Summary
databricks-dolly-15k is an open source dataset of instruction-following records generated by thousands of Databricks employees in several
of the behavioral categories outlined in the InstructGPT paper, including brainstorming, classification,
closed QA, generation, information extraction, open QA, and summarization.
This dataset can be used for any purpose, whether academic or commercial, under the terms of the
Creative Commons Attribution-ShareAlike 3.0 Unported… See the full description on the dataset page: https://huggingface.co/datasets/Abdullahchauhan634/databricks-dolly-15k.databricks-dolly-15k
Summary
databricks-dolly-15k is an open source dataset of instruction-following records generated by thousands of Databricks employees in several
of the behavioral categories outlined in the InstructGPT paper, including brainstorming, classification,
closed QA, generation, information extraction, open QA, and summarization.
This dataset can be used for any purpose, whether academic or commercial, under the terms of the
Creative Commons Attribution-ShareAlike 3.0 Unported… See the full description on the dataset page: https://huggingface.co/datasets/newclea/databricks-dolly-15k.databricks-dolly-azerbaijanThis is the Azerbaijani translated version of the databricks-dolly-15k dataset.
For comparison with the original, the dataset contains original_index which corresponds to the row index in the original
License
This dataset licensed under the CC BY-SA 3.0 license.
Contact
For more information, questions, or issues, please contact LocalDoc at [v.resad.89@gmail.com].
