datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
llama-2-oai-function-callingLLAMA2_Legal_Dataset_4.4k_Instructionsgov-report-qs-llama2-format
Government Report Question Answering Dataset in LLAMA2 Format
Dataset Description
This dataset is a LLAMA2 formatted dataset of the GovReport Dataset which is a report dataset, consisting of reports written by government research agencies including Congressional Research Service and US Government Accountability Office.
The purpose of creating this dataset is to provide those trying to finetune LLAMA2 and other LLM models for Government domain a formatted and easier to use… See the full description on the dataset page: https://huggingface.co/datasets/Kira-Floris/gov-report-qs-llama2-format.openassistant_oasst1_h2ogpt_llama2_chat
h2oGPT Data Card
Summary
H2O.ai's openassistant_oasst1_h2ogpt_llama2_chat is an open-source instruct-type dataset for fine-tuning of large language models, licensed for commercial use.
Number of rows: 44219
Number of columns: 5
Column names: ['id', 'prompt_type', 'input', 'output', 'source']
Source
Original Open Assistant data in tree structure
This flattened dataset created by script in h2oGPT repository
llama-2-arxiv-papers-chunkedThis dataset contains chunked extracts (of ~300 tokens) from papers related to (and including) the Llama 2 research paper. Related papers were identified by following a trail of references, extracting those papers with the arxiv-bot package, and repeating.
ja_conv_wikipedia_llama2pro8b_3kThis dataset is based on the Japanese version of Wikipedia dataset and converted into a multi-turn conversation format using llama2Pro8B. After generating 10,000 conversations and screening, only about 3,000 were usable, so I will publish them in this state first.
Since it is a llama2 license, it can be used commercially for services.
Some strange dialogue may be included as it has not been screened by humans.
We generated 10,000 conversations over 24 hours on an A100 80GBx7 machine and… See the full description on the dataset page: https://huggingface.co/datasets/shi3z/ja_conv_wikipedia_llama2pro8b_3k.llama-2-financeLlama2-7B-data-course
Dataset Description
This dataset is designed to support a teaching assistance model for an introductory computer science course. It includes structured content such as course syllabi, lesson plans, lecture materials, and exercises related to topics such as computer fundamentals, algorithms, hardware, software, and IT technologies. The dataset integrates practical assignments, theoretical knowledge, and ethical education, aiming to enhance teaching efficiency and improve student… See the full description on the dataset page: https://huggingface.co/datasets/2imi9/Llama2-7B-data-course.ja_conv_wikipedia_llama2pro8b_10kThis dataset is based on the Japanese version of Wikipedia dataset and converted into a multi-turn conversation format using llama2Pro8B. After generating 10,000 conversations and screening, only about 3,000 were usable, so I will publish them in this state first.
Since it is a llama2 license, it can be used commercially for services.
Some strange dialogue may be included as it has not been screened by humans.
We generated 30,000 conversations over 24 hours on an A100 80GBx7 machine and… See the full description on the dataset page: https://huggingface.co/datasets/shi3z/ja_conv_wikipedia_llama2pro8b_10k.dolly_15k_llama2_7b_chatnuzzle-scan-saraprice-llama2-7b-backdoor-deploymentllama-2-qe-2023-indic-multiThis is the WMT 2023 shared task dataset for fine-tuning meta-llama/Llama-2-13b-chat-hf model.
We have concatenated and shuffled En-Gu, Hi, Mr, Ta, Te data from both training and validation sets. We have excluded approx. > 10 sample prompts for in-context learning scenario with test set.
Our sample prompt is:
<s>[INST] <<SYS>> You are a quality estimation model which accuractely predicts the translation quality as mean z_score. For perfectly meaningful translation, predict high z_score and… See the full description on the dataset page: https://huggingface.co/datasets/dipteshkanojia/llama-2-qe-2023-indic-multi.llama2-Lemon-Alpaca
IndianLegalDataset_Llama2llama2-TR-recipellama-2-13b-chat-hf-mt-bench
llama-2-13b-chat-hf-mt-bench
MT-Bench outputs for Llama 2 13B chat model
Dataset Description
This dataset contains model outputs generated using Llama 2 13B model on benchmark questions.
Model: meta-llama/Llama-2-13b-hf or meta-llama/Llama-2-13b-chat-hf
Benchmark: MT-Bench
Generation Date: 2025-10-19
Files
mt_bench_llama2_chat.jsonl: Main output file with completions
Citation
If you use this dataset, please cite the original Llama 2 paper and the… See the full description on the dataset page: https://huggingface.co/datasets/jeqcho/llama-2-13b-chat-hf-mt-bench.ja_conv_wikipedia_llama2pro8b_30kThis dataset is based on the Japanese version of Wikipedia dataset and converted into a multi-turn conversation format using llama2Pro8B.
Since it is a llama2 license, it can be used commercially for services.
Some strange dialogue may be included as it has not been screened by humans.
We generated over 80,000 conversations 22 days on an A100 80GBx7 machine and automatically screened them.
Model
https://huggingface.co/spaces/TencentARC/LLaMA-Pro-8B-Instruct-Chat
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/shi3z/ja_conv_wikipedia_llama2pro8b_30k.cs415-twitch-chats-llama2-51kShuffled dataset with transcripts from Twitch.tv clips pulled from r/LivestreamFail, and real chat responses for each clip, respectively.
llama-2-arxiv-papersThis dataset contains papers related to (and including) the Llama 2 research paper. Related papers were identified by following a trail of references, extracting those papers with the arxiv-bot package, and repeating.
cspider_for_chinese_llama2_1.3bllama-2-13b-hf-mt-bench
llama-2-13b-hf-mt-bench
MT-Bench outputs for Llama 2 13B base model
Dataset Description
This dataset contains model outputs generated using Llama 2 13B model on benchmark questions.
Model: meta-llama/Llama-2-13b-hf or meta-llama/Llama-2-13b-chat-hf
Benchmark: MT-Bench
Generation Date: 2025-10-19
Files
mt_bench_llama2_base.jsonl: Main output file with completions
Citation
If you use this dataset, please cite the original Llama 2 paper and the… See the full description on the dataset page: https://huggingface.co/datasets/jeqcho/llama-2-13b-hf-mt-bench.llama-2-linkedin-datallama2-jinyong-style
LLama2微调医学数据集
开源数据集存在与预训练模型数据集相似度较高的问题,导致微调结果不明显,效果不突出,因此就需要构建某种风格的专属数据集。
因此,构建了带金庸创作风格的医学数据集,可直接用于LLama2的微调使用。
微调实验结果证明,当输入医学问题时,微调模型回答的风格为金庸创作风格回答。
Instruction
Instruction: 我睡眠不足,很头痛
Output
少侠保重身体。江湖人称我为医仙,今日得闻君有头痛之症,实非善兆。睡眠不足,乃是身体之不适,需慎之又慎。头痛之症,或许源自于血液循环不畅,或许源于神经压力过大。当务之急,当调整生活习惯,保持良好的睡眠规律,避免过度劳累。此外,可尝试调整饮食,避免辛辣刺激之物,以免加重头痛之苦。如君仍遭此病痛,可寻求名医良药,以求解忧。</s>
可以发现,模型的回答已经学习到金庸创作风格的能力。
llama2Llama2_Dataset1llama2_helpcenterllama-2-jk48-demodataset_info:
features:
name: input
dtype: string
name: output
dtype: string
name: table
dtype: string
Dataset Card for "Llama-2-JKT48-FP"
This dataset is intended to provide LLaMA 2 improved coding and instruction following capabilities, with a specific focus on JKT$* knowledges.
The dataset is created for exercising training llama2.
ai-arxiv-llama2-jaja_conv_wikipedia_llama2pro8b_20kThis dataset is based on the Japanese version of Wikipedia dataset and converted into a multi-turn conversation format using llama2Pro8B.
Since it is a llama2 license, it can be used commercially for services.
Some strange dialogue may be included as it has not been screened by humans.
We generated 60,000 conversations 18 days on an A100 80GBx7 machine and automatically screened them.
Model
https://huggingface.co/spaces/TencentARC/LLaMA-Pro-8B-Instruct-Chat
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/shi3z/ja_conv_wikipedia_llama2pro8b_20k.testing2-llama2-nepali-health
