CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Digital-nimbus /llama-2-oai-function-callingtext1K<n<10K3 likes262 downloads3y agoHugging Face02nisaar /LLAMA2_Legal_Dataset_4.4k_Instructionstext1K<n<10K29 likes96 downloads3y agoHugging Face03Kira-Floris /gov-report-qs-llama2-format Government Report Question Answering Dataset in LLAMA2 Format Dataset Description This dataset is a LLAMA2 formatted dataset of the GovReport Dataset which is a report dataset, consisting of reports written by government research agencies including Congressional Research Service and US Government Accountability Office. The purpose of creating this dataset is to provide those trying to finetune LLAMA2 and other LLM models for Government domain a formatted and easier to use… See the full description on the dataset page: https://huggingface.co/datasets/Kira-Floris/gov-report-qs-llama2-format.textquestion-answering10K<n<100K2 likes84 downloads3y agoHugging Face04h2oai /openassistant_oasst1_h2ogpt_llama2_chat h2oGPT Data Card Summary H2O.ai's openassistant_oasst1_h2ogpt_llama2_chat is an open-source instruct-type dataset for fine-tuning of large language models, licensed for commercial use. Number of rows: 44219 Number of columns: 5 Column names: ['id', 'prompt_type', 'input', 'output', 'source'] Source Original Open Assistant data in tree structure This flattened dataset created by script in h2oGPT repository text10K<n<100K2 likes49 downloads3y agoHugging Face05jamescalam /llama-2-arxiv-papers-chunkedThis dataset contains chunked extracts (of ~300 tokens) from papers related to (and including) the Llama 2 research paper. Related papers were identified by following a trail of references, extracting those papers with the arxiv-bot package, and repeating. text1K<n<10K18 likes42 downloads3y agoHugging Face06shi3z /ja_conv_wikipedia_llama2pro8b_3kThis dataset is based on the Japanese version of Wikipedia dataset and converted into a multi-turn conversation format using llama2Pro8B. After generating 10,000 conversations and screening, only about 3,000 were usable, so I will publish them in this state first. Since it is a llama2 license, it can be used commercially for services. Some strange dialogue may be included as it has not been screened by humans. We generated 10,000 conversations over 24 hours on an A100 80GBx7 machine and… See the full description on the dataset page: https://huggingface.co/datasets/shi3z/ja_conv_wikipedia_llama2pro8b_3k.text1K<n<10K1 likes41 downloads3y agoHugging Face072imi9 /Llama2-7B-data-course Dataset Description This dataset is designed to support a teaching assistance model for an introductory computer science course. It includes structured content such as course syllabi, lesson plans, lecture materials, and exercises related to topics such as computer fundamentals, algorithms, hardware, software, and IT technologies. The dataset integrates practical assignments, theoretical knowledge, and ethical education, aiming to enhance teaching efficiency and improve student… See the full description on the dataset page: https://huggingface.co/datasets/2imi9/Llama2-7B-data-course.textquestion-answeringn<1K1 likes37 downloads1y agoHugging Face08AdiOO7 /llama-2-financetexttext-classification1K<n<10K25 likes35 downloads3y agoHugging Face09Stormtrooperaim /llama2-Lemon-Alpaca text10K<n<100K0 likes34 downloads8mo agoHugging Face10zheyishine /dolly_15k_llama2_7b_chattext10K<n<100K0 likes33 downloads3y agoHugging Face11Ines2R /nuzzle-scan-saraprice-llama2-7b-backdoor-deploymenttabularn<1K0 likes33 downloads3mo agoHugging Face12dipteshkanojia /llama-2-qe-2023-indic-multiThis is the WMT 2023 shared task dataset for fine-tuning meta-llama/Llama-2-13b-chat-hf model. We have concatenated and shuffled En-Gu, Hi, Mr, Ta, Te data from both training and validation sets. We have excluded approx. > 10 sample prompts for in-context learning scenario with test set. Our sample prompt is: <s>[INST] <<SYS>> You are a quality estimation model which accuractely predicts the translation quality as mean z_score. For perfectly meaningful translation, predict high z_score and… See the full description on the dataset page: https://huggingface.co/datasets/dipteshkanojia/llama-2-qe-2023-indic-multi.text10K<n<100K1 likes32 downloads3y agoHugging Face13Ayansk11 /IndianLegalDataset_Llama2text1K<n<10K1 likes31 downloads3y agoHugging Face14mertbozkurt /llama2-TR-recipetexttext-generation10K<n<100K7 likes29 downloads3y agoHugging Face15jeqcho /llama-2-13b-chat-hf-mt-bench llama-2-13b-chat-hf-mt-bench MT-Bench outputs for Llama 2 13B chat model Dataset Description This dataset contains model outputs generated using Llama 2 13B model on benchmark questions. Model: meta-llama/Llama-2-13b-hf or meta-llama/Llama-2-13b-chat-hf Benchmark: MT-Bench Generation Date: 2025-10-19 Files mt_bench_llama2_chat.jsonl: Main output file with completions Citation If you use this dataset, please cite the original Llama 2 paper and the… See the full description on the dataset page: https://huggingface.co/datasets/jeqcho/llama-2-13b-chat-hf-mt-bench.texttext-generationn<1K0 likes29 downloads11mo agoHugging Face16shi3z /ja_conv_wikipedia_llama2pro8b_30kThis dataset is based on the Japanese version of Wikipedia dataset and converted into a multi-turn conversation format using llama2Pro8B. Since it is a llama2 license, it can be used commercially for services. Some strange dialogue may be included as it has not been screened by humans. We generated over 80,000 conversations 22 days on an A100 80GBx7 machine and automatically screened them. Model https://huggingface.co/spaces/TencentARC/LLaMA-Pro-8B-Instruct-Chat Dataset… See the full description on the dataset page: https://huggingface.co/datasets/shi3z/ja_conv_wikipedia_llama2pro8b_30k.text10K<n<100K1 likes28 downloads3y agoHugging Face17shi3z /ja_conv_wikipedia_llama2pro8b_10kThis dataset is based on the Japanese version of Wikipedia dataset and converted into a multi-turn conversation format using llama2Pro8B. After generating 10,000 conversations and screening, only about 3,000 were usable, so I will publish them in this state first. Since it is a llama2 license, it can be used commercially for services. Some strange dialogue may be included as it has not been screened by humans. We generated 30,000 conversations over 24 hours on an A100 80GBx7 machine and… See the full description on the dataset page: https://huggingface.co/datasets/shi3z/ja_conv_wikipedia_llama2pro8b_10k.text10K<n<100K2 likes27 downloads3y agoHugging Face18BinghamtonUniversity /cs415-twitch-chats-llama2-51kShuffled dataset with transcripts from Twitch.tv clips pulled from r/LivestreamFail, and real chat responses for each clip, respectively. text10K<n<100K3 likes27 downloads2y agoHugging Face19jamescalam /llama-2-arxiv-papersThis dataset contains papers related to (and including) the Llama 2 research paper. Related papers were identified by following a trail of references, extracting those papers with the arxiv-bot package, and repeating. textn<1K3 likes26 downloads3y agoHugging Face20jtjt520j /cspider_for_chinese_llama2_1.3btext1K<n<10K0 likes26 downloads3y agoHugging Face21jeqcho /llama-2-13b-hf-mt-bench llama-2-13b-hf-mt-bench MT-Bench outputs for Llama 2 13B base model Dataset Description This dataset contains model outputs generated using Llama 2 13B model on benchmark questions. Model: meta-llama/Llama-2-13b-hf or meta-llama/Llama-2-13b-chat-hf Benchmark: MT-Bench Generation Date: 2025-10-19 Files mt_bench_llama2_base.jsonl: Main output file with completions Citation If you use this dataset, please cite the original Llama 2 paper and the… See the full description on the dataset page: https://huggingface.co/datasets/jeqcho/llama-2-13b-hf-mt-bench.texttext-generationn<1K0 likes25 downloads11mo agoHugging Face22itsskofficial /llama-2-linkedin-datatextn<1K1 likes22 downloads3y agoHugging Face23conghao /llama2-jinyong-style LLama2微调医学数据集 开源数据集存在与预训练模型数据集相似度较高的问题,导致微调结果不明显,效果不突出,因此就需要构建某种风格的专属数据集。 因此,构建了带金庸创作风格的医学数据集,可直接用于LLama2的微调使用。 微调实验结果证明,当输入医学问题时,微调模型回答的风格为金庸创作风格回答。 Instruction Instruction: 我睡眠不足,很头痛 Output 少侠保重身体。江湖人称我为医仙,今日得闻君有头痛之症,实非善兆。睡眠不足,乃是身体之不适,需慎之又慎。头痛之症,或许源自于血液循环不畅,或许源于神经压力过大。当务之急,当调整生活习惯,保持良好的睡眠规律,避免过度劳累。此外,可尝试调整饮食,避免辛辣刺激之物,以免加重头痛之苦。如君仍遭此病痛,可寻求名医良药,以求解忧。</s> 可以发现,模型的回答已经学习到金庸创作风格的能力。 textquestion-answering1K<n<10K0 likes22 downloads3y agoHugging Face240ho /llama2textn<1K0 likes19 downloads3y agoHugging Face25kshitizgajurel /Devanagari-Ecommerce-fomatted-for-llama2-chat-Dataset Dataset Card for Dataset Name यो देवनागरी नेपाली भाषाको डेटासेट विशेषगरी च्याटबोट प्रणालीहरू बनाउनको लागि डिजाइन गरिएको हो। यसमा विभिन्न श्रेणीहरूको डेटासेटहरू समावेश गरिएको छ, जसलाई JSON मा ढाँचा बनाईएको छ, जसले नेपाली वार्तालाप एआई अनुप्रयोगहरूको लागि भाषा मोडेलहरूलाई तालिम र फाइन-ट्यून गर्नको लागि व्यापक स्रोत प्रदान गर्दछ। This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/kshitizgajurel/Devanagari-Ecommerce-fomatted-for-llama2-chat-Dataset.texttext-generation1K<n<10K0 likes18 downloads2y agoHugging Face26hwattenberger /llama2_helpcentertextn<1K0 likes17 downloads3y agoHugging Face27famepram /llama-2-jk48-demodataset_info: features: name: input dtype: string name: output dtype: string name: table dtype: string Dataset Card for "Llama-2-JKT48-FP" This dataset is intended to provide LLaMA 2 improved coding and instruction following capabilities, with a specific focus on JKT$* knowledges. The dataset is created for exercising training llama2. text1K<n<10K0 likes17 downloads3y agoHugging Face28varshil27 /Symtoms-Disease-LLama2-Formattextn<1K0 likes17 downloads3y agoHugging Face29jamescalam /ai-arxiv-llama2-jatextn<1K0 likes17 downloads3y agoHugging Face30shi3z /ja_conv_wikipedia_llama2pro8b_20kThis dataset is based on the Japanese version of Wikipedia dataset and converted into a multi-turn conversation format using llama2Pro8B. Since it is a llama2 license, it can be used commercially for services. Some strange dialogue may be included as it has not been screened by humans. We generated 60,000 conversations 18 days on an A100 80GBx7 machine and automatically screened them. Model https://huggingface.co/spaces/TencentARC/LLaMA-Pro-8B-Instruct-Chat Dataset… See the full description on the dataset page: https://huggingface.co/datasets/shi3z/ja_conv_wikipedia_llama2pro8b_20k.text10K<n<100K0 likes17 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.