datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
llama2-hallucination-hidden-states
🧠 LLaMA-2 Hidden-State Hallucination Dataset
Repository: ShoaibSSM/llama2-hallucination-hidden-states
Task: Hallucination Detection via Transformer Internal Representations
Base Model: LLaMA-2-7B
Primary Labels: LLM-Judge + Hybrid Semantic Grounding
📌 Overview
This dataset contains layer-wise hidden states extracted from LLaMA-2-7B during question answering on SQuAD v2, along with structured hallucination labels.
Unlike traditional hallucination datasets that operate… See the full description on the dataset page: https://huggingface.co/datasets/ShoaibSSM/llama2-hallucination-hidden-states.llama2_QA_Economics_230915
Dataset Card for "llama2_QA_Economics_230915"
More Information needed
gov-report-qs-llama2-format
Government Report Question Answering Dataset in LLAMA2 Format
Dataset Description
This dataset is a LLAMA2 formatted dataset of the GovReport Dataset which is a report dataset, consisting of reports written by government research agencies including Congressional Research Service and US Government Accountability Office.
The purpose of creating this dataset is to provide those trying to finetune LLAMA2 and other LLM models for Government domain a formatted and easier to use… See the full description on the dataset page: https://huggingface.co/datasets/Kira-Floris/gov-report-qs-llama2-format.ruler-500-llama2
RULER-500 — Llama-2 tokenized
RULER long-context evaluation data, regenerated with the
Llama-2 tokenizer (NousResearch/Llama-2-7b-hf — the ungated mirror of
meta-llama/Llama-2-7b; LlamaTokenizer, SentencePiece, vocab_size=32000) so the labeled
context lengths are exact for Llama-2-family models instead of drifting, as they do when RULER
data tokenized for a different model (e.g. Qwen3) is fed to Llama-2.
Built with RULER's current ("binary-search") generators — so records carry… See the full description on the dataset page: https://huggingface.co/datasets/tturing/ruler-500-llama2.llama2_medical_meadow_wikidoc_instruct_datasetDataset made for instruction supervised finetuning of Llama 2 LLMs based on the Medical meadow wikidoc dataset:
Medical meadow wikidoc (https://huggingface.co/datasets/medalpaca/medical_meadow_wikidoc/blob/main/README.md)
Medical meadow wikidoc
The Medical Meadow Wikidoc dataset comprises question-answer pairs sourced from WikiDoc, an online platform where medical professionals collaboratively contribute and share contemporary medical knowledge. WikiDoc features two primary… See the full description on the dataset page: https://huggingface.co/datasets/Shekswess/llama2_medical_meadow_wikidoc_instruct_dataset.dynamic_sonnet_llama2
Dynamic Sonnet - Llama2
Curated dataset for benchmarking LLM serving systems
In real-world service scenarios, each request comes with varying input token lengths.
Some requests generate only a few tokens, while others produce a significant number.
Traditional fixed-length benchmarks fail to capture this variability, making it difficult to accurately assess real-world throughput performance.
This dynamic nature of input token lengths is crucial as it directly affects key features of… See the full description on the dataset page: https://huggingface.co/datasets/squeezebits/dynamic_sonnet_llama2.llama2-TR-recipeLlama2-7B-data-course
Dataset Description
This dataset is designed to support a teaching assistance model for an introductory computer science course. It includes structured content such as course syllabi, lesson plans, lecture materials, and exercises related to topics such as computer fundamentals, algorithms, hardware, software, and IT technologies. The dataset integrates practical assignments, theoretical knowledge, and ethical education, aiming to enhance teaching efficiency and improve student… See the full description on the dataset page: https://huggingface.co/datasets/2imi9/Llama2-7B-data-course.medical_llama2_instruct_datasetDataset made for instruction supervised finetuning of Llama 2 LLMs, by combining of medical datasets:
Medical meadow wikidoc (https://huggingface.co/datasets/medalpaca/medical_meadow_wikidoc/blob/main/README.md)
Medquad (https://www.kaggle.com/datasets/jpmiller/layoutlm)
Medical meadow wikidoc
The Medical Meadow Wikidoc dataset comprises question-answer pairs sourced from WikiDoc, an online platform where medical professionals collaboratively contribute and share contemporary… See the full description on the dataset page: https://huggingface.co/datasets/Shekswess/medical_llama2_instruct_dataset.llama2_medquad_instruct_datasetDataset made for instruction supervised finetuning of Llama 2 LLMs based on the Medquad dataset:
Medquad dataset (https://www.kaggle.com/datasets/jpmiller/layoutlm)
Medquad
MedQuAD is a comprehensive collection consisting of 47,457 medical question-answer pairs compiled from 12 authoritative sources within the National Institutes of Health (NIH), including domains like cancer.gov, niddk.nih.gov, GARD, and MedlinePlus Health Topics. These question-answer pairs span 37 distinct… See the full description on the dataset page: https://huggingface.co/datasets/Shekswess/llama2_medquad_instruct_dataset.Llama-2-7b-chat-finetune
plagas y enfermedades en el cultivo del tomate Dataset 1000
Dataset de 1000 instrucciones sobre la plagas y enfermedades en el cultivo del tomate.
Uso
from datasets import load_dataset
dataset = load_dataset("anyerg21/plagas-enfermedades-tomate-1000")
Estructura
instruction: Pregunta sobre el cultivo del tomate
input: Campo vacio
output: Respuesta
category: Categoria tematica
question_type: Tipo de pregunta
difficulty: Nivel de dificultad
Ejemplo… See the full description on the dataset page: https://huggingface.co/datasets/anyerg21/Llama-2-7b-chat-finetune.llama2-jinyong-style
LLama2微调医学数据集
开源数据集存在与预训练模型数据集相似度较高的问题,导致微调结果不明显,效果不突出,因此就需要构建某种风格的专属数据集。
因此,构建了带金庸创作风格的医学数据集,可直接用于LLama2的微调使用。
微调实验结果证明,当输入医学问题时,微调模型回答的风格为金庸创作风格回答。
Instruction
Instruction: 我睡眠不足,很头痛
Output
少侠保重身体。江湖人称我为医仙,今日得闻君有头痛之症,实非善兆。睡眠不足,乃是身体之不适,需慎之又慎。头痛之症,或许源自于血液循环不畅,或许源于神经压力过大。当务之急,当调整生活习惯,保持良好的睡眠规律,避免过度劳累。此外,可尝试调整饮食,避免辛辣刺激之物,以免加重头痛之苦。如君仍遭此病痛,可寻求名医良药,以求解忧。</s>
可以发现,模型的回答已经学习到金庸创作风格的能力。
ecommerce-faq-llama2-QADevanagari-Ecommerce-fomatted-for-llama2-chat-Dataset
Dataset Card for Dataset Name
यो देवनागरी नेपाली भाषाको डेटासेट विशेषगरी च्याटबोट प्रणालीहरू बनाउनको लागि डिजाइन गरिएको हो। यसमा विभिन्न श्रेणीहरूको डेटासेटहरू समावेश गरिएको छ, जसलाई JSON मा ढाँचा बनाईएको छ, जसले नेपाली वार्तालाप एआई अनुप्रयोगहरूको लागि भाषा मोडेलहरूलाई तालिम र फाइन-ट्यून गर्नको लागि व्यापक स्रोत प्रदान गर्दछ।
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/kshitizgajurel/Devanagari-Ecommerce-fomatted-for-llama2-chat-Dataset.medical_llama2_instruct_dataset_shortDataset made for instruction supervised finetuning of Llama 2 LLMs, by combining of medical datasets and getting 2k entries from them:
Medical meadow wikidoc (https://huggingface.co/datasets/medalpaca/medical_meadow_wikidoc/blob/main/README.md)
Medquad (https://www.kaggle.com/datasets/jpmiller/layoutlm)
Medical meadow wikidoc
The Medical Meadow Wikidoc dataset comprises question-answer pairs sourced from WikiDoc, an online platform where medical professionals collaboratively… See the full description on the dataset page: https://huggingface.co/datasets/Shekswess/medical_llama2_instruct_dataset_short.NvidiaDocumentationQandApairs-llama2InnerILLM-Llama2-training-dataset
Inner I LLM Llama 2 Training Dataset
Overview
This dataset is designed for fine-tuning the Llama 2 model to explore, express, and expand upon concepts related to the True Self, the Inner 'I', the Impersonal 'I', 'I Am', and the singularity of human intelligence. The dataset aims to foster a deeper understanding and reflection on these themes, contributing to the development of an LLM that can engage in meaningful dialogues about self-awareness and consciousness.… See the full description on the dataset page: https://huggingface.co/datasets/InnerI/InnerILLM-Llama2-training-dataset.MedicalQnA-llama2
MedicalQnA-llama2 Dataset
This repository contains the MedQuad-MedicalQnADataset specifically formatted for use with the LLaMA 2 prompt template. The dataset consists of medical questions categorized by question type, along with their corresponding answers. It is designed for text-to-text generation and text-generation tasks, particularly focusing on the medical domain.
Dataset Structure
The dataset is structured to include prompts that follow the LLaMA 2 template. Each… See the full description on the dataset page: https://huggingface.co/datasets/randomani/MedicalQnA-llama2.Llama-2-fine-tuning
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/BhaskarAgrawal/Llama-2-fine-tuning.ecommerce-faq-llama2-chatiruca-llama2-1k
iruca-llama2-1k: Lazy Llama 2 Formatting
This is a subset (1000 samples) of the iruca.ai example dataset, processed to match Llama 2's prompt format as described in this article.
guanaco-llama2-2knbme-llama2baggageitems_rules_llama2ecommerce-faq-llama2-QAllama2
