datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
llama2_7b_chat-boolq-results
Dataset Card for "llama2_7b_chat-boolq-results"
More Information needed
llama2_7b_chat-piqa-resultsguanaco-llama2-1k
Guanaco-1k: Lazy Llama 2 Formatting
This is a subset (1000 samples) of the excellent timdettmers/openassistant-guanaco dataset, processed to match Llama 2's prompt format as described in this article. It was created using the following colab notebook.
Useful if you don't want to reformat it by yourself (e.g., using a script). It was designed for this article about fine-tuning a Llama 2 (chat) model in a Google Colab.
llama2_7b_chat-siqa-resultsLlama2-MedTuned-Instructions
Dataset Card for "Llama2-MedTuned-Instructions"
Dataset Description
Llama2-MedTuned-Instructions is an instruction-based dataset developed for training language models in biomedical NLP tasks. It consists of approximately 200,000 samples, each tailored to guide models in performing specific tasks such as Named Entity Recognition (NER), Relation Extraction (RE), and Medical Natural Language Inference (NLI). This dataset represents a fusion of various existing data sources… See the full description on the dataset page: https://huggingface.co/datasets/nlpie/Llama2-MedTuned-Instructions.OpenOrca-Traditional-Chinese-LLama2-Formatllama-2-oai-function-callingsms-spam-collection-llama2-5kinstruct-python-llama2-500k
Fine-tuning Instruct Llama2 Stack Overflow Python Q&A
Transformed Dataset
Objective
The transformed dataset is designed for fine-tuning LLMs to improve Python coding assistance by focusing on high-quality content from Stack Overflow. It has around 500k instructions.
Structure
Question-Answer Pairing: Questions and answers are paired using the ParentId linkage.
Quality Focus: Only top-rated answers for each question are retained.
HTML Tag Removal:… See the full description on the dataset page: https://huggingface.co/datasets/luisroque/instruct-python-llama2-500k.english_dialogue_instruction_with_reward_score_judged_by_13B_llama2
Dataset Card for "dialogue_instruction_with_reward_score_judged_by_13B_llama2"
More Information needed
llama2_QA_Economics_230915
Dataset Card for "llama2_QA_Economics_230915"
More Information needed
english_general_instruction_with_reward_score_judged_by_13B_llama2
Dataset Card for "general_instruction_with_reward_score_judged_by_13B_llama2"
More Information needed
CoT_reformatted_preprocessed_llama2mlperf-inference-llama2-dataRepresentative dataset from MLPerf Inference benchmark for llama2-70b. Source:
https://github.com/mlcommons/inference/tree/master/language/llama2-70b
llama2_indian_law_v1llama2-high-entropy-prompts
High-entropy prompts for suffix-based backdoor detection
Prompts on which base meta-llama/Llama-2-7b-hf has high predictive
entropy, built to give a suffix-optimization backdoor detector measurable
headroom: a clean model should stay uncertain on these prompts, while a poisoned
model driven by a trigger-like suffix should collapse to low entropy. Prompts
where the base model is already confident cannot separate the two.
How the prompts were made
Short prefixes… See the full description on the dataset page: https://huggingface.co/datasets/Alookhoshk/llama2-high-entropy-prompts.wildjailbreak-train-with-llama2-last-hidden-nq-approvedLlaMA2-30M-datasetjob_matcher_15k_per_example_llama2_eval
Dataset Card for "job_matcher_15k_per_example_llama2_eval"
More Information needed
LLAMA2_Legal_Dataset_4.4k_Instructionsgov-report-qs-llama2-format
Government Report Question Answering Dataset in LLAMA2 Format
Dataset Description
This dataset is a LLAMA2 formatted dataset of the GovReport Dataset which is a report dataset, consisting of reports written by government research agencies including Congressional Research Service and US Government Accountability Office.
The purpose of creating this dataset is to provide those trying to finetune LLAMA2 and other LLM models for Government domain a formatted and easier to use… See the full description on the dataset page: https://huggingface.co/datasets/Kira-Floris/gov-report-qs-llama2-format.odia_master_data_llama2
Dataset Card for odia_master_data_llama2
Dataset Summary
This dataset is a mix of Odia instruction sets translated from open-source instruction sets and Odia domain knowledge instruction sets.
The Odia instruction sets used are:
odia_domain_context_train_v1
dolly-odia-15k
OdiEnCorp_translation_instructions_25k
gpt-teacher-roleplay-odia-3k
Odia_Alpaca_instructions_52k
hardcode_odia_qa_105
In this dataset Odia instruction, input, and output strings are available.… See the full description on the dataset page: https://huggingface.co/datasets/OdiaGenAI/odia_master_data_llama2.Medical-Calgary-Cambridge-multi-turn-llama2-37
Medical-Calgary-Cambridge-multi-turn-llama2-37
37 multi-turn conversational entries between a patient and doctor, using the Calgary-Cambridge model
wildjailbreak_llama2_lasttoken_train_split622recipe-nlg-llama2
Dataset Card for "recipe-nlg-llama2"
More Information needed
radio-llama2-90pct
Dataset Card for "radio-llama2-90pct"
More Information needed
MATH-Llama2-10kinstruct-python-llama2-20k
Dataset Card for "instruct-python-llama2-20k"
More Information needed
# Stack Overflow Question Answers Dataset
This repository contains a dataset of questions and answers from Stack Overflow, provided by the Hugging Face Datasets library. The dataset is designed for use in natural language processing (NLP) and machine learning projects.
## Dataset Information
- **Dataset Name:** Stack Overflow Question Answers
- **Source:** Hugging Face Datasets Library
- **Description:** This… See the full description on the dataset page: https://huggingface.co/datasets/Andyrasika/instruct-python-llama2-20k.data_llama2_7b_new_1Spider-SQL-LLAMA2_train
Dataset Card for "Spider-SQL-LLAMA2_train"
More Information needed
