datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
llama2_indian_law_v1odia_master_data_llama2
Dataset Card for odia_master_data_llama2
Dataset Summary
This dataset is a mix of Odia instruction sets translated from open-source instruction sets and Odia domain knowledge instruction sets.
The Odia instruction sets used are:
odia_domain_context_train_v1
dolly-odia-15k
OdiEnCorp_translation_instructions_25k
gpt-teacher-roleplay-odia-3k
Odia_Alpaca_instructions_52k
hardcode_odia_qa_105
In this dataset Odia instruction, input, and output strings are available.… See the full description on the dataset page: https://huggingface.co/datasets/OdiaGenAI/odia_master_data_llama2.TexTrend-llama2
TextTrend Corpus: Exploring Linguistic Shifts and Semantic Patterns
Overview
The TextTrend Corpus is a unique dataset designed for fine-tuning language models. It consists of a diverse collection of text generated by AI over a span of approximately 19 hours, from 9 PM yesterday to 4 PM today. This dataset captures a snapshot of language evolution during this period, offering insights into linguistic trends and semantic shifts that can be explored and utilized for various… See the full description on the dataset page: https://huggingface.co/datasets/FinchResearch/TexTrend-llama2.sales-conversation-llama2enron_labeled_emails_with_subjects-llama2-7b_finetuningllama2_indian_law_v2Urgency-tone-topic-on-enron_labeled_emails_with_subjects-llama2-7b_finetuningllama2_legalturkish_agriculture_QA_llama2_22.6k
Dataset Card for Turkish_Agriculture_QA
This dataset contains question-answer pairs related to agriculture in Turkish. It has been translated and curated to fit the requirements for fine-tuning large language models such as LLaMA-2.
Dataset Details
Dataset Description
The Turkish_Agriculture_QA dataset contains a total of 22,615 question-answer pairs related to various aspects of agriculture. The dataset was originally curated from an English dataset and has… See the full description on the dataset page: https://huggingface.co/datasets/nieche/turkish_agriculture_QA_llama2_22.6k.Llama2TestingAmazonReviewsql-create-context-llama2-78kThis is dataset contain (78k samples) of the excellent b-mc2/sql-create-context and changed to derekiya/sql-create-context-llama2-78k dataset,
processed to match Llama 2's prompt format as described in this article.
Useful if you don't want to reformat it by yourself (e.g., using a script). It was designed for this article about fine-tuning a Llama 2 (chat)
enron_labeled_email-prompts-for-llama2_7bmath-llama2-20kWika-Llama2sentiment-analysis-llama2odia_context_10K_llama2_set
Dataset Card for odia_context_10k_llama2_set
Dataset Summary
This dataset contains 10K instructions that span various facets of Odisha's unique identity.
The instructions cover a wide array of subjects, ranging from the culinary delights in 'RECIPES,' the historical significance of 'HISTORICAL PLACES,' and
'TEMPLES OF ODISHA,' to the intellectual pursuits in 'ARITHMETIC,' 'HEALTH,' and 'GEOGRAPHY.'
It also explores the artistic tapestry of Odisha through 'ART AND… See the full description on the dataset page: https://huggingface.co/datasets/OdiaGenAI/odia_context_10K_llama2_set.llama_2_finetune_smallllama_2_finetunellama2-simulationbot-promptsmedtext-llama2Original data from:
https://huggingface.co/datasets/BI55/MedText
I just reformat it for fine tunning in lamma2 based on this article https://mlabonne.github.io/blog/posts/Fine_Tune_Your_Own_Llama_2_Model_in_a_Colab_Notebook.html
Another important point related to the data quality is the prompt template. Prompts are comprised of similar elements: system prompt (optional) to guide the model, user prompt (required) to give the instruction, additional inputs (optional) to take into consideration… See the full description on the dataset page: https://huggingface.co/datasets/ppdev/medtext-llama2.llama-2-7b-chat-hfMedQuad-MedicalQnADataset-Llama2-1k
MedQuad-1k: Llama 2 Formatting
This is a subset (1000 samples) of the keivalya/MedQuad-MedicalQnADataset dataset, processed to match Llama 2's prompt format as described in this article.
llama2_batterysparc-llama2math-llama2-1kLLAMA2medical-chat-llama2-datamath-llama2-8kllama_2_trainingllama2
