datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Llama3-SSL4EO-S12-v1.1-captions
Llama3-SSL4EO-S12-Captions
The captions are aligned with the SSL4EO-S12 v1.1 dataset and were automatically generated using the Llama3-LLaVA-Next-8B model.
Please find more information regarding the generation and evaluation in the Llama3-MS-CLIP paper.
Code: https://github.com/IBM/MS-CLIP
Data Structure
We provide the captions in two versions: As a single compressed Parquet file per split and as CSV files with 256 captions each that match the Zarr Zip files of the… See the full description on the dataset page: https://huggingface.co/datasets/ibm-esa-geospatial/Llama3-SSL4EO-S12-v1.1-captions.Llama-slideQA-Sample-FeaturesSIGNAL-Dataset-Hiddens-meta-llama_Meta-Llama-3-8B-InstructDeepSeek-R1-Distill-Qwen-32B_NUMINA_train_amc_aime-llama3.1winogrande-eval-for-llama.cppWinogrande evaluation dataset for llama.cpp
odia_master_data_llama2
Dataset Card for odia_master_data_llama2
Dataset Summary
This dataset is a mix of Odia instruction sets translated from open-source instruction sets and Odia domain knowledge instruction sets.
The Odia instruction sets used are:
odia_domain_context_train_v1
dolly-odia-15k
OdiEnCorp_translation_instructions_25k
gpt-teacher-roleplay-odia-3k
Odia_Alpaca_instructions_52k
hardcode_odia_qa_105
In this dataset Odia instruction, input, and output strings are available.… See the full description on the dataset page: https://huggingface.co/datasets/OdiaGenAI/odia_master_data_llama2.Politifact-fake-news-6-categories-for-llama3-1
Dataset compiled for the article "LLaMA 3 vs. State-of-the-Art LLMs: Performance in Detecting Nuanced Fake News"
based on Politifact Factcheck Data, available at https://www.kaggle.com/datasets/shivkumarganesh/politifact-factcheck-data
language:"
- en
license: llama3.1
TexTrend-llama2
TextTrend Corpus: Exploring Linguistic Shifts and Semantic Patterns
Overview
The TextTrend Corpus is a unique dataset designed for fine-tuning language models. It consists of a diverse collection of text generated by AI over a span of approximately 19 hours, from 9 PM yesterday to 4 PM today. This dataset captures a snapshot of language evolution during this period, offering insights into linguistic trends and semantic shifts that can be explored and utilized for various… See the full description on the dataset page: https://huggingface.co/datasets/FinchResearch/TexTrend-llama2.quantized-llama-3.1-humaneval-evals
Coding Benchmark Results
The coding benchmark results were obtained with the EvalPlus library.
HumanEvalpass@1
HumanEval+pass@1
meta-llama_Meta-Llama-3.1-405B-Instruct
67.3
67.5
neuralmagic_Meta-Llama-3.1-405B-Instruct-W8A8-FP8
66.7
66.6
neuralmagic_Meta-Llama-3.1-405B-Instruct-W4A16
66.5
66.4
neuralmagic_Meta-Llama-3.1-405B-Instruct-W8A8-INT8
64.3
64.8
neuralmagic_Meta-Llama-3.1-70B-Instruct-W8A8-FP8
58.1
57.7
neuralmagic_Meta-Llama-3.1-70B-Instruct-W4A16
57.1… See the full description on the dataset page: https://huggingface.co/datasets/neuralmagic/quantized-llama-3.1-humaneval-evals.llama_nmt
중-한 번역
subset: ch-ko_basic_science
length: 37.7k
subset: ch-ko_broadcast
length: 362k
subset: ch-ko_daily_colloquial
length: 600k
subset: ch-ko_food
length: 1.2M
subset: ch-ko_humanities
length: 33.4k
subset: ch-ko_utterance_type
length: 12k
영-한 번역
subset: en-ko_basic_science
length: 356
subset: en-ko_broadcast
length: 121k
subset: en-ko_daily_colloquial
length: 1.2M
subset: en-ko_food
length: 1.2M
subset: en-ko_humanities
length:… See the full description on the dataset page: https://huggingface.co/datasets/wisenut-nlp-team/llama_nmt.sales-conversation-llama2continued-pretraining-llama-format
Open Paws Continued Pretraining Llama Format
Overview
This dataset is part of the Open Paws initiative to develop AI training data aligned with animal liberation and advocacy principles. Created to train AI systems that understand and promote animal welfare, rights, and liberation.
Dataset Details
Dataset Type: Specialized Data
Format: CSV (Comma-separated values)
Languages: Multilingual (primarily English)
Focus: Animal advocacy and ethical reasoning… See the full description on the dataset page: https://huggingface.co/datasets/open-paws/continued-pretraining-llama-format.pku-llama3.1-8b-answers-features-trainllama2_indian_law_v1magpie-reasoning-v1-10k-step-by-step-rationale-alpaca-format-llama3.1Tacred_Llamaamazon-bedrock-ug-llama3-8B-Instruct-1k
Amazon Bedrock QandA Dataset for Llama3-8B-Instruct Fine-tuning
This dataset includes 988 QandA extracted from Amazon Bedrock Documentation. It is then processed to match llama3-8B-Instruct template format.
It can be used to fine-tune llama3 not hallucinating about Amazon Bedrock. Let's see what Llama3-8B says about Amazon Bedrock!
''' base_model = "meta-llama/Meta-Llama-3-8B-Instruct" tokenizer = AutoTokenizer.from_pretrained(base_model) pipe = pipeline(task="text-generation"… See the full description on the dataset page: https://huggingface.co/datasets/SepKeyPro/amazon-bedrock-ug-llama3-8B-Instruct-1k.llama_index_integration_dataQA-deepseek-r1-distill-llama-70b
DeepSeek-R1-LLama-70B Q&A Dataset
This repository contains a curated set of 484 questions and answers generated by the DeepSeek-R1-LLama-70B model. The main goal is to evaluate the quality, coherence, and factual correctness of the model’s responses under various scenarios. Before getting excited about it, let's be realistic—large language models can produce both impressive and abysmal results. This dataset is meant to help you figure out which side of that spectrum… See the full description on the dataset page: https://huggingface.co/datasets/MedSalim/QA-deepseek-r1-distill-llama-70b.openassistant-llama-style
Chat Fine-tuning Dataset - Llama 2 Style
This dataset allows for fine-tuning chat models using [INST] AND [/INST] to wrap user messages.
Preparation:
The dataset is cloned from TimDettmers, which itself is a subset of the Open Assistant dataset, which you can find here. This subset of the data only contains the highest-rated paths in the conversation tree, with a total of 9,846 samples.
The dataset was then filtered to:
replace instances of '### Human:' with '[INST]'
replace… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/openassistant-llama-style.enron_labeled_emails_with_subjects-llama2-7b_finetuningSelective-Context-Llama3.1-8B-resultsllama_1b_outputsLLMLingua2-Llama3.1-8B-resultsbl-conversation-dataset-for-llama3-finetune-v2pmc_llamaFormat updated from axiong/pmc_llama_instructions
Llama-Discore-enmovie-genre-llama-2llama2_indian_law_v2public-health-QA-handouts-instruct-Llama-2k
