datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Llama3-SSL4EO-S12-v1.1-captions
Llama3-SSL4EO-S12-Captions
The captions are aligned with the SSL4EO-S12 v1.1 dataset and were automatically generated using the Llama3-LLaVA-Next-8B model.
Please find more information regarding the generation and evaluation in the Llama3-MS-CLIP paper.
Code: https://github.com/IBM/MS-CLIP
Data Structure
We provide the captions in two versions: As a single compressed Parquet file per split and as CSV files with 256 captions each that match the Zarr Zip files of the… See the full description on the dataset page: https://huggingface.co/datasets/ibm-esa-geospatial/Llama3-SSL4EO-S12-v1.1-captions.DeepSeek-R1-Distill-Qwen-32B_NUMINA_train_amc_aime-llama3.1magpie-reasoning-v1-10k-step-by-step-rationale-alpaca-format-llama3.1amazon-bedrock-ug-llama3-8B-Instruct-1k
Amazon Bedrock QandA Dataset for Llama3-8B-Instruct Fine-tuning
This dataset includes 988 QandA extracted from Amazon Bedrock Documentation. It is then processed to match llama3-8B-Instruct template format.
It can be used to fine-tune llama3 not hallucinating about Amazon Bedrock. Let's see what Llama3-8B says about Amazon Bedrock!
''' base_model = "meta-llama/Meta-Llama-3-8B-Instruct" tokenizer = AutoTokenizer.from_pretrained(base_model) pipe = pipeline(task="text-generation"… See the full description on the dataset page: https://huggingface.co/datasets/SepKeyPro/amazon-bedrock-ug-llama3-8B-Instruct-1k.Politifact-fake-news-6-categories-for-llama3-1
Dataset compiled for the article "LLaMA 3 vs. State-of-the-Art LLMs: Performance in Detecting Nuanced Fake News"
based on Politifact Factcheck Data, available at https://www.kaggle.com/datasets/shivkumarganesh/politifact-factcheck-data
language:"
- en
license: llama3.1
ossat1_8k_llama3bl-conversation-dataset-for-llama3-finetune-v2decoding_llama3Optimizer-llama370bgeneratedquestionLlama3_8b-emotion_multiclass-Plutchik
Description
This is a dataset for emotion classification of text sentences.
The dataset is a CSV file with 6,540 sentences. Each row has two columns: the first one has the sentence text, and the second one has its main emotion:
"text";"emotion"
The emotion can be one of Plutchik's eight emotion groups plus a neutral category. The sentence counts for each emotion are:
joy: 611 (9.34%)
sadness: 748 (11.44%)
trust: 735 (11.24%)
disgust: 838 (12.81%)
fear: 579 (8.85%)
anger: 743… See the full description on the dataset page: https://huggingface.co/datasets/uavster/Llama3_8b-emotion_multiclass-Plutchik.DPO_Llama3pku-llama3.1-8b-dataset-train-generationsllama3_dataset_1kpku-llama3.1-8b-dataset-test-generationsreasoning-0.01-content-llama3.1medical_llama3_instruct_datareasoning-base-20k-llama3.1llama-3.1-dataset-001energy_llama3.1-8B_multiple_batcheswater_dataset_llama3pku-llama3.1-8b-dataset-featuresLlama-3.1-8B-Instruct-evalllama3-8b-base
llama3-8b-base
Vietnamese labor-law raw document corpus prepared for continued pretraining.
Files
documents.csv
Columns
text
id
so_ky_hieu
Source
Local file: /home/thaivv/hehe/data/processed/labor_source_pack/core_relationship_cleaned_text_dataset_dict_fix/documents.csv
Rows: 3368
Notes
This dataset is document-level text.
so_ky_hieu is preserved as metadata for each document.
persian-alpaca-Llama3-70BAkka_Finetuning_Llama3.2train-llama3-lawscribeT2G-1k-Llama3.2-3B
T2G
Overview
T2G is a synthetic data consisting of text-graph pairs designed to finetune LLMs on information extraction tasks, specifically text-to-graph conversion.
Dataset Structure
The dataset is organized into the following main components:
Train Set: 800 instances for training models.
Validation Set: 100 for validating model performance.
Test Set: 100 instances for final evaluation.
Data Fields
Each instance in the dataset contains the… See the full description on the dataset page: https://huggingface.co/datasets/ESITime/T2G-1k-Llama3.2-3B.RankingSentences-NLI-LLaMA3-8B-32
RankingSentences-NLI-LLaMA3-8B-32
Dataset Details
RankingSentences-NLI-LLaMA3-8B-32 is a dataset crafted by LLaMA3-8B-Instruct. Its distinctive feature lies in organizing sentences within the semantic space according to their semantic order. For further details, please refer to our paper (https://arxiv.org/pdf/2502.13656) and code (https://github.com/hly1998/RankingSentenceGeneration).
Llama_3_70b_biologyRaiden-DeepSeek-R1-llama3.1-v1
