CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ibm-esa-geospatial /Llama3-SSL4EO-S12-v1.1-captions Llama3-SSL4EO-S12-Captions The captions are aligned with the SSL4EO-S12 v1.1 dataset and were automatically generated using the Llama3-LLaVA-Next-8B model. Please find more information regarding the generation and evaluation in the Llama3-MS-CLIP paper. Code: https://github.com/IBM/MS-CLIP Data Structure We provide the captions in two versions: As a single compressed Parquet file per split and as CSV files with 256 captions each that match the Zarr Zip files of the… See the full description on the dataset page: https://huggingface.co/datasets/ibm-esa-geospatial/Llama3-SSL4EO-S12-v1.1-captions.tabularzero-shot-image-classification100K<n<1M5 likes1.4k downloads1y agoHugging Face02CreitinGameplays /DeepSeek-R1-Distill-Qwen-32B_NUMINA_train_amc_aime-llama3.1tabular1K<n<10K0 likes167 downloads2y agoHugging Face03CreitinGameplays /magpie-reasoning-v1-10k-step-by-step-rationale-alpaca-format-llama3.1text10K<n<100K2 likes63 downloads2y agoHugging Face04SepKeyPro /amazon-bedrock-ug-llama3-8B-Instruct-1k Amazon Bedrock QandA Dataset for Llama3-8B-Instruct Fine-tuning This dataset includes 988 QandA extracted from Amazon Bedrock Documentation. It is then processed to match llama3-8B-Instruct template format. It can be used to fine-tune llama3 not hallucinating about Amazon Bedrock. Let's see what Llama3-8B says about Amazon Bedrock! ''' base_model = "meta-llama/Meta-Llama-3-8B-Instruct" tokenizer = AutoTokenizer.from_pretrained(base_model) pipe = pipeline(task="text-generation"… See the full description on the dataset page: https://huggingface.co/datasets/SepKeyPro/amazon-bedrock-ug-llama3-8B-Instruct-1k.textn<1K0 likes49 downloads2y agoHugging Face05Phoenyx83 /Politifact-fake-news-6-categories-for-llama3-1 Dataset compiled for the article "LLaMA 3 vs. State-of-the-Art LLMs: Performance in Detecting Nuanced Fake News" based on Politifact Factcheck Data, available at https://www.kaggle.com/datasets/shivkumarganesh/politifact-factcheck-data language:" - en license: llama3.1 tabular10K<n<100K1 likes41 downloads2y agoHugging Face06deepanshdj /ossat1_8k_llama3text1K<n<10K2 likes32 downloads2y agoHugging Face07SarwarShafee /bl-conversation-dataset-for-llama3-finetune-v2textn<1K0 likes32 downloads2y agoHugging Face08yoonLM /decoding_llama3text100K<n<1M0 likes25 downloads2y agoHugging Face09AIAT /Optimizer-llama370bgeneratedquestiontext1K<n<10K0 likes24 downloads2y agoHugging Face10uavster /Llama3_8b-emotion_multiclass-Plutchik Description This is a dataset for emotion classification of text sentences. The dataset is a CSV file with 6,540 sentences. Each row has two columns: the first one has the sentence text, and the second one has its main emotion: "text";"emotion" The emotion can be one of Plutchik's eight emotion groups plus a neutral category. The sentence counts for each emotion are: joy: 611 (9.34%) sadness: 748 (11.44%) trust: 735 (11.24%) disgust: 838 (12.81%) fear: 579 (8.85%) anger: 743… See the full description on the dataset page: https://huggingface.co/datasets/uavster/Llama3_8b-emotion_multiclass-Plutchik.texttext-classification1K<n<10K3 likes23 downloads2y agoHugging Face11vinven7 /DPO_Llama3text10K<n<100K0 likes19 downloads2y agoHugging Face12luckeciano /pku-llama3.1-8b-dataset-train-generationstabular1M<n<10M0 likes19 downloads2y agoHugging Face13deepanshdj /llama3_dataset_1ktext1K<n<10K0 likes18 downloads2y agoHugging Face14luckeciano /pku-llama3.1-8b-dataset-test-generationstext1M<n<10M0 likes17 downloads2y agoHugging Face15CreitinGameplays /reasoning-0.01-content-llama3.1text10K<n<100K0 likes16 downloads2y agoHugging Face16SURESHBEEKHANI /medical_llama3_instruct_datatext10K<n<100K0 likes15 downloads2y agoHugging Face17CreitinGameplays /reasoning-base-20k-llama3.1text10K<n<100K0 likes15 downloads2y agoHugging Face18orionai /llama-3.1-dataset-001textn<1K0 likes13 downloads2y agoHugging Face19Anon152425 /energy_llama3.1-8B_multiple_batchestabular10K<n<100K0 likes13 downloads1y agoHugging Face20parthrautV /water_dataset_llama3text10K<n<100K0 likes12 downloads2y agoHugging Face21luckeciano /pku-llama3.1-8b-dataset-featurestabular10K<n<100K0 likes12 downloads2y agoHugging Face22crosslingual-em /Llama-3.1-8B-Instruct-evaldocumentn<1K0 likes12 downloads5mo agoHugging Face23vinhguy /llama3-8b-base llama3-8b-base Vietnamese labor-law raw document corpus prepared for continued pretraining. Files documents.csv Columns text id so_ky_hieu Source Local file: /home/thaivv/hehe/data/processed/labor_source_pack/core_relationship_cleaned_text_dataset_dict_fix/documents.csv Rows: 3368 Notes This dataset is document-level text. so_ky_hieu is preserved as metadata for each document. texttext-generation1K<n<10K0 likes12 downloads4mo agoHugging Face24myrkur /persian-alpaca-Llama3-70Btext10K<n<100K2 likes11 downloads2y agoHugging Face25lorixmassello /Akka_Finetuning_Llama3.2textquestion-answeringn<1K0 likes11 downloads2y agoHugging Face26saruo06 /train-llama3-lawscribetextn<1K0 likes9 downloads2y agoHugging Face27ESITime /T2G-1k-Llama3.2-3B T2G Overview T2G is a synthetic data consisting of text-graph pairs designed to finetune LLMs on information extraction tasks, specifically text-to-graph conversion. Dataset Structure The dataset is organized into the following main components: Train Set: 800 instances for training models. Validation Set: 100 for validating model performance. Test Set: 100 instances for final evaluation. Data Fields Each instance in the dataset contains the… See the full description on the dataset page: https://huggingface.co/datasets/ESITime/T2G-1k-Llama3.2-3B.text1K<n<10K0 likes9 downloads2y agoHugging Face28leoner24 /RankingSentences-NLI-LLaMA3-8B-32 RankingSentences-NLI-LLaMA3-8B-32 Dataset Details RankingSentences-NLI-LLaMA3-8B-32 is a dataset crafted by LLaMA3-8B-Instruct. Its distinctive feature lies in organizing sentences within the semantic space according to their semantic order. For further details, please refer to our paper (https://arxiv.org/pdf/2502.13656) and code (https://github.com/hly1998/RankingSentenceGeneration). texttext-classification10K<n<100K1 likes9 downloads2y agoHugging Face29tourist800 /Llama_3_70b_biologytext1K<n<10K0 likes8 downloads2y agoHugging Face30CreitinGameplays /Raiden-DeepSeek-R1-llama3.1-v1text10K<n<100K0 likes8 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.