CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ibm-esa-geospatial /Llama3-SSL4EO-S12-v1.1-captions Llama3-SSL4EO-S12-Captions The captions are aligned with the SSL4EO-S12 v1.1 dataset and were automatically generated using the Llama3-LLaVA-Next-8B model. Please find more information regarding the generation and evaluation in the Llama3-MS-CLIP paper. Code: https://github.com/IBM/MS-CLIP Data Structure We provide the captions in two versions: As a single compressed Parquet file per split and as CSV files with 256 captions each that match the Zarr Zip files of the… See the full description on the dataset page: https://huggingface.co/datasets/ibm-esa-geospatial/Llama3-SSL4EO-S12-v1.1-captions.tabularzero-shot-image-classification100K<n<1M5 likes1.3k downloads1y agoHugging Face02CNX-PathLLM /Llama-slideQA-Sample-Featurestextn<1K0 likes746 downloads4mo agoHugging Face03mikheevshow /SIGNAL-Dataset-Hiddens-meta-llama_Meta-Llama-3-8B-Instructtextn<1K0 likes193 downloads11mo agoHugging Face04CreitinGameplays /DeepSeek-R1-Distill-Qwen-32B_NUMINA_train_amc_aime-llama3.1tabular1K<n<10K0 likes153 downloads2y agoHugging Face05ikawrakow /winogrande-eval-for-llama.cppWinogrande evaluation dataset for llama.cpp tabular1K<n<10K1 likes103 downloads3y agoHugging Face06OdiaGenAI /odia_master_data_llama2 Dataset Card for odia_master_data_llama2 Dataset Summary This dataset is a mix of Odia instruction sets translated from open-source instruction sets and Odia domain knowledge instruction sets. The Odia instruction sets used are: odia_domain_context_train_v1 dolly-odia-15k OdiEnCorp_translation_instructions_25k gpt-teacher-roleplay-odia-3k Odia_Alpaca_instructions_52k hardcode_odia_qa_105 In this dataset Odia instruction, input, and output strings are available.… See the full description on the dataset page: https://huggingface.co/datasets/OdiaGenAI/odia_master_data_llama2.texttext-generation100K<n<1M1 likes83 downloads3y agoHugging Face07Phoenyx83 /Politifact-fake-news-6-categories-for-llama3-1 Dataset compiled for the article "LLaMA 3 vs. State-of-the-Art LLMs: Performance in Detecting Nuanced Fake News" based on Politifact Factcheck Data, available at https://www.kaggle.com/datasets/shivkumarganesh/politifact-factcheck-data language:" - en license: llama3.1 tabular10K<n<100K1 likes76 downloads2y agoHugging Face08FinchResearch /TexTrend-llama2 TextTrend Corpus: Exploring Linguistic Shifts and Semantic Patterns Overview The TextTrend Corpus is a unique dataset designed for fine-tuning language models. It consists of a diverse collection of text generated by AI over a span of approximately 19 hours, from 9 PM yesterday to 4 PM today. This dataset captures a snapshot of language evolution during this period, offering insights into linguistic trends and semantic shifts that can be explored and utilized for various… See the full description on the dataset page: https://huggingface.co/datasets/FinchResearch/TexTrend-llama2.text10K<n<100K0 likes72 downloads3y agoHugging Face09neuralmagic /quantized-llama-3.1-humaneval-evals Coding Benchmark Results The coding benchmark results were obtained with the EvalPlus library. HumanEvalpass@1 HumanEval+pass@1 meta-llama_Meta-Llama-3.1-405B-Instruct 67.3 67.5 neuralmagic_Meta-Llama-3.1-405B-Instruct-W8A8-FP8 66.7 66.6 neuralmagic_Meta-Llama-3.1-405B-Instruct-W4A16 66.5 66.4 neuralmagic_Meta-Llama-3.1-405B-Instruct-W8A8-INT8 64.3 64.8 neuralmagic_Meta-Llama-3.1-70B-Instruct-W8A8-FP8 58.1 57.7 neuralmagic_Meta-Llama-3.1-70B-Instruct-W4A16 57.1… See the full description on the dataset page: https://huggingface.co/datasets/neuralmagic/quantized-llama-3.1-humaneval-evals.text10K<n<100K0 likes68 downloads2y agoHugging Face10wisenut-nlp-team /llama_nmt 중-한 번역 subset: ch-ko_basic_science length: 37.7k subset: ch-ko_broadcast length: 362k subset: ch-ko_daily_colloquial length: 600k subset: ch-ko_food length: 1.2M subset: ch-ko_humanities length: 33.4k subset: ch-ko_utterance_type length: 12k 영-한 번역 subset: en-ko_basic_science length: 356 subset: en-ko_broadcast length: 121k subset: en-ko_daily_colloquial length: 1.2M subset: en-ko_food length: 1.2M subset: en-ko_humanities length:… See the full description on the dataset page: https://huggingface.co/datasets/wisenut-nlp-team/llama_nmt.text10M<n<100M0 likes65 downloads2y agoHugging Face11InHawK /sales-conversation-llama2textn<1K2 likes63 downloads3y agoHugging Face12open-paws /continued-pretraining-llama-format Open Paws Continued Pretraining Llama Format Overview This dataset is part of the Open Paws initiative to develop AI training data aligned with animal liberation and advocacy principles. Created to train AI systems that understand and promote animal welfare, rights, and liberation. Dataset Details Dataset Type: Specialized Data Format: CSV (Comma-separated values) Languages: Multilingual (primarily English) Focus: Animal advocacy and ethical reasoning… See the full description on the dataset page: https://huggingface.co/datasets/open-paws/continued-pretraining-llama-format.texttext-generation10K<n<100K2 likes63 downloads1y agoHugging Face13luckeciano /pku-llama3.1-8b-answers-features-traintabular1M<n<10M0 likes62 downloads2y agoHugging Face14jizzu /llama2_indian_law_v1text10K<n<100K6 likes61 downloads2y agoHugging Face15CreitinGameplays /magpie-reasoning-v1-10k-step-by-step-rationale-alpaca-format-llama3.1text10K<n<100K1 likes61 downloads2y agoHugging Face16AmirLayegh /Tacred_Llamatext100K<n<1M0 likes60 downloads3y agoHugging Face17SepKeyPro /amazon-bedrock-ug-llama3-8B-Instruct-1k Amazon Bedrock QandA Dataset for Llama3-8B-Instruct Fine-tuning This dataset includes 988 QandA extracted from Amazon Bedrock Documentation. It is then processed to match llama3-8B-Instruct template format. It can be used to fine-tune llama3 not hallucinating about Amazon Bedrock. Let's see what Llama3-8B says about Amazon Bedrock! ''' base_model = "meta-llama/Meta-Llama-3-8B-Instruct" tokenizer = AutoTokenizer.from_pretrained(base_model) pipe = pipeline(task="text-generation"… See the full description on the dataset page: https://huggingface.co/datasets/SepKeyPro/amazon-bedrock-ug-llama3-8B-Instruct-1k.textn<1K0 likes49 downloads2y agoHugging Face18shanghong /llama_index_integration_datatext10M<n<100M0 likes49 downloads1y agoHugging Face19MedSalim /QA-deepseek-r1-distill-llama-70b DeepSeek-R1-LLama-70B Q&A Dataset This repository contains a curated set of 484 questions and answers generated by the DeepSeek-R1-LLama-70B model. The main goal is to evaluate the quality, coherence, and factual correctness of the model’s responses under various scenarios. Before getting excited about it, let's be realistic—large language models can produce both impressive and abysmal results. This dataset is meant to help you figure out which side of that spectrum… See the full description on the dataset page: https://huggingface.co/datasets/MedSalim/QA-deepseek-r1-distill-llama-70b.textn<1K1 likes46 downloads2y agoHugging Face20Trelis /openassistant-llama-style Chat Fine-tuning Dataset - Llama 2 Style This dataset allows for fine-tuning chat models using [INST] AND [/INST] to wrap user messages. Preparation: The dataset is cloned from TimDettmers, which itself is a subset of the Open Assistant dataset, which you can find here. This subset of the data only contains the highest-rated paths in the conversation tree, with a total of 9,846 samples. The dataset was then filtered to: replace instances of '### Human:' with '[INST]' replace… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/openassistant-llama-style.text10K<n<100K9 likes45 downloads3y agoHugging Face21cheekymachine /enron_labeled_emails_with_subjects-llama2-7b_finetuningtexttext-classification1K<n<10K5 likes44 downloads3y agoHugging Face22Itz-Amethyst /Selective-Context-Llama3.1-8B-resultstabular10K<n<100K0 likes38 downloads3mo agoHugging Face23PhotonTJ /llama_1b_outputsimagen<1K0 likes36 downloads5mo agoHugging Face24Itz-Amethyst /LLMLingua2-Llama3.1-8B-resultstabular10K<n<100K0 likes35 downloads3mo agoHugging Face25SarwarShafee /bl-conversation-dataset-for-llama3-finetune-v2textn<1K0 likes34 downloads2y agoHugging Face26Hmehdi515 /pmc_llamaFormat updated from axiong/pmc_llama_instructions text100K<n<1M0 likes34 downloads2y agoHugging Face27OnlyCheeini /Llama-Discore-entextn<1K0 likes33 downloads3y agoHugging Face28abhirajeshbhai /movie-genre-llama-2text10K<n<100K0 likes32 downloads3y agoHugging Face29jizzu /llama2_indian_law_v2text10K<n<100K7 likes32 downloads2y agoHugging Face30sambanankhu /public-health-QA-handouts-instruct-Llama-2ktext1K<n<10K1 likes31 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.