CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ibm-esa-geospatial /Llama3-SSL4EO-S12-v1.1-captions Llama3-SSL4EO-S12-Captions The captions are aligned with the SSL4EO-S12 v1.1 dataset and were automatically generated using the Llama3-LLaVA-Next-8B model. Please find more information regarding the generation and evaluation in the Llama3-MS-CLIP paper. Code: https://github.com/IBM/MS-CLIP Data Structure We provide the captions in two versions: As a single compressed Parquet file per split and as CSV files with 256 captions each that match the Zarr Zip files of the… See the full description on the dataset page: https://huggingface.co/datasets/ibm-esa-geospatial/Llama3-SSL4EO-S12-v1.1-captions.tabularzero-shot-image-classification100K<n<1M5 likes1.3k downloads1y agoHugging Face02CNX-PathLLM /Llama-slideQA-Sample-Featurestextn<1K0 likes746 downloads4mo agoHugging Face03mikheevshow /SIGNAL-Dataset-Hiddens-meta-llama_Meta-Llama-3-8B-Instructtextn<1K0 likes200 downloads11mo agoHugging Face04CreitinGameplays /DeepSeek-R1-Distill-Qwen-32B_NUMINA_train_amc_aime-llama3.1tabular1K<n<10K0 likes156 downloads2y agoHugging Face05jizzu /llama2_indian_law_v1text10K<n<100K6 likes111 downloads2y agoHugging Face06ikawrakow /winogrande-eval-for-llama.cppWinogrande evaluation dataset for llama.cpp tabular1K<n<10K1 likes101 downloads3y agoHugging Face07OdiaGenAI /odia_master_data_llama2 Dataset Card for odia_master_data_llama2 Dataset Summary This dataset is a mix of Odia instruction sets translated from open-source instruction sets and Odia domain knowledge instruction sets. The Odia instruction sets used are: odia_domain_context_train_v1 dolly-odia-15k OdiEnCorp_translation_instructions_25k gpt-teacher-roleplay-odia-3k Odia_Alpaca_instructions_52k hardcode_odia_qa_105 In this dataset Odia instruction, input, and output strings are available.… See the full description on the dataset page: https://huggingface.co/datasets/OdiaGenAI/odia_master_data_llama2.texttext-generation100K<n<1M1 likes83 downloads3y agoHugging Face08Phoenyx83 /Politifact-fake-news-6-categories-for-llama3-1 Dataset compiled for the article "LLaMA 3 vs. State-of-the-Art LLMs: Performance in Detecting Nuanced Fake News" based on Politifact Factcheck Data, available at https://www.kaggle.com/datasets/shivkumarganesh/politifact-factcheck-data language:" - en license: llama3.1 tabular10K<n<100K1 likes77 downloads2y agoHugging Face09wisenut-nlp-team /llama_nmt 중-한 번역 subset: ch-ko_basic_science length: 37.7k subset: ch-ko_broadcast length: 362k subset: ch-ko_daily_colloquial length: 600k subset: ch-ko_food length: 1.2M subset: ch-ko_humanities length: 33.4k subset: ch-ko_utterance_type length: 12k 영-한 번역 subset: en-ko_basic_science length: 356 subset: en-ko_broadcast length: 121k subset: en-ko_daily_colloquial length: 1.2M subset: en-ko_food length: 1.2M subset: en-ko_humanities length:… See the full description on the dataset page: https://huggingface.co/datasets/wisenut-nlp-team/llama_nmt.text10M<n<100M0 likes69 downloads2y agoHugging Face10FinchResearch /TexTrend-llama2 TextTrend Corpus: Exploring Linguistic Shifts and Semantic Patterns Overview The TextTrend Corpus is a unique dataset designed for fine-tuning language models. It consists of a diverse collection of text generated by AI over a span of approximately 19 hours, from 9 PM yesterday to 4 PM today. This dataset captures a snapshot of language evolution during this period, offering insights into linguistic trends and semantic shifts that can be explored and utilized for various… See the full description on the dataset page: https://huggingface.co/datasets/FinchResearch/TexTrend-llama2.text10K<n<100K0 likes68 downloads3y agoHugging Face11neuralmagic /quantized-llama-3.1-humaneval-evals Coding Benchmark Results The coding benchmark results were obtained with the EvalPlus library. HumanEvalpass@1 HumanEval+pass@1 meta-llama_Meta-Llama-3.1-405B-Instruct 67.3 67.5 neuralmagic_Meta-Llama-3.1-405B-Instruct-W8A8-FP8 66.7 66.6 neuralmagic_Meta-Llama-3.1-405B-Instruct-W4A16 66.5 66.4 neuralmagic_Meta-Llama-3.1-405B-Instruct-W8A8-INT8 64.3 64.8 neuralmagic_Meta-Llama-3.1-70B-Instruct-W8A8-FP8 58.1 57.7 neuralmagic_Meta-Llama-3.1-70B-Instruct-W4A16 57.1… See the full description on the dataset page: https://huggingface.co/datasets/neuralmagic/quantized-llama-3.1-humaneval-evals.text10K<n<100K0 likes68 downloads2y agoHugging Face12AmirLayegh /Tacred_Llamatext100K<n<1M0 likes67 downloads3y agoHugging Face13InHawK /sales-conversation-llama2textn<1K2 likes64 downloads3y agoHugging Face14CreitinGameplays /magpie-reasoning-v1-10k-step-by-step-rationale-alpaca-format-llama3.1text10K<n<100K1 likes61 downloads2y agoHugging Face15open-paws /continued-pretraining-llama-format Open Paws Continued Pretraining Llama Format Overview This dataset is part of the Open Paws initiative to develop AI training data aligned with animal liberation and advocacy principles. Created to train AI systems that understand and promote animal welfare, rights, and liberation. Dataset Details Dataset Type: Specialized Data Format: CSV (Comma-separated values) Languages: Multilingual (primarily English) Focus: Animal advocacy and ethical reasoning… See the full description on the dataset page: https://huggingface.co/datasets/open-paws/continued-pretraining-llama-format.texttext-generation10K<n<100K2 likes60 downloads1y agoHugging Face16SepKeyPro /amazon-bedrock-ug-llama3-8B-Instruct-1k Amazon Bedrock QandA Dataset for Llama3-8B-Instruct Fine-tuning This dataset includes 988 QandA extracted from Amazon Bedrock Documentation. It is then processed to match llama3-8B-Instruct template format. It can be used to fine-tune llama3 not hallucinating about Amazon Bedrock. Let's see what Llama3-8B says about Amazon Bedrock! ''' base_model = "meta-llama/Meta-Llama-3-8B-Instruct" tokenizer = AutoTokenizer.from_pretrained(base_model) pipe = pipeline(task="text-generation"… See the full description on the dataset page: https://huggingface.co/datasets/SepKeyPro/amazon-bedrock-ug-llama3-8B-Instruct-1k.textn<1K0 likes49 downloads2y agoHugging Face17Trelis /openassistant-llama-style Chat Fine-tuning Dataset - Llama 2 Style This dataset allows for fine-tuning chat models using [INST] AND [/INST] to wrap user messages. Preparation: The dataset is cloned from TimDettmers, which itself is a subset of the Open Assistant dataset, which you can find here. This subset of the data only contains the highest-rated paths in the conversation tree, with a total of 9,846 samples. The dataset was then filtered to: replace instances of '### Human:' with '[INST]' replace… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/openassistant-llama-style.text10K<n<100K9 likes47 downloads3y agoHugging Face18MedSalim /QA-deepseek-r1-distill-llama-70b DeepSeek-R1-LLama-70B Q&A Dataset This repository contains a curated set of 484 questions and answers generated by the DeepSeek-R1-LLama-70B model. The main goal is to evaluate the quality, coherence, and factual correctness of the model’s responses under various scenarios. Before getting excited about it, let's be realistic—large language models can produce both impressive and abysmal results. This dataset is meant to help you figure out which side of that spectrum… See the full description on the dataset page: https://huggingface.co/datasets/MedSalim/QA-deepseek-r1-distill-llama-70b.textn<1K1 likes47 downloads2y agoHugging Face19shanghong /llama_index_integration_datatext10M<n<100M0 likes47 downloads1y agoHugging Face20cheekymachine /enron_labeled_emails_with_subjects-llama2-7b_finetuningtexttext-classification1K<n<10K5 likes44 downloads3y agoHugging Face21mikheevshow /SIGNAL-Dataset-Hiddens-meta-llama_Meta-Llama-3-8Btextn<1K0 likes44 downloads11mo agoHugging Face22open-paws /conversational-finetuning-llama-format Open Paws Conversational Finetuning Llama Format This dataset is part of the Open Paws initiative to develop AI training data aligned with animal liberation and advocacy principles. Created to train AI systems that understand and promote animal welfare, rights, and liberation. Dataset Details Dataset Type: Training Data Format: CSV (Comma-separated values) Languages: Multilingual (primarily English) Focus: Animal advocacy and ethical reasoning Organization: Open Paws… See the full description on the dataset page: https://huggingface.co/datasets/open-paws/conversational-finetuning-llama-format.texttext-generation10K<n<100K2 likes40 downloads1y agoHugging Face23PhotonTJ /llama_1b_outputsimagen<1K0 likes37 downloads5mo agoHugging Face24OnlyCheeini /Llama-Discore-entextn<1K0 likes36 downloads3y agoHugging Face25jizzu /llama2_indian_law_v2text10K<n<100K7 likes35 downloads2y agoHugging Face26SarwarShafee /bl-conversation-dataset-for-llama3-finetune-v2textn<1K0 likes35 downloads2y agoHugging Face27Hmehdi515 /pmc_llamaFormat updated from axiong/pmc_llama_instructions text100K<n<1M0 likes35 downloads2y agoHugging Face28wisenut-nlp-team /llama_jp chat jmultiwoz (chat-pred) length: 3.54k real-persona-chat (chat-pred) length: 13.58k multiple Bactrian-X length: 67k databricks-dolly-15k-ja length: 15k guanaco_ja length: 100.63k llm-japanese-dataset-vanilla length: 2.52M OpenOrcaJapanese length: 573.62k qa AutoGeneratedJapaneseQA (open-qa) length: 93k JAQKET (closed-qa) length: 13.33k JaQuAD (closed-qa) length: 35.69k smr dialogsum-ja (chat-smr) length: 20.28k… See the full description on the dataset page: https://huggingface.co/datasets/wisenut-nlp-team/llama_jp.text1M<n<10M0 likes33 downloads2y agoHugging Face29abhirajeshbhai /movie-genre-llama-2text10K<n<100K0 likes32 downloads3y agoHugging Face30deepanshdj /ossat1_8k_llama3text1K<n<10K2 likes31 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.