CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01donghyunli /Llama-2-7b-KronQ-HG Llama-2-7b — KronQ H_G (output-side gradient covariance) Paper: arXiv:2607.07964 · Code: GitHub Pre-computed H_G for Llama-2-7b, the output-side curvature factor used by KronQ under the K-FAC factorization H ≈ H_X ⊗ H_G. H_G is the per-sublayer sampled-Fisher gradient covariance (labels drawn from the model distribution) (E[g gᵀ] over the layer output), distinct from the standard input-side Hessian H_X (which GPTQ/GPTAQ build online during calibration). Publishing this lets you… See the full description on the dataset page: https://huggingface.co/datasets/donghyunli/Llama-2-7b-KronQ-HG.text-generation0 likes2.5k downloads2mo agoHugging Face02KristianS7 /prepacked-fineweb-edu-llama2-32K-T2048 prepacked-fineweb-edu-llama2-32K-T2048 Pre-tokenized and BOS-aligned best-fit packed version of FineWeb-Edu for training with looped nanochat. Tokenized with the Llama 2 tokenizer (32,000 base vocab + 8 special tokens = 32,008). Stats Train split Source karpathy/fineweb-edu-100b-shuffle (1,821 shards) Total tokens 63.26B Total docs 97.1M Rows 30,873,598 Shards 2,059 (train-00000 to train-02058) Rows per shard ~15,000… See the full description on the dataset page: https://huggingface.co/datasets/KristianS7/prepacked-fineweb-edu-llama2-32K-T2048.text-generation0 likes1.6k downloads6mo agoHugging Face03donghyunli /Llama-2-70b-KronQ-HG Llama-2-70b-hf — KronQ H_G (output-side gradient covariance) Paper: arXiv:2607.07964 · Code: GitHub Pre-computed H_G for Llama-2-70b-hf, the output-side curvature factor used by KronQ under the K-FAC factorization H ≈ H_X ⊗ H_G. Per-sublayer empirical-Fisher gradient covariance (E[g gᵀ] over the layer output), distinct from the input-side Hessian H_X. Publishing this lets you reproduce KronQ quantization without the offline Fisher precompute step. Contents (80 layers… See the full description on the dataset page: https://huggingface.co/datasets/donghyunli/Llama-2-70b-KronQ-HG.text-generation0 likes1.5k downloads2mo agoHugging Face04donghyunli /Llama-2-13b-KronQ-HG Llama-2-13b — KronQ H_G (output-side gradient covariance) Paper: arXiv:2607.07964 · Code: GitHub Pre-computed H_G for Llama-2-13b, the output-side curvature factor used by KronQ under the K-FAC factorization H ≈ H_X ⊗ H_G. H_G is the per-sublayer empirical-Fisher gradient covariance (E[g gᵀ] over the layer output), distinct from the standard input-side Hessian H_X (built online during calibration). Publishing this lets you reproduce KronQ quantization without the offline Fisher… See the full description on the dataset page: https://huggingface.co/datasets/donghyunli/Llama-2-13b-KronQ-HG.text-generation0 likes1.1k downloads2mo agoHugging Face05luisroque /instruct-python-llama2-500k Fine-tuning Instruct Llama2 Stack Overflow Python Q&A Transformed Dataset Objective The transformed dataset is designed for fine-tuning LLMs to improve Python coding assistance by focusing on high-quality content from Stack Overflow. It has around 500k instructions. Structure Question-Answer Pairing: Questions and answers are paired using the ParentId linkage. Quality Focus: Only top-rated answers for each question are retained. HTML Tag Removal:… See the full description on the dataset page: https://huggingface.co/datasets/luisroque/instruct-python-llama2-500k.texttext-generation100K<n<1M4 likes131 downloads3y agoHugging Face06tim9510019 /llama2_QA_Economics_230915 Dataset Card for "llama2_QA_Economics_230915" More Information needed tabularquestion-answering1K<n<10K13 likes129 downloads2y agoHugging Face07emozilla /dolma-v1_7-30B-tokenized-llama2-nanosetTokenized (Llama 2) verison of emozilla/dolma-v1_7-30B as a Nanotron dataset split into 10 GB chunks. To download: huggingface-cli download --repo-type dataset --local-dir dolma-v1_7-30B-tokenized-llama2-nanoset --local-dir-use-symlinks False emozilla/dolma-v1_7-30B-tokenized-llama2-nanoset To recombine: cat dolma-v1_7-30B-tokenized-llama2-nanoset/dolma-v1_7-30B-tokenized-llama2-nanoset_input_ids.npy.* > dolma-v1_7-30B-tokenized-llama2-nanoset.npy rm -rf… See the full description on the dataset page: https://huggingface.co/datasets/emozilla/dolma-v1_7-30B-tokenized-llama2-nanoset.text-generation100B<n<1T1 likes121 downloads2y agoHugging Face08Alookhoshk /llama2-high-entropy-prompts High-entropy prompts for suffix-based backdoor detection Prompts on which base meta-llama/Llama-2-7b-hf has high predictive entropy, built to give a suffix-optimization backdoor detector measurable headroom: a clean model should stay uncertain on these prompts, while a poisoned model driven by a trigger-like suffix should collapse to low entropy. Prompts where the base model is already confident cannot separate the two. How the prompts were made Short prefixes… See the full description on the dataset page: https://huggingface.co/datasets/Alookhoshk/llama2-high-entropy-prompts.tabulartext-generation1K<n<10K0 likes109 downloads17d agoHugging Face09jsun /Prolong_64K_v2_Llama2_Tokenizer Prolong_64K_v2_Llama2_Tokenizer This is the Prolong_64K dataset, tokenized using the Llama-2-7b-hf tokenizer for use in Samba-style training. This dataset was used in the research paper: Rethinking Language Model Scaling under Transferable Hypersphere Optimization. The official training codebase can be found at GitHub - microsoft/ArchScale. Download 👉 You can download and unzip the dataset from: prolong_64K_v2.zip wget -c… See the full description on the dataset page: https://huggingface.co/datasets/jsun/Prolong_64K_v2_Llama2_Tokenizer.text-generation3 likes101 downloads6mo agoHugging Face10OdiaGenAI /odia_master_data_llama2 Dataset Card for odia_master_data_llama2 Dataset Summary This dataset is a mix of Odia instruction sets translated from open-source instruction sets and Odia domain knowledge instruction sets. The Odia instruction sets used are: odia_domain_context_train_v1 dolly-odia-15k OdiEnCorp_translation_instructions_25k gpt-teacher-roleplay-odia-3k Odia_Alpaca_instructions_52k hardcode_odia_qa_105 In this dataset Odia instruction, input, and output strings are available.… See the full description on the dataset page: https://huggingface.co/datasets/OdiaGenAI/odia_master_data_llama2.texttext-generation100K<n<1M1 likes83 downloads3y agoHugging Face11baoanhtran /guanaco-llama2-200 CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages \text-generationn<1K4 likes67 downloads3y agoHugging Face12squeezebits /dynamic_sonnet_llama2 Dynamic Sonnet - Llama2 Curated dataset for benchmarking LLM serving systems In real-world service scenarios, each request comes with varying input token lengths. Some requests generate only a few tokens, while others produce a significant number. Traditional fixed-length benchmarks fail to capture this variability, making it difficult to accurately assess real-world throughput performance. This dynamic nature of input token lengths is crucial as it directly affects key features of… See the full description on the dataset page: https://huggingface.co/datasets/squeezebits/dynamic_sonnet_llama2.textquestion-answering1K<n<10K1 likes59 downloads2y agoHugging Face13ravikumarmn /guanaco-llama2texttext-generation10K<n<100K3 likes35 downloads2y agoHugging Face14jeqcho /llama-2-13b-outputs Llama 2 13B Model Outputs This repository contains all concatenated text outputs from Llama 2 13B models (base and chat) for MT-Bench and AlpacaEval benchmarks. Quick Download Download the output files directly: # Base model outputs (885 completions: 80 MT-Bench + 805 AlpacaEval) wget https://huggingface.co/datasets/jeqcho/llama-2-13b-outputs/resolve/main/llama-2-13b-base.txt # Chat model outputs (885 completions: 80 MT-Bench + 805 AlpacaEval) wget… See the full description on the dataset page: https://huggingface.co/datasets/jeqcho/llama-2-13b-outputs.texttext-generation10K<n<100K0 likes35 downloads11mo agoHugging Face15luisroque /instruct-python-llama2-20k Fine-tuning Instruct Llama2 Stack Overflow Python Q&A Transformed Dataset Objective The transformed dataset is designed for fine-tuning LLMs to improve Python coding assistance by focusing on high-quality content from Stack Overflow. It has around 20k instructions. Structure Question-Answer Pairing: Questions and answers are paired using the ParentId linkage. Quality Focus: Only top-rated answers for each question are retained. HTML Tag Removal: All… See the full description on the dataset page: https://huggingface.co/datasets/luisroque/instruct-python-llama2-20k.texttext-generation10K<n<100K1 likes32 downloads3y agoHugging Face16LakshayRahal /linkedin-llama2-datasettexttext-generationn<1K0 likes31 downloads1y agoHugging Face17jeqcho /llama-2-13b-chat-hf-mt-bench llama-2-13b-chat-hf-mt-bench MT-Bench outputs for Llama 2 13B chat model Dataset Description This dataset contains model outputs generated using Llama 2 13B model on benchmark questions. Model: meta-llama/Llama-2-13b-hf or meta-llama/Llama-2-13b-chat-hf Benchmark: MT-Bench Generation Date: 2025-10-19 Files mt_bench_llama2_chat.jsonl: Main output file with completions Citation If you use this dataset, please cite the original Llama 2 paper and the… See the full description on the dataset page: https://huggingface.co/datasets/jeqcho/llama-2-13b-chat-hf-mt-bench.texttext-generationn<1K0 likes30 downloads11mo agoHugging Face18mertbozkurt /llama2-TR-recipetexttext-generation10K<n<100K7 likes28 downloads3y agoHugging Face19anyerg21 /Llama-2-7b-chat-finetune plagas y enfermedades en el cultivo del tomate Dataset 1000 Dataset de 1000 instrucciones sobre la plagas y enfermedades en el cultivo del tomate. Uso from datasets import load_dataset dataset = load_dataset("anyerg21/plagas-enfermedades-tomate-1000") Estructura instruction: Pregunta sobre el cultivo del tomate input: Campo vacio output: Respuesta category: Categoria tematica question_type: Tipo de pregunta difficulty: Nivel de dificultad Ejemplo… See the full description on the dataset page: https://huggingface.co/datasets/anyerg21/Llama-2-7b-chat-finetune.textquestion-answeringn<1K0 likes28 downloads1y agoHugging Face20gpjt /openassistant-guanaco-llama2-formatThis dataset is timdettmers/openassistant-guanaco converted to what I believe to be the Llama 2 prompt format (based on this Reddit post). It is otherwise unchanged. The format is like this: <s>[INST] <<SYS>> You are a helpful, respectful and honest assistant. Always answer as helpfully as possible, while being safe. Your answers should not include any harmful, unethical, racist, sexist, toxic, dangerous, or illegal content. Please ensure that your responses are socially unbiased and… See the full description on the dataset page: https://huggingface.co/datasets/gpjt/openassistant-guanaco-llama2-format.texttext-generation10K<n<100K0 likes27 downloads2y agoHugging Face21jeqcho /llama-2-13b-hf-mt-bench llama-2-13b-hf-mt-bench MT-Bench outputs for Llama 2 13B base model Dataset Description This dataset contains model outputs generated using Llama 2 13B model on benchmark questions. Model: meta-llama/Llama-2-13b-hf or meta-llama/Llama-2-13b-chat-hf Benchmark: MT-Bench Generation Date: 2025-10-19 Files mt_bench_llama2_base.jsonl: Main output file with completions Citation If you use this dataset, please cite the original Llama 2 paper and the… See the full description on the dataset page: https://huggingface.co/datasets/jeqcho/llama-2-13b-hf-mt-bench.texttext-generationn<1K0 likes25 downloads11mo agoHugging Face22cheekymachine /enron_labeled_email-prompts-for-llama2_7btexttext-classification1K<n<10K0 likes19 downloads3y agoHugging Face23OdiaGenAI /odia_context_10K_llama2_set Dataset Card for odia_context_10k_llama2_set Dataset Summary This dataset contains 10K instructions that span various facets of Odisha's unique identity. The instructions cover a wide array of subjects, ranging from the culinary delights in 'RECIPES,' the historical significance of 'HISTORICAL PLACES,' and 'TEMPLES OF ODISHA,' to the intellectual pursuits in 'ARITHMETIC,' 'HEALTH,' and 'GEOGRAPHY.' It also explores the artistic tapestry of Odisha through 'ART AND… See the full description on the dataset page: https://huggingface.co/datasets/OdiaGenAI/odia_context_10K_llama2_set.texttext-generation10K<n<100K1 likes18 downloads3y agoHugging Face24VedCodes /llama2_project Dataset Card for Dataset Name Dataset Summary This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/VedCodes/llama2_project.text-generationn<1K0 likes18 downloads3y agoHugging Face25kshitizgajurel /Devanagari-Ecommerce-fomatted-for-llama2-chat-Dataset Dataset Card for Dataset Name यो देवनागरी नेपाली भाषाको डेटासेट विशेषगरी च्याटबोट प्रणालीहरू बनाउनको लागि डिजाइन गरिएको हो। यसमा विभिन्न श्रेणीहरूको डेटासेटहरू समावेश गरिएको छ, जसलाई JSON मा ढाँचा बनाईएको छ, जसले नेपाली वार्तालाप एआई अनुप्रयोगहरूको लागि भाषा मोडेलहरूलाई तालिम र फाइन-ट्यून गर्नको लागि व्यापक स्रोत प्रदान गर्दछ। This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/kshitizgajurel/Devanagari-Ecommerce-fomatted-for-llama2-chat-Dataset.texttext-generation1K<n<10K0 likes17 downloads2y agoHugging Face26rbx-imarcin /llama2-ft-test-datasettexttext-generationn<1K0 likes15 downloads3y agoHugging Face27ar-modeling /illumicore-llama2-1k IllumiCore-1k: Llama2 Formatting This is a VNF resource allocation dataset (1000 samples) generated by IllumiCore [1], processed to match Llama 2's prompt format [2]: <s>[INST] <<SYS>> {{ system_prompt }} <</SYS>> {{ user_msg_1 }} [/INST] {{ model_answer_1 }} </s><s>[INST] {{ user_msg_2 }} [/INST] {{ model_answer_1 }} </s> Here is an example of a dataset record: <s>[INST] <<SYS>> As a telecommunication realm expert with professional knowledge of network function virtualization and… See the full description on the dataset page: https://huggingface.co/datasets/ar-modeling/illumicore-llama2-1k.texttext-generation1K<n<10K1 likes15 downloads3y agoHugging Face28jeqcho /llama-2-13b-chat-hf-alpaca-eval llama-2-13b-chat-hf-alpaca-eval AlpacaEval outputs for Llama 2 13B chat model Dataset Description This dataset contains model outputs generated using Llama 2 13B model on benchmark questions. Model: meta-llama/Llama-2-13b-hf or meta-llama/Llama-2-13b-chat-hf Benchmark: AlpacaEval Generation Date: 2025-10-19 Files alpaca_eval_llama2_chat.json: Main output file with completions Citation If you use this dataset, please cite the original Llama 2 paper… See the full description on the dataset page: https://huggingface.co/datasets/jeqcho/llama-2-13b-chat-hf-alpaca-eval.texttext-generationn<1K0 likes14 downloads11mo agoHugging Face29harishvs /ecommerce-faq-llama2-chattextquestion-answeringn<1K1 likes12 downloads3y agoHugging Face30anuzb /humorchains-llama2-1k 🤖 HumorChains - LLaMA2-1k A dataset of 2,000 humorous one-liners, jokes, and witty responses formatted for instruction-tuned language models (e.g., LLaMA 2, GPT-style).The dataset is designed to help train and fine-tune models that can generate short, punchy, and context-aware humor. 📂 Dataset Summary Name: humorchains-llama2-1k Modality: Text Size: 2,000 samples (~324 KB) Format: Instruction-style (<s>[INST] ... [/INST] ... </s>) Use Case: Fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/anuzb/humorchains-llama2-1k.texttext-generation1K<n<10K0 likes11 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.