CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01llamafactory /demo_data 1,000 examples from https://huggingface.co/datasets/llamafactory/alpaca_gpt4_en 1,000 examples from https://huggingface.co/datasets/llamafactory/alpaca_gpt4_zh 300 examples from https://huggingface.co/datasets/llamafactory/glaive_toolcall_en 300 examples from https://huggingface.co/datasets/llamafactory/glaive_toolcall_zh 91 examples for identity learning 300 examples from https://huggingface.co/datasets/cognitivecomputations/SystemChat-2.0 6 examples for multimodal supervised… See the full description on the dataset page: https://huggingface.co/datasets/llamafactory/demo_data.texttext-generation1K<n<10K1 likes116k downloads2y agoHugging Face02llamafactory /tiny-supervised-datasettexttext-generationn<1K4 likes42k downloads2y agoHugging Face03AIencoder /llama-cpp-wheelsIf you like this please consider liking and donating (https://buymeacoffee.com/aiencoder) 🏭 llama-cpp-python Mega-Factory Wheels "Stop waiting for pip to compile. Just install and run." The most complete collection of pre-built llama-cpp-python wheels in existence — 8,333 wheels across every platform, Python version, backend, and CPU optimization level. No more cmake, gcc, or compilation hell. No more waiting 10 minutes for a build that might fail. Just find your wheel and… See the full description on the dataset page: https://huggingface.co/datasets/AIencoder/llama-cpp-wheels.text-generation1K<n<10K4 likes4.6k downloads1d agoHugging Face04RUC-AIBOX /Llama-3-SynE-Dataset 📄 Report   |   💻 GitHub Repo 🔍 English  |  简体中文 Here is the continual pre-training dataset. The Llama-3-SynE model is available here. News 🌟🌟 2024/12/17: We released the code used for continual pre-training and data preparation. The code contains detailed documentation comments. ✨✨ 2024/08/12: We released the continual pre-training dataset. ✨✨ 2024/08/10: We released the Llama-3-SynE model. ✨ 2024/07/26: We released the technical report, welcome to check it… See the full description on the dataset page: https://huggingface.co/datasets/RUC-AIBOX/Llama-3-SynE-Dataset.texttext-generation100M<n<1B10 likes2.9k downloads1y agoHugging Face05donghyunli /Llama-2-7b-KronQ-HG Llama-2-7b — KronQ H_G (output-side gradient covariance) Paper: arXiv:2607.07964 · Code: GitHub Pre-computed H_G for Llama-2-7b, the output-side curvature factor used by KronQ under the K-FAC factorization H ≈ H_X ⊗ H_G. H_G is the per-sublayer sampled-Fisher gradient covariance (labels drawn from the model distribution) (E[g gᵀ] over the layer output), distinct from the standard input-side Hessian H_X (which GPTQ/GPTAQ build online during calibration). Publishing this lets you… See the full description on the dataset page: https://huggingface.co/datasets/donghyunli/Llama-2-7b-KronQ-HG.text-generation0 likes2.7k downloads2mo agoHugging Face06donghyunli /Meta-Llama-3-70B-KronQ-HG Meta-Llama-3-70B — KronQ H_G (output-side gradient covariance) Paper: arXiv:2607.07964 · Code: GitHub Pre-computed H_G for Meta-Llama-3-70B, the output-side curvature factor used by KronQ under the K-FAC factorization H ≈ H_X ⊗ H_G. Per-sublayer empirical-Fisher gradient covariance (E[g gᵀ] over the layer output), distinct from the input-side Hessian H_X. Publishing this lets you reproduce KronQ quantization without the offline Fisher precompute step. Contents (80… See the full description on the dataset page: https://huggingface.co/datasets/donghyunli/Meta-Llama-3-70B-KronQ-HG.text-generation0 likes2k downloads2mo agoHugging Face07donghyunli /Llama-2-70b-KronQ-HG Llama-2-70b-hf — KronQ H_G (output-side gradient covariance) Paper: arXiv:2607.07964 · Code: GitHub Pre-computed H_G for Llama-2-70b-hf, the output-side curvature factor used by KronQ under the K-FAC factorization H ≈ H_X ⊗ H_G. Per-sublayer empirical-Fisher gradient covariance (E[g gᵀ] over the layer output), distinct from the input-side Hessian H_X. Publishing this lets you reproduce KronQ quantization without the offline Fisher precompute step. Contents (80 layers… See the full description on the dataset page: https://huggingface.co/datasets/donghyunli/Llama-2-70b-KronQ-HG.text-generation0 likes1.7k downloads2mo agoHugging Face08KristianS7 /prepacked-fineweb-edu-llama2-32K-T2048 prepacked-fineweb-edu-llama2-32K-T2048 Pre-tokenized and BOS-aligned best-fit packed version of FineWeb-Edu for training with looped nanochat. Tokenized with the Llama 2 tokenizer (32,000 base vocab + 8 special tokens = 32,008). Stats Train split Source karpathy/fineweb-edu-100b-shuffle (1,821 shards) Total tokens 63.26B Total docs 97.1M Rows 30,873,598 Shards 2,059 (train-00000 to train-02058) Rows per shard ~15,000… See the full description on the dataset page: https://huggingface.co/datasets/KristianS7/prepacked-fineweb-edu-llama2-32K-T2048.text-generation0 likes1.6k downloads6mo agoHugging Face09Magpie-Align /Magpie-Llama-3.1-Pro-300K-Filtered Project Web: https://magpie-align.github.io/ Arxiv Technical Report: https://arxiv.org/abs/2406.08464 Codes: https://github.com/magpie-align/magpie Abstract Click Here High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent… See the full description on the dataset page: https://huggingface.co/datasets/Magpie-Align/Magpie-Llama-3.1-Pro-300K-Filtered.tabulartext-generation100K<n<1M17 likes1.4k downloads2y agoHugging Face10donghyunli /Llama-2-13b-KronQ-HG Llama-2-13b — KronQ H_G (output-side gradient covariance) Paper: arXiv:2607.07964 · Code: GitHub Pre-computed H_G for Llama-2-13b, the output-side curvature factor used by KronQ under the K-FAC factorization H ≈ H_X ⊗ H_G. H_G is the per-sublayer empirical-Fisher gradient covariance (E[g gᵀ] over the layer output), distinct from the standard input-side Hessian H_X (built online during calibration). Publishing this lets you reproduce KronQ quantization without the offline Fisher… See the full description on the dataset page: https://huggingface.co/datasets/donghyunli/Llama-2-13b-KronQ-HG.text-generation0 likes1.3k downloads2mo agoHugging Face11Magpie-Align /Magpie-Llama-3.1-Pro-MT-300K-Filtered Project Web: https://magpie-align.github.io/ Arxiv Technical Report: https://arxiv.org/abs/2406.08464 Codes: https://github.com/magpie-align/magpie Abstract Click Here High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent… See the full description on the dataset page: https://huggingface.co/datasets/Magpie-Align/Magpie-Llama-3.1-Pro-MT-300K-Filtered.tabulartext-generation100K<n<1M17 likes1.2k downloads2y agoHugging Face12botsi /trust-game-llama-2-chat-historytexttext-generationn<1K0 likes1.1k downloads2y agoHugging Face13Arushhh /Llama-HybridDiffusion-processed-data-run1 Llama-HybridDiffusion processed training mixture — run 1 Built with Llama. This repository preserves the exact Hugging Face Dataset.save_to_disk Arrow snapshot used by run 1 of a Qwen3.5-2B HybridDiffusion reproduction. The directory names, dataset_info.json, state.json, and Arrow shard boundaries are retained so the data can be downloaded and supplied to the existing training configuration without a lossy format conversion. Exact snapshot inventory Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Arushhh/Llama-HybridDiffusion-processed-data-run1.text-generation1M<n<10M0 likes1k downloads1mo agoHugging Face14llamafactory /alpaca_gpt4_zhBorrowed from: https://github.com/Instruction-Tuning-with-GPT-4/GPT-4-LLM Removed 6,103 mistruncated examples. You can use it in LLaMA Factory by specifying dataset: alpaca_gpt4_zh. texttext-generation10K<n<100K20 likes931 downloads2y agoHugging Face15donghyunli /Meta-Llama-3-8B-KronQ-HG Meta-Llama-3-8B — KronQ H_G (output-side gradient covariance) Paper: arXiv:2607.07964 · Code: GitHub Pre-computed H_G for Meta-Llama-3-8B, the output-side curvature factor used by KronQ under the K-FAC factorization H ≈ H_X ⊗ H_G. H_G is the per-sublayer empirical-Fisher gradient covariance (E[g gᵀ] over the layer output), distinct from the standard input-side Hessian H_X (built online during calibration). Publishing this lets you reproduce KronQ quantization without the offline… See the full description on the dataset page: https://huggingface.co/datasets/donghyunli/Meta-Llama-3-8B-KronQ-HG.text-generation0 likes835 downloads2mo agoHugging Face16arcee-ai /LLama-405B-Logits Llama-405B-Logits Dataset The Llama-405B-Logits Dataset is a curated subset of logits extracted from the Llama-405B model, created to distill high-performance language models such as Arcee AI's SuperNova using DistillKit. This dataset was also instrumental in the training of the groundbreaking INTELLECT-1 model, demonstrating the effectiveness of leveraging distilled knowledge for enhancing model performance. About the Dataset This dataset contains a carefully… See the full description on the dataset page: https://huggingface.co/datasets/arcee-ai/LLama-405B-Logits.text-generation10K<n<100K14 likes651 downloads2y agoHugging Face17LLaMAX /BenchMAX_Rule-based Dataset Sources Paper: BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models Link: https://huggingface.co/papers/2502.07346 Repository: https://github.com/CONE-MT/BenchMAX Dataset Description BenchMAX_Rule-based is a dataset of BenchMAX, sourcing from IFEval, which is a rule-based benchmark for evaluating the instruction following capabilities in multilingual scenarios. We extend the original dataset to 16 non-English languages by first… See the full description on the dataset page: https://huggingface.co/datasets/LLaMAX/BenchMAX_Rule-based.texttext-generation1K<n<10K1 likes623 downloads2y agoHugging Face18open-paws /tool-use-llama-format Open Paws Tool Use Llama Format This dataset is part of the Open Paws initiative to develop AI training data aligned with animal liberation and advocacy principles. Created to train AI systems that understand and promote animal welfare, rights, and liberation. Dataset Details Dataset Type: Tool Use Data Format: JSONL (JSON Lines) Languages: Multilingual (primarily English) Focus: Animal advocacy and ethical reasoning Organization: Open Paws License: Apache 2.0… See the full description on the dataset page: https://huggingface.co/datasets/open-paws/tool-use-llama-format.texttext-generation1M<n<10M3 likes620 downloads1y agoHugging Face19llamafactory /alpaca_zhBorrowed from: https://huggingface.co/datasets/hfl/alpaca_zh_51k Removed some examples with empty output. You can use it in LLaMA Factory by specifying dataset: alpaca_zh. texttext-generation10K<n<100K4 likes549 downloads2y agoHugging Face20open-paws /visual-qa-llama-format Open Paws Visual Qa Llama Format This dataset is part of the Open Paws initiative to develop AI training data aligned with animal liberation and advocacy principles. Created to train AI systems that understand and promote animal welfare, rights, and liberation. Dataset Details Dataset Type: Multimodal Data Format: JSONL (JSON Lines) Languages: Multilingual (primarily English) Focus: Animal advocacy and ethical reasoning Organization: Open Paws License: Apache 2.0… See the full description on the dataset page: https://huggingface.co/datasets/open-paws/visual-qa-llama-format.imagetext-generation1M<n<10M1 likes541 downloads1y agoHugging Face21LLaMAX /BenchMAX_Problem_Solving Dataset Sources Paper: BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models Link: https://huggingface.co/papers/2502.07346 Repository: https://github.com/CONE-MT/BenchMAX Dataset Description BenchMAX_Problem_Solving is a dataset of BenchMAX, sourcing from LiveCodeBench_v4, which evaluates the code generation capability for solving multilingual competitive code problems. We extend the original English dataset by 16 non-English languages. The… See the full description on the dataset page: https://huggingface.co/datasets/LLaMAX/BenchMAX_Problem_Solving.texttext-generation10K<n<100K1 likes532 downloads2y agoHugging Face22llamafactory /glaive_toolcall_enBorrowed from: https://huggingface.co/datasets/glaiveai/glaive-function-calling-v2 You can use it in LLaMA Factory by specifying dataset: glaive_toolcall_en. texttext-generation1K<n<10K10 likes489 downloads2y agoHugging Face23cartersusi /zig-llama Zig LLama This dataset is used to fine-tune meta-llama/Meta-Llama-3.1-8B-Instruct. Dataset Details The dataset uses ~1100 of the most popular and recently updated Zig repos on GitHub. Dataset Sources The full list of source repos used. The folder of source repos used. texttext-generation100K<n<1M2 likes464 downloads2y agoHugging Face24axiong /pmc_llama_instructionsThis repo provides part of the dataset used for PMC-LLaMA-13B's instruction tuning. Data Size Link ChatDoctor 100K https://www.yunxiangli.top/ChatDoctor/ MedQA 10.2K https://huggingface.co/datasets/GBaker/MedQA-USMLE-4-options MedMCQA 183K https://huggingface.co/datasets/medmcqa PubmedQA 211K https://huggingface.co/datasets/pubmed_qa LiveQA 635 https://huggingface.co/datasets/truehealth/liveqa MedicationQA 690 https://huggingface.co/datasets/truehealth/medicationqa UMLS… See the full description on the dataset page: https://huggingface.co/datasets/axiong/pmc_llama_instructions.textquestion-answering100K<n<1M33 likes453 downloads3y agoHugging Face25emozilla /dolma-v1_7-305B-tokenized-llama3-nanosetTokenized (Llama 3) verison of NousResearch/dolma-v1_7-305B as a Nanotron dataset split into 10 GB chunks. To download: huggingface-cli download --repo-type dataset --local-dir dolma-v1_7-305B-tokenized-llama3-nanoset --local-dir-use-symlinks False NousResearch/dolma-v1_7-305B-tokenized-llama3-nanoset To recombine: cat dolma-v1_7-305B-tokenized-llama3-nanoset/dolma-v1_7-305B-tokenized-llama3-nanoset.npy.* > dolma-v1_7-305B-tokenized-llama3-nanoset.npy rm -rf… See the full description on the dataset page: https://huggingface.co/datasets/emozilla/dolma-v1_7-305B-tokenized-llama3-nanoset.text-generation100B<n<1T1 likes434 downloads2y agoHugging Face26rizerphe /glaive-function-calling-v2-llama Glaive's Function Calling V2 for Llama2 Glaive's Function Calling V2 dataset, formatted according to the Llama2 chat schema, with all the data that I wasn't able to automatically convert removed manually. Adds a special <function> token. Here's an example prompt: <s>[INST] <<SYS>> <function>Available functions: <function>{ "name": "generate_password", "description": "Generate a random password with specified criteria", "parameters": { "type": "object"… See the full description on the dataset page: https://huggingface.co/datasets/rizerphe/glaive-function-calling-v2-llama.texttext-generation100K<n<1M22 likes431 downloads3y agoHugging Face27llamafactory /DPO-En-Zh-20kThis dataset is composed by 4,000 examples of argilla/distilabel-capybara-dpo-7k-binarized with chosen score>=4. 3,000 examples of argilla/distilabel-intel-orca-dpo-pairs with chosen score>=8. 3,000 examples of argilla/ultrafeedback-binarized-preferences-cleaned with chosen score>=4. 10,000 examples of wenbopan/Chinese-dpo-pairs. You can use it in LLaMA Factory by specifying dataset: dpo_mix_en,dpo_mix_zh. texttext-generation10K<n<100K104 likes396 downloads2y agoHugging Face28laion /llama-nemotron-science-reasoning-on-canonical-think-full Llama-Nemotron science reasoning — Delphi canonical-think (COMPLETE, no length filter) The complete reasoning:on science split of nvidia/Llama-Nemotron-Post-Training-Dataset, converted once into the canonical Delphi chat-template thinking format. 708,920 rows. Unlike the cold-start warmup slice open-athena/llama-nemotron-science-reasoning-on-le3000tok-100k (and its -canonical-think variant), this build applies no length cap and no subsample — every long-CoT science example is… See the full description on the dataset page: https://huggingface.co/datasets/laion/llama-nemotron-science-reasoning-on-canonical-think-full.texttext-generation100K<n<1M0 likes395 downloads21d agoHugging Face29Magpie-Align /Magpie-Llama-3.3-Pro-1M-v0.1 Project Web: https://magpie-align.github.io/ Arxiv Technical Report: https://arxiv.org/abs/2406.08464 Codes: https://github.com/magpie-align/magpie Abstract Click Here High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent… See the full description on the dataset page: https://huggingface.co/datasets/Magpie-Align/Magpie-Llama-3.3-Pro-1M-v0.1.tabulartext-generation1M<n<10M5 likes370 downloads2y agoHugging Face30llamafactory /alpaca_enBorrowed from: https://github.com/tatsu-lab/stanford_alpaca Removed some erroneous examples. You can use it in LLaMA Factory by specifying dataset: alpaca_en. texttext-generation10K<n<100K5 likes354 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.