CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01llamafactory /demo_data 1,000 examples from https://huggingface.co/datasets/llamafactory/alpaca_gpt4_en 1,000 examples from https://huggingface.co/datasets/llamafactory/alpaca_gpt4_zh 300 examples from https://huggingface.co/datasets/llamafactory/glaive_toolcall_en 300 examples from https://huggingface.co/datasets/llamafactory/glaive_toolcall_zh 91 examples for identity learning 300 examples from https://huggingface.co/datasets/cognitivecomputations/SystemChat-2.0 6 examples for multimodal supervised… See the full description on the dataset page: https://huggingface.co/datasets/llamafactory/demo_data.texttext-generation1K<n<10K1 likes119k downloads2y agoHugging Face02llamafactory /tiny-supervised-datasettexttext-generationn<1K4 likes45k downloads2y agoHugging Face03llamaindex /ParseBench ParseBench Quick links: [🌐 Website] [📜 Paper] [💻 Code] ParseBench is a benchmark for evaluating document parsing systems on real-world enterprise documents, with the following characteristics: Multi-dimensional evaluation. The benchmark is stratified into five capability dimensions — tables, charts, content faithfulness, semantic formatting, and visual grounding — each with task-specific metrics designed to capture what agentic workflows depend on. Real-world enterprise… See the full description on the dataset page: https://huggingface.co/datasets/llamaindex/ParseBench.document100K<n<1M130 likes23k downloads5mo agoHugging Face04llamaindex /ExtractBench ExtractBench Quick links: [🌐 Website] [📜 Paper] [💻 Code] Given a document and a schema, a system returns structured data with evidence. The input is a full document, born-digital or scanned, and a schema written by the user. The output is a schema-valid JSON object, with the source page and a bounding box for each value as evidence. It must return correct, exhaustive values (including repeated records), correctly use null for absent information, and ground each extracted… See the full description on the dataset page: https://huggingface.co/datasets/llamaindex/ExtractBench.documentn<1K32 likes19k downloads1mo agoHugging Face05scaleinvariant /paired-llama-3.2-1b-embeddings-lmsys-chat-1m Paired Llama 3.2 1B Token Embeddings (LMSYS-Chat-1M) This dataset contains paired activations corresponding to single token locations extracted from Meta's Llama 3.2 1B Instruct on conversations from LMSYS-Chat-1M. Embeddings are provided for layers 5 through 14, which capture the most interesting intermediate representations. This dataset was built to study things like: Learning different basis for activations at a given layer Studying if there are cases where position encodes… See the full description on the dataset page: https://huggingface.co/datasets/scaleinvariant/paired-llama-3.2-1b-embeddings-lmsys-chat-1m.tabularfeature-extraction100M<n<1B3 likes6.6k downloads7mo agoHugging Face06nvidia /Llama-Nemotron-Post-Training-Dataset Llama-Nemotron-Post-Training-Dataset-v1.1 Release Update [4/8/2025]: v1.1: We are releasing an additional 2.2M Math and 500K Code Reasoning Data in support of our release of Llama-3.1-Nemotron-Ultra-253B-v1. 🎉 Data Overview This dataset is a compilation of SFT and RL data that supports improvements of math, code, general reasoning, and instruction following capabilities of the original Llama instruct model, in support of NVIDIA’s release of… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Llama-Nemotron-Post-Training-Dataset.text1M<n<10M709 likes6.5k downloads1y agoHugging Face07Mechanistic-Anomaly-Detection /llama3-jailbreakstext10K<n<100K4 likes5k downloads2y agoHugging Face08SaylorTwift /details_meta-llama__Llama-3.1-8B-Instruct_private Dataset Card for Evaluation run of meta-llama/Llama-3.1-8B-Instruct Dataset automatically created during the evaluation run of model meta-llama/Llama-3.1-8B-Instruct. The dataset is composed of 78 configuration, each one corresponding to one of the evaluated task. The dataset has been created from 20 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/SaylorTwift/details_meta-llama__Llama-3.1-8B-Instruct_private.textn<1K0 likes4.5k downloads1y agoHugging Face09Realmbird /nla-av-responses-llama-70b-layer53tabular1K<n<10K0 likes4.4k downloads4mo agoHugging Face10kevin009 /olympiad-math-contest-llama3-78ktext10K<n<100K1 likes3.7k downloads2y agoHugging Face11llamafactory /v1-sft-demotextn<1K0 likes3.5k downloads10mo agoHugging Face12allenai /llama-3.1-tulu-3-8b-preference-mixture Tulu 3 8B Preference Mixture Note that this collection is licensed under ODC-BY-1.0 license; different licenses apply to subsets of the data. Some portions of the dataset are non-commercial. We present the mixture as a research artifact. This mix is made up from the following preference datasets: https://huggingface.co/datasets/allenai/tulu-3-sft-reused-off-policy https://huggingface.co/datasets/allenai/tulu-3-sft-reused-on-policy-8b… See the full description on the dataset page: https://huggingface.co/datasets/allenai/llama-3.1-tulu-3-8b-preference-mixture.text100K<n<1M27 likes2.9k downloads2y agoHugging Face13RUC-AIBOX /Llama-3-SynE-Dataset 📄 Report   |   💻 GitHub Repo 🔍 English  |  简体中文 Here is the continual pre-training dataset. The Llama-3-SynE model is available here. News 🌟🌟 2024/12/17: We released the code used for continual pre-training and data preparation. The code contains detailed documentation comments. ✨✨ 2024/08/12: We released the continual pre-training dataset. ✨✨ 2024/08/10: We released the Llama-3-SynE model. ✨ 2024/07/26: We released the technical report, welcome to check it… See the full description on the dataset page: https://huggingface.co/datasets/RUC-AIBOX/Llama-3-SynE-Dataset.texttext-generation100M<n<1B10 likes2.8k downloads1y agoHugging Face14nvidia /Llama-Nemotron-VLM-Dataset-v1 Llama-Nemotron-VLM-Dataset v1 Versions Date Commit Changes 2025-08-11 bdb3899 Initial release 2025-08-18 5abc7df Fixes bug (ocr_1 and ocr_3 images were swapped) 2025-08-19 ef85bef Update instructions for ocr_9 2025-08-25 4e46f2b Added example for Megatron Energon 2025-09-02 head Update license headers Quickstart If you want to dive in right away and load some samples using Megatron Energon, check out this section below. Data… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Llama-Nemotron-VLM-Dataset-v1.textvisual-question-answering1M<n<10M168 likes2.7k downloads11mo agoHugging Face15llamastack /mmlu_cottext100K<n<1M0 likes2.3k downloads1y agoHugging Face16OALL /details_grimjim__Llama-3-Instruct-8B-SimPO-SPPO-Iter3-merge Dataset Card for Evaluation run of grimjim/Llama-3-Instruct-8B-SimPO-SPPO-Iter3-merge Dataset automatically created during the evaluation run of model grimjim/Llama-3-Instruct-8B-SimPO-SPPO-Iter3-merge. The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_grimjim__Llama-3-Instruct-8B-SimPO-SPPO-Iter3-merge.tabular100K<n<1M0 likes2.2k downloads2y agoHugging Face17neuralmagic /quantized-llama-3.1-leaderboard-v2-evals Open LLM Leaderboard v2 Benchmark Results This artifact contains all the data from evaluations of Neural Magic's quantized Llama-3.1 models. These evaluations were produced with lm-evaluation-harness by running the following command: lm_eval \ --model vllm \ --model_args pretrained="<model_path>",dtype=auto,add_bos_token=False,max_model_len=4096,tensor_parallel_size="<num_gpus>",gpu_memory_utilization=0.8,enable_chunked_prefill=True \ --apply_chat_template \… See the full description on the dataset page: https://huggingface.co/datasets/neuralmagic/quantized-llama-3.1-leaderboard-v2-evals.tabular100K<n<1M0 likes2.1k downloads2y agoHugging Face18aidando73 /llama-codes-swe-bench-evalstext100K<n<1M0 likes2.1k downloads2y agoHugging Face19llama-farm /military-labeled-yolo Military-Labeled YOLO Dataset (DVIDS sourced) Multi-class military object detection in YOLO format. Source images pulled from the Defense Visual Information Distribution Service (DVIDS) public domain library; labeled via in-house Gemini-VLM-assisted pipeline with human-in-the-loop correction. Classes (12) ID Name 0 soldier 1 tank 2 apc 3 artillery 4 mlrs 5 military_truck 6 helicopter 7 aircraft 8 warship 9 missile_launcher 10 car… See the full description on the dataset page: https://huggingface.co/datasets/llama-farm/military-labeled-yolo.imageobject-detection1K<n<10K2 likes2.1k downloads4mo agoHugging Face20automated-research-group /llama2_7b_chat-boolq-results Dataset Card for "llama2_7b_chat-boolq-results" More Information needed text100K<n<1M1 likes2k downloads3y agoHugging Face21HuggingFaceTB /everyday-conversations-llama3.1-2k Everyday conversations for Smol LLMs finetunings This dataset contains 2.2k multi-turn conversations generated by Llama-3.1-70B-Instruct. We ask the LLM to generate a simple multi-turn conversation, with 3-4 short exchanges, between a User and an AI Assistant about a certain topic. The topics are chosen to be simple to understand by smol LLMs and cover everyday topics + elementary science. We include: 20 everyday topics with 100 subtopics each 43 elementary science topics with 10… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceTB/everyday-conversations-llama3.1-2k.text1K<n<10K138 likes2k downloads2y agoHugging Face22kevin009 /olympiad-math-stepwise-solutions-llama3-20kThe MATH dataset is a collection of 20,300 problems from AMC and AIME competitions covering algebra, number theory, geometry, and precalculus problems and solution sets. Problems and solutions are formatted in LATEX. Step-by-step solutions and insight sections have been added in order to use a chain of thought to clarify the problem and solution. text10K<n<100K1 likes2k downloads2y agoHugging Face23Yhyu13 /glaive-function-calling-v2-llama-factory-convertThis is a converted dataset for https://huggingface.co/datasets/glaiveai/glaive-function-calling-v2 that allows sft in https://github.com/hiyouga/LLaMA-Factory for function calling fine tuning. You need to add the following to the datasets.json file, and changed the file_name to your local path. "glaive-function-calling-v2": { "file_name": "./glaive-function-calling-v2/simple-function-calling-v2_converted.json", "columns": { "prompt": "instruction", "query": "input"… See the full description on the dataset page: https://huggingface.co/datasets/Yhyu13/glaive-function-calling-v2-llama-factory-convert.text100K<n<1M6 likes1.9k downloads3y agoHugging Face24Muesli1 /dclm-baseline-1.0-llama3-tokenized-shuffled !! Note: this dataset is currently being uploaded and processed. The .bin files are intermediate files to allow shuffling. !! DCLM-Baseline Pretokenized (LLaMA 3.1, 8192 context) This dataset is a pretokenized and globally shuffled version of DCLM-Baseline (mlfoundations/dclm-baseline-1.0), prepared for large-scale language model pretraining. It is intended to be used as a direct drop-in pretraining corpus for LLaMA 3.1 style training pipelines. The original DCLM-Baseline… See the full description on the dataset page: https://huggingface.co/datasets/Muesli1/dclm-baseline-1.0-llama3-tokenized-shuffled.text0 likes1.7k downloads5mo agoHugging Face25TAUR-Lab /Taur_CoT_Analysis_Project___meta-llama__Meta-Llama-3.1-8B-Instructtext10K<n<100K0 likes1.5k downloads2y agoHugging Face26Magpie-Align /Magpie-Llama-3.1-Pro-300K-Filtered Project Web: https://magpie-align.github.io/ Arxiv Technical Report: https://arxiv.org/abs/2406.08464 Codes: https://github.com/magpie-align/magpie Abstract Click Here High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent… See the full description on the dataset page: https://huggingface.co/datasets/Magpie-Align/Magpie-Llama-3.1-Pro-300K-Filtered.tabulartext-generation100K<n<1M17 likes1.5k downloads2y agoHugging Face27ENSEONG /full-math-private-n256-Llama-3.2-3B-Instruct-bontabular100K<n<1M0 likes1.4k downloads5mo agoHugging Face28llamafactory /reason-tool-use-demo-1500 Dataset info The dataset is a selection of reasoning toolcalls data from https://huggingface.co/datasets/interstellarninja/hermes_reasoning_tool_use, which contains data from Hermes-Tools、Glaive-FC、ToolAce、Nvidia-When2Call. The format has been transformed to adapt llama-factory v1 training pipeline. textquestion-answering1K<n<10K1 likes1.4k downloads9mo agoHugging Face29ibm-esa-geospatial /Llama3-SSL4EO-S12-v1.1-captions Llama3-SSL4EO-S12-Captions The captions are aligned with the SSL4EO-S12 v1.1 dataset and were automatically generated using the Llama3-LLaVA-Next-8B model. Please find more information regarding the generation and evaluation in the Llama3-MS-CLIP paper. Code: https://github.com/IBM/MS-CLIP Data Structure We provide the captions in two versions: As a single compressed Parquet file per split and as CSV files with 256 captions each that match the Zarr Zip files of the… See the full description on the dataset page: https://huggingface.co/datasets/ibm-esa-geospatial/Llama3-SSL4EO-S12-v1.1-captions.tabularzero-shot-image-classification100K<n<1M5 likes1.4k downloads1y agoHugging Face30Magpie-Align /Magpie-Llama-3.1-Pro-MT-300K-Filtered Project Web: https://magpie-align.github.io/ Arxiv Technical Report: https://arxiv.org/abs/2406.08464 Codes: https://github.com/magpie-align/magpie Abstract Click Here High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent… See the full description on the dataset page: https://huggingface.co/datasets/Magpie-Align/Magpie-Llama-3.1-Pro-MT-300K-Filtered.tabulartext-generation100K<n<1M17 likes1.3k downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.