CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01EleutherAI /rpj-v2-sampleThis is a mirror of the sample-10B subset of RedPajama-Data-V2 which we have re-uploaded in order to resolve issues with the original download script. Getting Started RedPajama-V2 is an open dataset for training large language models. The dataset includes over 100B text documents coming from 84 CommonCrawl snapshots and processed using the CCNet pipeline. Out of these, there are 30B documents in the corpus that additionally come with quality signals. In addition, we also provide the… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/rpj-v2-sample.texttext-generation1M<n<10M2 likes4.9k downloads2y agoHugging Face02MelissaJ /SampleToHiyori SampleToHiyori '모모세 히요리(桃瀬 ひより)' 페르소나 학습용 한국어 데이터셋. config 두 개로 이루어진다. config split 행 수 내용 default train 4,837 단일 턴 한국어 페르소나 대화 (instruction / response) tools train / eval 5,495 / 322 도구 호출(function calling) 대화 히요리는 상대를 항상 "오빠" 라고 부르고, 일인칭은 "히요리", 말투는 "인걸" / "인거야" 다. tools OpenMascotAI 마스코트의 자비스 모드(윈도우를 실제로 조작하는 모드)에서 쓰기 위한 도구 호출 학습 데이터. 페르소나 LoRA를 얹으면 베이스 모델이 도구를 전혀 호출하지 않게 되는 현상을 고치려고 만들었다. 시나리오(도구·인자·결과·브리프)는 자매 데이터셋 MelissaJ/ProjectLucia_Hera… See the full description on the dataset page: https://huggingface.co/datasets/MelissaJ/SampleToHiyori.texttext-generation10K<n<100K0 likes1.9k downloads22d agoHugging Face03lemoncmd /lldms-associative-memory-samples LLDMs Associative Memory — Generated Samples Model-generated text for the paper: Language Diffusion Models are Associative Memories Capable of Retrieving Unseen Data Bao Pham, Mohammed J. Zaki, Luca Ambrogioni, Dmitry Krotov, Matteo Negri Accepted to EMNLP 2026 (Main Conference). arXiv:2604.26841 · paper · code · checkpoints 29.5 million generated sequences (~3.8B tokens) sampled from the released checkpoints — one generation run per (model size, training-set fraction). These… See the full description on the dataset page: https://huggingface.co/datasets/lemoncmd/lldms-associative-memory-samples.text-generation10M<n<100M0 likes1.8k downloads27d agoHugging Face04JamesConley /fineweb-sample-22.95B-512 FineWeb-Sample-22.95B-512 Dataset Description This dataset contains approximately 22.95 billion tokens (22,948,244,480 tokens) sampled from the FineWeb-Edu dataset. Each text sample is capped at a maximum of 512 tokens. Dataset Statistics Total Tokens: ~22.95B (22,948,244,480) Max Tokens per Sample: 512 Max Characters per Sample: 5,120 (10 chars/token estimate) Source Dataset: FineWeb-Edu 350BT Random Seed: 42 Dataset Structure The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/JamesConley/fineweb-sample-22.95B-512.texttext-generation10M<n<100M0 likes884 downloads11mo agoHugging Face05DynaMath /DynaMath_Sample Dataset Card for DynaMath [💻 Github] [🌐 Homepage][📖 Preprint Paper] Dataset Details 🔈 Notice DynaMath is a dynamic benchmark with 501 seed question generators. This dataset is only a sample of 10 variants generated by DynaMath. We encourage you to use the dataset generator on our github site to generate random datasets to test. 🌟 About DynaMath The rapid advancements in Vision-Language Models (VLMs) have shown significant potential in tackling… See the full description on the dataset page: https://huggingface.co/datasets/DynaMath/DynaMath_Sample.imagemultiple-choice1K<n<10K9 likes748 downloads2y agoHugging Face06BEE-spoke-data /TxT360-5M-sample-en BEE-spoke-data/TxT360-5M-sample-en english only sample from LLM360/TxT360: min length 256 GPT-4 tokens max length 24576 GPT-4 tokens GPT-4 tiktoken token count: token_count count 5.000000e+06 mean 1.003614e+03 std 1.424231e+03 min 2.570000e+02 25% 4.020000e+02 50% 6.220000e+02 75% 1.050000e+03 max 2.457400e+04 Total count: 5018.07 M tokens texttext-generation10M<n<100M3 likes736 downloads9mo agoHugging Face07aklein4 /fineweb-edu-sample-10BT-shuffled 📚 FineWeb-Edu (Shuffled) The samples in HuggingFaceFW/fineweb-edu don't appear to be fully shuffled, leading to oscillating loss curves. This dataset contains a shuffled version of the sample-10BT sample from HuggingFaceFW/fineweb-edu. Shuffling was performed using the following script: import datasets data = datasets.load_dataset( "HuggingFaceFW/fineweb-edu", "sample-10BT", split="train", streaming=False, ) data_shuffled = data.shuffle(seed=42)… See the full description on the dataset page: https://huggingface.co/datasets/aklein4/fineweb-edu-sample-10BT-shuffled.tabulartext-generation1M<n<10M1 likes396 downloads1y agoHugging Face08voidful /agent-sft-stitch-zh-tts-taste-codec-chat-sample Gemma 4 E2B Taste-S multi-turn codec SFT This dataset contains 37,362 complete Traditional Chinese agent dialogues selected from voidful/agent-sft-stitch-zh-tts. It covers 229,434 synthesized speech segments, approximately 520.5 hours of audio before codec extraction. Every assistant speech segment is represented without Gemma native audio tags: <SAY> text_token <a_code> <b_code> ... <p_code> ... </SAY> The first assistant output starts immediately with <SAY>. [SOPR]...[EOPR]… See the full description on the dataset page: https://huggingface.co/datasets/voidful/agent-sft-stitch-zh-tts-taste-codec-chat-sample.tabulartext-generation10K<n<100K0 likes309 downloads2mo agoHugging Face09ks46 /urls-sampled URLs (hash-sampled) The same 74,918,894,107 URLs as ks46/urls, partitioned by xxh3_64 range into 2,048 chunks of ≈36.6 M rows instead of by SURT key range. Each chunk is a uniform random sample of the whole corpus, and a URL's chunk depends on nothing but the URL itself. Why this exists The SURT layout groups the web by host: shard 1,000 is a contiguous slice of the key space, so it holds whole sites and nothing about any other site. That is what you want for… See the full description on the dataset page: https://huggingface.co/datasets/ks46/urls-sampled.texttext-generation10B<n<100B0 likes254 downloads1mo agoHugging Face10JamesConley /fineweb-sample-5.97B-512 FineWeb-Sample-5.97B-512 Dataset Description This dataset contains approximately 5.97 billion tokens (5,968,954,880 tokens) sampled from the FineWeb-Edu dataset. Each text sample is capped at a maximum of 512 tokens. Dataset Statistics Total Tokens: ~5.97B (5,968,954,880) Max Tokens per Sample: 512 Max Characters per Sample: 5,120 (10 chars/token estimate) Source Dataset: FineWeb-Edu 350BT Random Seed: 42 Dataset Structure The dataset is stored in… See the full description on the dataset page: https://huggingface.co/datasets/JamesConley/fineweb-sample-5.97B-512.texttext-generation10M<n<100M0 likes241 downloads11mo agoHugging Face11ll922 /RedPajama-Data-1T-Sample-Backup RedPajama Data 1T Sample Backup This dataset is a backup mirror of togethercomputer/RedPajama-Data-1T-Sample. It is provided for easier access when the original dataset is unavailable or difficult to download. Usage Original: from datasets import load_dataset ds = load_dataset( "togethercomputer/RedPajama-Data-1T-Sample", split="train", trust_remote_code=True, ) Backup: from datasets import load_dataset ds = load_dataset(… See the full description on the dataset page: https://huggingface.co/datasets/ll922/RedPajama-Data-1T-Sample-Backup.texttext-generation100K<n<1M0 likes193 downloads5mo agoHugging Face12injaeryou /mermaid_samples_13k mermaid_samples_13k Mermaid chart dataset samples about 13k unwrapped: graph TD ... wrapped: ```mermaid graph TD ... ``` Checked Mermaid chart's validation Mermaid validation : 2024/09/10 Mermaid version : 11.0.2 Mermaid visualization : Live Editor Datasets from Mixed dataset and select only valid mermaid chart Celiadraw/text-to-mermaid Celiadraw/text-to-mermaid-2 rakitha/mermaid-flowchart-transformer bucaro/mermaid_code… See the full description on the dataset page: https://huggingface.co/datasets/injaeryou/mermaid_samples_13k.texttext-generation10K<n<100K2 likes187 downloads2y agoHugging Face13openeurollm /nemotron-cc-10K-sample-translated Translated Nemotron-cc-hq samples This dataset contains translated samples from https://huggingface.co/datasets/spyysalo/nemotron-cc-10K-sample Currently, the following are available, we will add other models and languages: Model Languages Gemma-3-4b-it ["Bulgarian", "Czech", "Danish", "German", "Estonian", "Finnish", "French", "Croatian", "Dutch"] EuroLLM-9B-Instruct ["Bulgarian", "Czech", "Danish", "German", "Greek", "Estonia", "Finnish", "French", "Irish"… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/nemotron-cc-10K-sample-translated.texttext-generation100K<n<1M1 likes174 downloads1y agoHugging Face14Mindgard /evaded-prompt-injection-and-jailbreak-samplesgatedThis dataset originates from our paper 'Bypassing Prompt Injection and Jailbreak Detection in LLM Guardrails'. The dataset contains a mixture of prompt injections and jailbreak samples modified via character injection and adversarial ML evasion techniques (Techniques can be found within the paper above). For each sample we provide the original unaltered prompt and a modified prompt, the attack_name outlines which attack technique was used to modify the sample. Acknowledgements… See the full description on the dataset page: https://huggingface.co/datasets/Mindgard/evaded-prompt-injection-and-jailbreak-samples.texttext-classification10K<n<100K20 likes168 downloads1y agoHugging Face15EliMC /TxT360-5M-sample-en BEE-spoke-data/TxT360-5M-sample-en english only sample from LLM360/TxT360: min length 256 GPT-4 tokens max length 24576 GPT-4 tokens GPT-4 tiktoken token count: token_count count 5.000000e+06 mean 1.003614e+03 std 1.424231e+03 min 2.570000e+02 25% 4.020000e+02 50% 6.220000e+02 75% 1.050000e+03 max 2.457400e+04 Total count: 5018.07 M tokens texttext-generation10M<n<100M0 likes155 downloads10mo agoHugging Face16david-thrower /tiny-stories-mini-96-seq-len-50000-samples Source: noanabeshima/TinyStoriesV2 Purpose: The purpose of this dataset is for proof of concept smoke - testing of generative architectures from a cold start at the 96 token sequence length on 50,000 text samples. Description: A clone of noanabeshima/TinyStoriesV2 that separates the paragraphs into individual text samples, selects samples at or under 96 tokens of length (as determined by the tokenizer HuggingFaceTB/SmolLM3-3B) texttext-generation10K<n<100K0 likes148 downloads8mo agoHugging Face17WhissleAI /egocentric-activity-sample Egocentric Activity Sample Dataset A small-scale egocentric (first-person) video dataset with Ego4D-style annotations, designed for quick prototyping and experimentation with egocentric video understanding tasks. Dataset Summary Metric Value Video clips 19 Total duration ~9.5 minutes Resolution 960x540 (540p) FPS 30 Narrations 99 NLQ queries 57 Moment annotations 19 FHO actions 57 Total size ~54 MB Activities Covered… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/egocentric-activity-sample.tabularvideo-classificationn<1K0 likes101 downloads5mo agoHugging Face18DarjaCore /algerian-darja-sample Algerian Darja Sample A growing Algerian Darja text corpus collected for NLP and language-modeling research. This dataset is updated incrementally as new sources are collected and processed. Exact sample counts, file sizes, and statistics change between releases — refer to the Dataset Viewer on the repository page for current figures rather than any numbers in this card. Dataset at a Glance Sample count, file size, character/word counts, and other metrics are… See the full description on the dataset page: https://huggingface.co/datasets/DarjaCore/algerian-darja-sample.texttext-generation1M<n<10M1 likes98 downloads6d agoHugging Face19BEE-spoke-data /TxT360-1M-sample BEE-spoke-data/TxT360-1M-sample One million row sample from LLM360/TxT360: min length 256 GPT-4 tokens max length 8192 GPT-4 tokens texttext-generation1M<n<10M0 likes89 downloads9mo agoHugging Face20BoomQ /fineweb-edu-2016-qwen2-sample FineWeb-Edu 2016 / Qwen2 Consistency sample — not the completed year. Documents: 900. Actual recounted Qwen2 tokens: 937,977. Source: HuggingFaceFW/fineweb-edu, revision 87f09149ef4734204d70ed1d046ddc9ca3f2b8f9. The input inventory covers 9 crawl directories. date is the integer crawl year 2016, not an article publication date. Original text is preserved without cleaning, normalization, deduplication, truncation, or added formatting. Source token counts are not used. Optional… See the full description on the dataset page: https://huggingface.co/datasets/BoomQ/fineweb-edu-2016-qwen2-sample.tabulartext-generationn<1K0 likes79 downloads15d agoHugging Face21mashu-data /reddit-comments-sample Reddit Comment Trees Sample — Initial snapshot Initial sample: 43,913 posts and 59,874 comments across three communities. This is a selected, structurally checked snapshot, with incomplete subreddit coverage. It is not a complete three-month archive. Overview Posts and associated comments from r/LocalLLaMA, r/wallstreetbets, and r/SkincareAddiction. The requested post window is June 1–August 31, 2026 in Asia/Shanghai, with UTC bounds 2026-05-31 16:00:00… See the full description on the dataset page: https://huggingface.co/datasets/mashu-data/reddit-comments-sample.tabulartext-generation100K<n<1M0 likes68 downloads13d agoHugging Face22shihanlin /ambig-iac-sample Ambig-IaC Random 50 A random subset of 50 rows sampled without replacement from all 300 rows of the default/train split of znyang/ambig-iac. Source revision: 96429693e7164b024333e6d02f8c6b2d017e4ecb. Sampling: Python random.Random(42).sample(range(300), 50). Rows are stored in random draw order. All seven original columns, their types, values, and original IDs are preserved. No filtering or text changes were made. The split is named train and contains exactly 50 rows. The… See the full description on the dataset page: https://huggingface.co/datasets/shihanlin/ambig-iac-sample.texttext-generationn<1K1 likes58 downloads9d agoHugging Face23brandolorian /nemotron-post-training-samples-splits Nemotron Post-Training Samples with Train/Val/Test Splits This dataset contains structured train/validation/test splits from the nvidia/Llama-Nemotron-Post-Training-Dataset, with both tagged and untagged versions for different training scenarios. Attribution This work is derived from the Llama-Nemotron-Post-Training-Dataset-v1.1 by NVIDIA Corporation, licensed under CC BY 4.0. Original Dataset: nvidia/Llama-Nemotron-Post-Training-Dataset Original Authors: NVIDIA… See the full description on the dataset page: https://huggingface.co/datasets/brandolorian/nemotron-post-training-samples-splits.texttext-generation10K<n<100K0 likes48 downloads1y agoHugging Face24Tinuade /common-crawl-docx-sample Common Crawl DOCX Sample A sample of normalized text extracted from DOCX records in Common Crawl. Source Common Crawl release: CC-MAIN-YYYY-NN Source index: Common Crawl URL Index Pipeline: marin-community/marin Pipeline revision: REPLACE_WITH_GIT_SHA Records were selected using declared DOCX MIME type, detected DOCX MIME type, or a .docx URL suffix. Only successful, non-truncated index records were eligible. Processing The pipeline:… See the full description on the dataset page: https://huggingface.co/datasets/Tinuade/common-crawl-docx-sample.tabulartext-generation1K<n<10K0 likes48 downloads9d agoHugging Face25debolut /amazon-reviews-2023-all-beauty-sample Amazon Reviews 2023 – All_Beauty (Sampled) This dataset is a sampled subset of the McAuley-Lab/Amazon-Reviews-2023 All_Beauty category, prepared for the YZM2022 Data Mining homework (Assoc. Prof. Dr. Arzu Kakisim). Sampling strategy Source: full All_Beauty reviews (701K) and metadata (112K items). 3-core filtering (each user and item has at least 3 interactions, iterated to convergence). Cap to the most recent 60 000 interactions, re-applied 3-core. Metadata restricted… See the full description on the dataset page: https://huggingface.co/datasets/debolut/amazon-reviews-2023-all-beauty-sample.tabulartext-classification10K<n<100K0 likes47 downloads4mo agoHugging Face26TheFinAI /dolma3_300B_sample_shuffled dolma3_300B_sample_shuffled Global row-level shuffle of TheFinAI/dolma3_300B_sample. Source data uses per-row Bernoulli sampling (p ≈ 0.0506) from allenai/dolma3_mix-6T-1025-7B to produce ~300B cl100k tokens preserving the original Dolma3 mix ratios. However the source parquets cluster records by sub-source on disk (each ~100K-row parquet groups rows from the same input shard contiguously), which means a small training shuffle buffer would see a non-uniform source mix per… See the full description on the dataset page: https://huggingface.co/datasets/TheFinAI/dolma3_300B_sample_shuffled.tabulartext-generation100M<n<1B0 likes46 downloads4mo agoHugging Face27yongqiqng /CivilInstruct-Sample CivilInstruct-Sample A 10% sample of the CivilInstruct dataset from the paper Rethinking Scientific Modeling: Toward Physically Consistent and Simulation-Executable Programmatic Generation. This sample is released for demonstration and reproducibility of the AutoBM training pipeline. The full dataset (10,912 samples) will be released upon paper publication. Overview CivilInstruct is a domain-specific instruction dataset for training LLMs to generate executable, physically… See the full description on the dataset page: https://huggingface.co/datasets/yongqiqng/CivilInstruct-Sample.texttext-generation1K<n<10K0 likes44 downloads6mo agoHugging Face28brandolorian /nemotron-post-training-samples Nemotron Post-Training Samples This dataset contains random samples extracted from the nvidia/Llama-Nemotron-Post-Training-Dataset. Attribution This work is derived from the Llama-Nemotron-Post-Training-Dataset-v1.1 by NVIDIA Corporation, licensed under CC BY 4.0. Original Dataset: nvidia/Llama-Nemotron-Post-Training-DatasetOriginal Authors: NVIDIA CorporationOriginal License: CC BY 4.0 Dataset Details Source: nvidia/Llama-Nemotron-Post-Training-Dataset… See the full description on the dataset page: https://huggingface.co/datasets/brandolorian/nemotron-post-training-samples.texttext-generation10K<n<100K0 likes40 downloads1y agoHugging Face29AmanPriyanshu /reasoning-sft-poor-quality-reasoning-sample-mix Reasoning SFT Sample Mix A mixed-domain reasoning SFT dataset with multiple response variants per input at varying levels of verbosity and style. Format Each row contains an input conversation and several response columns representing different generation strategies applied to the same prompt. Usage from datasets import load_dataset ds = load_dataset("AmanPriyanshu/reasoning-sft-poor-quality-reasoning-sample-mix", split="train") License Apache 2.0 texttext-generation100K<n<1M0 likes39 downloads6mo agoHugging Face30iarbel /amazon-product-data-sample Dataset Card for "amazon-product-data-filter" Dataset Summary The Amazon Product Dataset contains product listing data from the Amazon US website. It can be used for various NLP and classification tasks, such as text generation, product type classification, attribute extraction, image recognition and more. NOTICE: This is a sample of the full Amazon Product Dataset, which contains 1K examples. Follow the link to gain access to the full dataset. Languages… See the full description on the dataset page: https://huggingface.co/datasets/iarbel/amazon-product-data-sample.imagetext-generationn<1K1 likes35 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.