datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
rpj-v2-sampleThis is a mirror of the sample-10B subset of RedPajama-Data-V2 which we have re-uploaded in order to resolve issues with the original download script.
Getting Started
RedPajama-V2 is an open dataset for training large language models. The dataset includes over 100B text
documents coming from 84 CommonCrawl snapshots and processed using
the CCNet pipeline. Out of these, there are 30B documents in the corpus
that additionally come with quality signals. In addition, we also provide the… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/rpj-v2-sample.SampleToHiyori
SampleToHiyori
'모모세 히요리(桃瀬 ひより)' 페르소나 학습용 한국어 데이터셋. config 두 개로 이루어진다.
config
split
행 수
내용
default
train
4,837
단일 턴 한국어 페르소나 대화 (instruction / response)
tools
train / eval
5,495 / 322
도구 호출(function calling) 대화
히요리는 상대를 항상 "오빠" 라고 부르고, 일인칭은 "히요리", 말투는 "인걸" / "인거야" 다.
tools
OpenMascotAI 마스코트의 자비스 모드(윈도우를 실제로 조작하는 모드)에서 쓰기 위한 도구 호출
학습 데이터. 페르소나 LoRA를 얹으면 베이스 모델이 도구를 전혀 호출하지 않게 되는 현상을 고치려고
만들었다. 시나리오(도구·인자·결과·브리프)는 자매 데이터셋 MelissaJ/ProjectLucia_Hera… See the full description on the dataset page: https://huggingface.co/datasets/MelissaJ/SampleToHiyori.lldms-associative-memory-samples
LLDMs Associative Memory — Generated Samples
Model-generated text for the paper:
Language Diffusion Models are Associative Memories Capable of Retrieving Unseen Data
Bao Pham, Mohammed J. Zaki, Luca Ambrogioni, Dmitry Krotov, Matteo Negri
Accepted to EMNLP 2026 (Main Conference).
arXiv:2604.26841 · paper · code · checkpoints
29.5 million generated sequences (~3.8B tokens) sampled from the released checkpoints — one
generation run per (model size, training-set fraction). These… See the full description on the dataset page: https://huggingface.co/datasets/lemoncmd/lldms-associative-memory-samples.fineweb-sample-22.95B-512
FineWeb-Sample-22.95B-512
Dataset Description
This dataset contains approximately 22.95 billion tokens (22,948,244,480 tokens) sampled from the FineWeb-Edu dataset. Each text sample is capped at a maximum of 512 tokens.
Dataset Statistics
Total Tokens: ~22.95B (22,948,244,480)
Max Tokens per Sample: 512
Max Characters per Sample: 5,120 (10 chars/token estimate)
Source Dataset: FineWeb-Edu 350BT
Random Seed: 42
Dataset Structure
The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/JamesConley/fineweb-sample-22.95B-512.DynaMath_Sample
Dataset Card for DynaMath
[💻 Github] [🌐 Homepage][📖 Preprint Paper]
Dataset Details
🔈 Notice
DynaMath is a dynamic benchmark with 501 seed question generators. This dataset is only a sample of 10 variants generated by DynaMath. We encourage you to use the dataset generator on our github site to generate random datasets to test.
🌟 About DynaMath
The rapid advancements in Vision-Language Models (VLMs) have shown significant potential in tackling… See the full description on the dataset page: https://huggingface.co/datasets/DynaMath/DynaMath_Sample.TxT360-5M-sample-en
BEE-spoke-data/TxT360-5M-sample-en
english only sample from LLM360/TxT360:
min length 256 GPT-4 tokens
max length 24576 GPT-4 tokens
GPT-4 tiktoken token count:
token_count
count 5.000000e+06
mean 1.003614e+03
std 1.424231e+03
min 2.570000e+02
25% 4.020000e+02
50% 6.220000e+02
75% 1.050000e+03
max 2.457400e+04
Total count: 5018.07 M tokens
fineweb-edu-sample-10BT-shuffled
📚 FineWeb-Edu (Shuffled)
The samples in HuggingFaceFW/fineweb-edu don't appear to be fully shuffled, leading to oscillating loss curves.
This dataset contains a shuffled version of the sample-10BT sample from HuggingFaceFW/fineweb-edu.
Shuffling was performed using the following script:
import datasets
data = datasets.load_dataset(
"HuggingFaceFW/fineweb-edu",
"sample-10BT",
split="train",
streaming=False,
)
data_shuffled = data.shuffle(seed=42)… See the full description on the dataset page: https://huggingface.co/datasets/aklein4/fineweb-edu-sample-10BT-shuffled.agent-sft-stitch-zh-tts-taste-codec-chat-sample
Gemma 4 E2B Taste-S multi-turn codec SFT
This dataset contains 37,362 complete Traditional Chinese agent
dialogues selected from voidful/agent-sft-stitch-zh-tts. It covers
229,434 synthesized speech segments, approximately
520.5 hours of audio before codec extraction.
Every assistant speech segment is represented without Gemma native audio tags:
<SAY> text_token <a_code> <b_code> ... <p_code> ... </SAY>
The first assistant output starts immediately with <SAY>.
[SOPR]...[EOPR]… See the full description on the dataset page: https://huggingface.co/datasets/voidful/agent-sft-stitch-zh-tts-taste-codec-chat-sample.urls-sampled
URLs (hash-sampled)
The same 74,918,894,107 URLs as
ks46/urls, partitioned by
xxh3_64 range into 2,048 chunks of ≈36.6 M rows instead of by SURT key range.
Each chunk is a uniform random sample of the whole corpus, and a URL's chunk
depends on nothing but the URL itself.
Why this exists
The SURT layout groups the web by host: shard 1,000 is a contiguous slice of the
key space, so it holds whole sites and nothing about any other site. That is
what you want for… See the full description on the dataset page: https://huggingface.co/datasets/ks46/urls-sampled.fineweb-sample-5.97B-512
FineWeb-Sample-5.97B-512
Dataset Description
This dataset contains approximately 5.97 billion tokens (5,968,954,880 tokens) sampled from the FineWeb-Edu dataset. Each text sample is capped at a maximum of 512 tokens.
Dataset Statistics
Total Tokens: ~5.97B (5,968,954,880)
Max Tokens per Sample: 512
Max Characters per Sample: 5,120 (10 chars/token estimate)
Source Dataset: FineWeb-Edu 350BT
Random Seed: 42
Dataset Structure
The dataset is stored in… See the full description on the dataset page: https://huggingface.co/datasets/JamesConley/fineweb-sample-5.97B-512.RedPajama-Data-1T-Sample-Backup
RedPajama Data 1T Sample Backup
This dataset is a backup mirror of togethercomputer/RedPajama-Data-1T-Sample.
It is provided for easier access when the original dataset is unavailable or difficult to download.
Usage
Original:
from datasets import load_dataset
ds = load_dataset(
"togethercomputer/RedPajama-Data-1T-Sample",
split="train",
trust_remote_code=True,
)
Backup:
from datasets import load_dataset
ds = load_dataset(… See the full description on the dataset page: https://huggingface.co/datasets/ll922/RedPajama-Data-1T-Sample-Backup.mermaid_samples_13k
mermaid_samples_13k
Mermaid chart dataset samples about 13k
unwrapped: graph TD ...
wrapped: ```mermaid graph TD ... ```
Checked Mermaid chart's validation
Mermaid validation : 2024/09/10
Mermaid version : 11.0.2
Mermaid visualization : Live Editor
Datasets from
Mixed dataset and select only valid mermaid chart
Celiadraw/text-to-mermaid
Celiadraw/text-to-mermaid-2
rakitha/mermaid-flowchart-transformer
bucaro/mermaid_code… See the full description on the dataset page: https://huggingface.co/datasets/injaeryou/mermaid_samples_13k.nemotron-cc-10K-sample-translated
Translated Nemotron-cc-hq samples
This dataset contains translated samples from https://huggingface.co/datasets/spyysalo/nemotron-cc-10K-sample
Currently, the following are available, we will add other models and languages:
Model
Languages
Gemma-3-4b-it
["Bulgarian", "Czech", "Danish", "German", "Estonian", "Finnish", "French", "Croatian", "Dutch"]
EuroLLM-9B-Instruct
["Bulgarian", "Czech", "Danish", "German", "Greek", "Estonia", "Finnish", "French", "Irish"… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/nemotron-cc-10K-sample-translated.evaded-prompt-injection-and-jailbreak-samplesThis dataset originates from our paper 'Bypassing Prompt Injection and Jailbreak Detection in LLM Guardrails'.
The dataset contains a mixture of prompt injections and jailbreak samples modified via character injection and adversarial ML evasion techniques (Techniques can be found within the paper above). For each sample we provide the original unaltered prompt and a modified prompt, the attack_name outlines which attack technique was used to modify the sample.
Acknowledgements… See the full description on the dataset page: https://huggingface.co/datasets/Mindgard/evaded-prompt-injection-and-jailbreak-samples.TxT360-5M-sample-en
BEE-spoke-data/TxT360-5M-sample-en
english only sample from LLM360/TxT360:
min length 256 GPT-4 tokens
max length 24576 GPT-4 tokens
GPT-4 tiktoken token count:
token_count
count 5.000000e+06
mean 1.003614e+03
std 1.424231e+03
min 2.570000e+02
25% 4.020000e+02
50% 6.220000e+02
75% 1.050000e+03
max 2.457400e+04
Total count: 5018.07 M tokens
tiny-stories-mini-96-seq-len-50000-samples
Source:
noanabeshima/TinyStoriesV2
Purpose:
The purpose of this dataset is for proof of concept smoke - testing of generative architectures from a cold start at the 96 token sequence length on 50,000 text samples.
Description:
A clone of noanabeshima/TinyStoriesV2 that separates the paragraphs into individual text samples, selects samples at or under 96 tokens of length (as determined by the tokenizer HuggingFaceTB/SmolLM3-3B)
egocentric-activity-sample
Egocentric Activity Sample Dataset
A small-scale egocentric (first-person) video dataset with Ego4D-style annotations, designed for quick prototyping and experimentation with egocentric video understanding tasks.
Dataset Summary
Metric
Value
Video clips
19
Total duration
~9.5 minutes
Resolution
960x540 (540p)
FPS
30
Narrations
99
NLQ queries
57
Moment annotations
19
FHO actions
57
Total size
~54 MB
Activities Covered… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/egocentric-activity-sample.algerian-darja-sample
Algerian Darja Sample
A growing Algerian Darja text corpus collected for NLP and language-modeling research.
This dataset is updated incrementally as new sources are collected and processed. Exact sample counts, file sizes, and statistics change between releases — refer to the Dataset Viewer on the repository page for current figures rather than any numbers in this card.
Dataset at a Glance
Sample count, file size, character/word counts, and other metrics are… See the full description on the dataset page: https://huggingface.co/datasets/DarjaCore/algerian-darja-sample.TxT360-1M-sample
BEE-spoke-data/TxT360-1M-sample
One million row sample from LLM360/TxT360:
min length 256 GPT-4 tokens
max length 8192 GPT-4 tokens
fineweb-edu-2016-qwen2-sample
FineWeb-Edu 2016 / Qwen2
Consistency sample — not the completed year.
Documents: 900. Actual recounted Qwen2 tokens: 937,977.
Source: HuggingFaceFW/fineweb-edu, revision 87f09149ef4734204d70ed1d046ddc9ca3f2b8f9.
The input inventory covers 9 crawl directories. date is the integer crawl year 2016,
not an article publication date. Original text is preserved without cleaning,
normalization, deduplication, truncation, or added formatting. Source token counts
are not used. Optional… See the full description on the dataset page: https://huggingface.co/datasets/BoomQ/fineweb-edu-2016-qwen2-sample.reddit-comments-sample
Reddit Comment Trees Sample — Initial snapshot
Initial sample: 43,913 posts and 59,874 comments across three communities.
This is a selected, structurally checked snapshot, with incomplete subreddit coverage. It is not a complete three-month archive.
Overview
Posts and associated comments from r/LocalLLaMA, r/wallstreetbets, and r/SkincareAddiction.
The requested post window is June 1–August 31, 2026 in Asia/Shanghai, with UTC bounds 2026-05-31 16:00:00… See the full description on the dataset page: https://huggingface.co/datasets/mashu-data/reddit-comments-sample.ambig-iac-sample
Ambig-IaC Random 50
A random subset of 50 rows sampled without replacement from all 300 rows of
the default/train split of znyang/ambig-iac.
Source revision: 96429693e7164b024333e6d02f8c6b2d017e4ecb.
Sampling: Python random.Random(42).sample(range(300), 50).
Rows are stored in random draw order. All seven original columns, their types,
values, and original IDs are preserved. No filtering or text changes were made.
The split is named train and contains exactly 50 rows. The… See the full description on the dataset page: https://huggingface.co/datasets/shihanlin/ambig-iac-sample.nemotron-post-training-samples-splits
Nemotron Post-Training Samples with Train/Val/Test Splits
This dataset contains structured train/validation/test splits from the nvidia/Llama-Nemotron-Post-Training-Dataset, with both tagged and untagged versions for different training scenarios.
Attribution
This work is derived from the Llama-Nemotron-Post-Training-Dataset-v1.1 by NVIDIA Corporation, licensed under CC BY 4.0.
Original Dataset: nvidia/Llama-Nemotron-Post-Training-Dataset
Original Authors: NVIDIA… See the full description on the dataset page: https://huggingface.co/datasets/brandolorian/nemotron-post-training-samples-splits.common-crawl-docx-sample
Common Crawl DOCX Sample
A sample of normalized text extracted from DOCX records in Common Crawl.
Source
Common Crawl release: CC-MAIN-YYYY-NN
Source index: Common Crawl URL Index
Pipeline: marin-community/marin
Pipeline revision: REPLACE_WITH_GIT_SHA
Records were selected using declared DOCX MIME type, detected DOCX MIME type,
or a .docx URL suffix. Only successful, non-truncated index records were
eligible.
Processing
The pipeline:… See the full description on the dataset page: https://huggingface.co/datasets/Tinuade/common-crawl-docx-sample.amazon-reviews-2023-all-beauty-sample
Amazon Reviews 2023 – All_Beauty (Sampled)
This dataset is a sampled subset of the McAuley-Lab/Amazon-Reviews-2023
All_Beauty category, prepared for the YZM2022 Data Mining homework
(Assoc. Prof. Dr. Arzu Kakisim).
Sampling strategy
Source: full All_Beauty reviews (701K) and metadata (112K items).
3-core filtering (each user and item has at least 3 interactions, iterated to convergence).
Cap to the most recent 60 000 interactions, re-applied 3-core.
Metadata restricted… See the full description on the dataset page: https://huggingface.co/datasets/debolut/amazon-reviews-2023-all-beauty-sample.dolma3_300B_sample_shuffled
dolma3_300B_sample_shuffled
Global row-level shuffle of TheFinAI/dolma3_300B_sample.
Source data uses per-row Bernoulli sampling (p ≈ 0.0506) from
allenai/dolma3_mix-6T-1025-7B to produce ~300B cl100k tokens preserving
the original Dolma3 mix ratios. However the source parquets cluster
records by sub-source on disk (each ~100K-row parquet groups rows from the
same input shard contiguously), which means a small training shuffle
buffer would see a non-uniform source mix per… See the full description on the dataset page: https://huggingface.co/datasets/TheFinAI/dolma3_300B_sample_shuffled.CivilInstruct-Sample
CivilInstruct-Sample
A 10% sample of the CivilInstruct dataset from the paper Rethinking Scientific Modeling: Toward Physically Consistent and Simulation-Executable Programmatic Generation.
This sample is released for demonstration and reproducibility of the AutoBM training pipeline. The full dataset (10,912 samples) will be released upon paper publication.
Overview
CivilInstruct is a domain-specific instruction dataset for training LLMs to generate executable, physically… See the full description on the dataset page: https://huggingface.co/datasets/yongqiqng/CivilInstruct-Sample.nemotron-post-training-samples
Nemotron Post-Training Samples
This dataset contains random samples extracted from the nvidia/Llama-Nemotron-Post-Training-Dataset.
Attribution
This work is derived from the Llama-Nemotron-Post-Training-Dataset-v1.1 by NVIDIA Corporation, licensed under CC BY 4.0.
Original Dataset: nvidia/Llama-Nemotron-Post-Training-DatasetOriginal Authors: NVIDIA CorporationOriginal License: CC BY 4.0
Dataset Details
Source: nvidia/Llama-Nemotron-Post-Training-Dataset… See the full description on the dataset page: https://huggingface.co/datasets/brandolorian/nemotron-post-training-samples.reasoning-sft-poor-quality-reasoning-sample-mix
Reasoning SFT Sample Mix
A mixed-domain reasoning SFT dataset with multiple response variants per input at varying levels of verbosity and style.
Format
Each row contains an input conversation and several response columns representing different generation strategies applied to the same prompt.
Usage
from datasets import load_dataset
ds = load_dataset("AmanPriyanshu/reasoning-sft-poor-quality-reasoning-sample-mix", split="train")
License
Apache 2.0
amazon-product-data-sample
Dataset Card for "amazon-product-data-filter"
Dataset Summary
The Amazon Product Dataset contains product listing data from the Amazon US website. It can be used for various NLP and classification tasks, such as text generation, product type classification, attribute extraction, image recognition and more.
NOTICE: This is a sample of the full Amazon Product Dataset, which contains 1K examples. Follow the link to gain access to the full dataset.
Languages… See the full description on the dataset page: https://huggingface.co/datasets/iarbel/amazon-product-data-sample.
