datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
smollm-chunked
FAISS Indices and Chunked Datasets for SmolLM and SmolLM2 corpora
This repository contains part of the FAISS indices and chunked datasets used for novelty detection for SmolLM and SmolLM2, as presented in the paper LLM generation novelty through the lens of semantic similarity.
Full Documentation
For complete usage instructions, installation guide, and tutorial, please refer to:
Main Tutorial README
Data Distribution
Due to Hugging Face storage quota… See the full description on the dataset page: https://huggingface.co/datasets/enguyen/smollm-chunked.dl3dv_chunked
DL3DV Post-processed for Less3Depend
This dataset is a post-processed version of the DL3DV-10K dataset, specifically prepared for the repository 👉 Less3Depend.
The goal of this release is to provide a clean, unified, and research-ready variant of DL3DV that is directly usable in 👉 PixelSplat style.
Acknowledgement
If you find this dataset useful in your research, please consider citing:
1️⃣ DL3DV Original Dataset:
@inproceedings{ling2024dl3dv,
title={Dl3dv-10k: A… See the full description on the dataset page: https://huggingface.co/datasets/littlekoyo/dl3dv_chunked.xd-violence-rgb-videomae-chunked-testIndian-Supreme-Court-Judgements-Chunked
Indian Supreme Court Judgements Chunked
Executive Summary
The dataset aims to address the chronic backlog in the Indian judiciary system, particularly in the Supreme Court, by creating a dataset optimized for legal language models (LLMs). The dataset will consist of pre-processed, chunked, and embedded textual data derived from the Supreme Court's judgment PDFs.
Problem and Importance - Motivation
Indian courts are overwhelmed with pending cases, with the… See the full description on the dataset page: https://huggingface.co/datasets/vihaannnn/Indian-Supreme-Court-Judgements-Chunked.nsw-caselaw-chunkedwiki-chunked-mxbai-embed-large-v1pg_books-tokenized-bos-eos-chunked-65536
Dataset Card for "pg_books-tokenized-bos-eos-chunked-65536"
The pg19 dataset tokenized under LLaMA into 64k chunks, bookended with BOS and EOS
wikipedia_chunked
Dataset Card for "wikipedia_chunked"
More Information needed
openiti_chunked
Description
This dataset is derived from the 2023.1.8 release of the OpenITI corpus and is intended to pretrain small language models with short context lengths (<2048 Unicode code points).
Processing
The markdown files were converted into raw text by stripping all code points neither classified as whitespace nor found in the Arabic Unicode code pages. Each document was then chunked by randomly sampling sequences of 2048 character length with a number of samples selected… See the full description on the dataset page: https://huggingface.co/datasets/mittagessen/openiti_chunked.nasle-mana-clean-chunked-30s
Nasl-e-Mana Clean Speech Corpus — Sentence-Safe 30s Chunks
Training-oriented WAV chunks derived from the public Nasl-e-Mana magazine audio corpus. Chunks target approximately 30 seconds and are cut at detected acoustic pauses; the labeled configuration additionally assigns only complete source-text sentences to each chunk.
Configuration
Rows
Columns
Meaning
labeled (train/)
9,886
audio, label
Sentence-grouped text/audio pairs from duration-compatible source-text… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/nasle-mana-clean-chunked-30s.nasle-mana-clean-chunked-30s-avasanj
Nasl-e-Mana Clean Persian Speech — corrected 30-second chunks
Corrected, provenance-preserving audio chunks collected from the Nasl-e-Mana magazine website, generated on 2026-08-30. This release supersedes the earlier unreliable proportional-mapping chunk export; that older release was not used here.
Splits
Split
Rows
Audio
Columns
labeled
4,981
41.41 hours
audio, label
to_transcribe
11,127
92.72 hours
audio
The labeled split contains the… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/nasle-mana-clean-chunked-30s-avasanj.SlimPajama-chunked
SlimPajama-Chunked
Dataset Description
This is a chunked re-upload of Cerebras' SlimPajama-627B. The original upload has split
the dataset into 10 chunks, with each containing upwards of 5,000 files. This makes it cumbersome to download and process. We've downloaded the entire
dataset for our own purposes, and decided to upload the chunked version for easier usage.
Each file is ~45GB due to HuggingFace's limitation of 50GB per LFS file.
fineweb-edu-sample-10BT-chunked-500-nomic-text-v1.5
FineWeb-edu 10BT Sample embedded with nomic-text-v1.5
The FineWeb-edu 10BT sample was first chunked into 500 tokens (using bert-base-uncased) with 10% overlap resulting in 25 million rows and 10.5BT.
The chunks were then embedded using nomic-text-v1.5.
Dataset Details
Dataset Sources
Repository: https://github.com/enjalot/fineweb-modal
Uses
Direct Use
The dataset was embedded with the clustering: prefix, so the main… See the full description on the dataset page: https://huggingface.co/datasets/enjalot/fineweb-edu-sample-10BT-chunked-500-nomic-text-v1.5.Earnings22-Cleaned-AA-chunked
Earnings22-Cleaned-AA-chunked
Quick links: AA Streaming Speech to Text Leaderboard | Speech to Text methodology
Earnings22-Cleaned-AA-chunked is a chunked version of Earnings22-Cleaned-AA, the cleaned Earnings-22 subset used by Artificial Analysis for streaming Speech to Text evaluation.
The original Earnings-22 data comes from esb/datasets, a corpus of corporate earnings calls. Artificial Analysis manually reviewed and corrected the reference transcripts in the cleaned subset… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/Earnings22-Cleaned-AA-chunked.so101_pick_cube_chunked
SO101 Pick Cube Dataset (Chunked)
This is a restructured version of the gpudad/so101_pick_cube dataset with episode-level video files for faster data loading during training.
Why Chunked?
The original dataset has 3 monolithic video files (one per camera, 13+ hours each). Random access during training is slow because the decoder must seek through huge files.
This version splits videos into 1000 episodes per chunk, making data loading ~50x faster.
Dataset Info… See the full description on the dataset page: https://huggingface.co/datasets/gpudad/so101_pick_cube_chunked.dolma-books-chunked-4kchunked-wikipedia20220301en-bookcorpusopen
Dataset Card for "chunked-wikipedia20220301en-bookcorpusopen"
num_examples: 33.5 million
download_size: 15.3 GB
dataset_size: 26.1 GB
This dataset combines wikipedia20220301.en and bookcorpusopen,
and splits the data into smaller chunks, of size ~820 chars
(such that each item will be at least ~128 tokens for the average tokenizer).
The logic only splits on spaces, so the chunks are likely to be slightly larger than 820 chars.
The dataset has been normalized into lower case… See the full description on the dataset page: https://huggingface.co/datasets/sradc/chunked-wikipedia20220301en-bookcorpusopen.arXiv-full-text-chunked
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]:… See the full description on the dataset page: https://huggingface.co/datasets/amrachraf/arXiv-full-text-chunked.audiobook_chunked_tts_train
audiobook_chunked_tts_train
This is a gated Uzbek TTS training dataset from instinct-org.
This repository contains tokenized or prepared speech data for text-to-speech training workflows.
Language
Primary language: uz (Uzbek)
Intended Use
text-to-speech training
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data Notes
Prepared for TTS training workflows.… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/audiobook_chunked_tts_train.audio_youtube_chunked_tts_train
audio_youtube_chunked_tts_train
This is a gated Uzbek TTS training dataset from instinct-org.
This repository contains tokenized or prepared speech data for text-to-speech training workflows.
Language
Primary language: uz (Uzbek)
Intended Use
text-to-speech training
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data Notes
Prepared for TTS training… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/audio_youtube_chunked_tts_train.amber-data-arxiv-chunked-360chunked-shuffled-wikipedia20220301en-bookcorpusopen
Dataset Card for "wikipedia20220301en-bookcorpusopen-chunked-shuffled"
num_examples: 33.5 million
download_size: 15.3 GB
dataset_size: 26.1 GB
This dataset combines wikipedia20220301.en and bookcorpusopen,
and splits the data into smaller chunks, of size ~820 chars
(such that each item will be at least ~128 tokens for the average tokenizer).
The order of the items in this dataset has been shuffled,
meaning you don't have to use dataset.shuffle,
which is slower to iterate over.… See the full description on the dataset page: https://huggingface.co/datasets/sradc/chunked-shuffled-wikipedia20220301en-bookcorpusopen.espeech_podcasts_chunked_tts_train
espeech_podcasts_chunked_tts_train
This is a gated Russian TTS training dataset from instinct-org.
This repository contains tokenized or prepared speech data for text-to-speech training workflows.
Language
Primary language: ru (Russian)
Intended Use
text-to-speech training
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data Notes
Prepared for TTS training… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/espeech_podcasts_chunked_tts_train.zy_chunked_tts_train
zy_chunked_tts_train
This is a gated Uzbek TTS training dataset from instinct-org.
This repository contains tokenized or prepared speech data for text-to-speech training workflows.
Language
Primary language: uz (Uzbek)
Intended Use
text-to-speech training
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data Notes
Prepared for TTS training workflows.
Derived… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/zy_chunked_tts_train.ganjoor-recitations-chunked
🗂️ ganjoor-recitations-chunked
English + فارسی · Part of Shenava 1.0 · Project hub · SLT paper submission
🌟 At a glance | معرفی سریع
English
فارسی
🎯 Purpose
Ganjoor recitation chunked ASR dataset.
قطعههای تلاوت و خوانش گنجور برای آموزش و ارزیابی گفتار ادبی، شعر و خوانش رسمی فارسی.
🧩 Role
Persian speech dataset
مجموعهدادهٔ گفتار فارسی
📦 Snapshot
64 files; approximately 118.09 GB
64 فایل؛ حدود 118.09 GB
🧱 Packaging
61 Parquet files and 0… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/ganjoor-recitations-chunked.wikipedia-longest-stride-chunked-500
Wikipedia-Longest-Stride-Chunked-500
This is the chunked version of Wikipedia HF Dataset.
{
"article_hash": "... hashes ...",
"language": 'en',
"chunks": ["chunks1", "chunks2", "chunks3"],
"num_chunks": 3
}
chunking strategy is in the scripts/chunking.py.
detailed description will be provided later. Please stay tuned!
All licence reserved to the original author.
phylo-dna-archaea-bacteria-0.2-tokenized-chunked-non-overlap-32768asr-farsi-youtube-chunked-10-secondstrain_group_theory_cpt_chunked_4096Chunked-Indian-Supreme-Court-Judgements
Indian Supreme Court Judgements Chunked
