datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
smollm-chunked
FAISS Indices and Chunked Datasets for SmolLM and SmolLM2 corpora
This repository contains part of the FAISS indices and chunked datasets used for novelty detection for SmolLM and SmolLM2, as presented in the paper LLM generation novelty through the lens of semantic similarity.
Full Documentation
For complete usage instructions, installation guide, and tutorial, please refer to:
Main Tutorial README
Data Distribution
Due to Hugging Face storage quota… See the full description on the dataset page: https://huggingface.co/datasets/enguyen/smollm-chunked.xd-violence-rgb-videomae-chunked-testIndian-Supreme-Court-Judgements-Chunked
Indian Supreme Court Judgements Chunked
Executive Summary
The dataset aims to address the chronic backlog in the Indian judiciary system, particularly in the Supreme Court, by creating a dataset optimized for legal language models (LLMs). The dataset will consist of pre-processed, chunked, and embedded textual data derived from the Supreme Court's judgment PDFs.
Problem and Importance - Motivation
Indian courts are overwhelmed with pending cases, with the… See the full description on the dataset page: https://huggingface.co/datasets/vihaannnn/Indian-Supreme-Court-Judgements-Chunked.nsw-caselaw-chunkedwiki-chunked-mxbai-embed-large-v1wikipedia_chunked
Dataset Card for "wikipedia_chunked"
More Information needed
openiti_chunked
Description
This dataset is derived from the 2023.1.8 release of the OpenITI corpus and is intended to pretrain small language models with short context lengths (<2048 Unicode code points).
Processing
The markdown files were converted into raw text by stripping all code points neither classified as whitespace nor found in the Arabic Unicode code pages. Each document was then chunked by randomly sampling sequences of 2048 character length with a number of samples selected… See the full description on the dataset page: https://huggingface.co/datasets/mittagessen/openiti_chunked.nasle-mana-clean-chunked-30s
Nasl-e-Mana Clean Speech Corpus — Sentence-Safe 30s Chunks
Training-oriented WAV chunks derived from the public Nasl-e-Mana magazine audio corpus. Chunks target approximately 30 seconds and are cut at detected acoustic pauses; the labeled configuration additionally assigns only complete source-text sentences to each chunk.
Configuration
Rows
Columns
Meaning
labeled (train/)
9,886
audio, label
Sentence-grouped text/audio pairs from duration-compatible source-text… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/nasle-mana-clean-chunked-30s.nasle-mana-clean-chunked-30s-avasanj
Nasl-e-Mana Clean Persian Speech — corrected 30-second chunks
Corrected, provenance-preserving audio chunks collected from the Nasl-e-Mana magazine website, generated on 2026-08-30. This release supersedes the earlier unreliable proportional-mapping chunk export; that older release was not used here.
Splits
Split
Rows
Audio
Columns
labeled
4,981
41.41 hours
audio, label
to_transcribe
11,127
92.72 hours
audio
The labeled split contains the… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/nasle-mana-clean-chunked-30s-avasanj.SlimPajama-chunked
SlimPajama-Chunked
Dataset Description
This is a chunked re-upload of Cerebras' SlimPajama-627B. The original upload has split
the dataset into 10 chunks, with each containing upwards of 5,000 files. This makes it cumbersome to download and process. We've downloaded the entire
dataset for our own purposes, and decided to upload the chunked version for easier usage.
Each file is ~45GB due to HuggingFace's limitation of 50GB per LFS file.
fineweb-edu-sample-10BT-chunked-500-nomic-text-v1.5
FineWeb-edu 10BT Sample embedded with nomic-text-v1.5
The FineWeb-edu 10BT sample was first chunked into 500 tokens (using bert-base-uncased) with 10% overlap resulting in 25 million rows and 10.5BT.
The chunks were then embedded using nomic-text-v1.5.
Dataset Details
Dataset Sources
Repository: https://github.com/enjalot/fineweb-modal
Uses
Direct Use
The dataset was embedded with the clustering: prefix, so the main… See the full description on the dataset page: https://huggingface.co/datasets/enjalot/fineweb-edu-sample-10BT-chunked-500-nomic-text-v1.5.Earnings22-Cleaned-AA-chunked
Earnings22-Cleaned-AA-chunked
Quick links: AA Streaming Speech to Text Leaderboard | Speech to Text methodology
Earnings22-Cleaned-AA-chunked is a chunked version of Earnings22-Cleaned-AA, the cleaned Earnings-22 subset used by Artificial Analysis for streaming Speech to Text evaluation.
The original Earnings-22 data comes from esb/datasets, a corpus of corporate earnings calls. Artificial Analysis manually reviewed and corrected the reference transcripts in the cleaned subset… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/Earnings22-Cleaned-AA-chunked.dolma-books-chunked-4kchunked-wikipedia20220301en-bookcorpusopen
Dataset Card for "chunked-wikipedia20220301en-bookcorpusopen"
num_examples: 33.5 million
download_size: 15.3 GB
dataset_size: 26.1 GB
This dataset combines wikipedia20220301.en and bookcorpusopen,
and splits the data into smaller chunks, of size ~820 chars
(such that each item will be at least ~128 tokens for the average tokenizer).
The logic only splits on spaces, so the chunks are likely to be slightly larger than 820 chars.
The dataset has been normalized into lower case… See the full description on the dataset page: https://huggingface.co/datasets/sradc/chunked-wikipedia20220301en-bookcorpusopen.arXiv-full-text-chunked
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]:… See the full description on the dataset page: https://huggingface.co/datasets/amrachraf/arXiv-full-text-chunked.amber-data-arxiv-chunked-360chunked-shuffled-wikipedia20220301en-bookcorpusopen
Dataset Card for "wikipedia20220301en-bookcorpusopen-chunked-shuffled"
num_examples: 33.5 million
download_size: 15.3 GB
dataset_size: 26.1 GB
This dataset combines wikipedia20220301.en and bookcorpusopen,
and splits the data into smaller chunks, of size ~820 chars
(such that each item will be at least ~128 tokens for the average tokenizer).
The order of the items in this dataset has been shuffled,
meaning you don't have to use dataset.shuffle,
which is slower to iterate over.… See the full description on the dataset page: https://huggingface.co/datasets/sradc/chunked-shuffled-wikipedia20220301en-bookcorpusopen.ganjoor-recitations-chunked
🗂️ ganjoor-recitations-chunked
English + فارسی · Part of Shenava 1.0 · Project hub · SLT paper submission
🌟 At a glance | معرفی سریع
English
فارسی
🎯 Purpose
Ganjoor recitation chunked ASR dataset.
قطعههای تلاوت و خوانش گنجور برای آموزش و ارزیابی گفتار ادبی، شعر و خوانش رسمی فارسی.
🧩 Role
Persian speech dataset
مجموعهدادهٔ گفتار فارسی
📦 Snapshot
64 files; approximately 118.09 GB
64 فایل؛ حدود 118.09 GB
🧱 Packaging
61 Parquet files and 0… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/ganjoor-recitations-chunked.wikipedia-longest-stride-chunked-500
Wikipedia-Longest-Stride-Chunked-500
This is the chunked version of Wikipedia HF Dataset.
{
"article_hash": "... hashes ...",
"language": 'en',
"chunks": ["chunks1", "chunks2", "chunks3"],
"num_chunks": 3
}
chunking strategy is in the scripts/chunking.py.
detailed description will be provided later. Please stay tuned!
All licence reserved to the original author.
asr-farsi-youtube-chunked-10-secondstrain_group_theory_cpt_chunked_4096Chunked-Indian-Supreme-Court-Judgements
Indian Supreme Court Judgements Chunked
asr-farsi-youtube-chunked-30-seconds
How To Use
from datasets import load_dataset
train = load_dataset('pourmand1376/asr-farsi-youtube-chunked-30-seconds', split='train+val')
test =load_dataset('pourmand1376/asr-farsi-youtube-chunked-30-seconds', split='test')
+300 Hours ASR dataset generated from this kaggle dataset
bookcorpus-wikitext-ccnews-tinystories-chunkedbookcorpus-wikitext-ccnews-sometinystories-chunkedai-arxiv-chunkedcnn-dailymail-chunked-512-embeddingsxd-violence-20pct-audio-ast-chunked-traincosmopedia-wikihow-chunked
Overview
This dataset is a chunked version of a subset of data in the Cosmopedia dataset curated by Hugging Face.
Specifically, we have only used a subset of Wikihow articles from the Cosmopedia dataset, and each article has been split into chunks containing no more than 2 paragraphs.
Dataset Structure
Each record in the dataset represents a chunk of a larger article, and contains the following fields:
doc_id: A unique identifier for the parent article
chunk_id: A unique… See the full description on the dataset page: https://huggingface.co/datasets/MongoDB/cosmopedia-wikihow-chunked.xd-violence-i3d-flow-chunked-test
