datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SlimPajama-chunked
SlimPajama-Chunked
Dataset Description
This is a chunked re-upload of Cerebras' SlimPajama-627B. The original upload has split
the dataset into 10 chunks, with each containing upwards of 5,000 files. This makes it cumbersome to download and process. We've downloaded the entire
dataset for our own purposes, and decided to upload the chunked version for easier usage.
Each file is ~45GB due to HuggingFace's limitation of 50GB per LFS file.
arXiv-full-text-chunked
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]:… See the full description on the dataset page: https://huggingface.co/datasets/amrachraf/arXiv-full-text-chunked.fineweb-edu-10B-lfm25-chunked-512
FineWeb-Edu 10BT — LFM2.5 chunked (512 tokens)
Pre-tokenized, packed chunks of HuggingFaceFW/fineweb-edu sample-10BT, tokenized with the LFM2.5 tokenizer and cut into fixed 512-token sequences.
What's in it
Field
Value
Source
HuggingFaceFW/fineweb-edu sample-10BT (train)
Tokenizer
LiquidAI/LFM2.5-1.2B-Base
Chunk size
512 tokens
Total chunks
19,300,635
Total tokens
9,881,925,120
Columns
input_ids (int32, length 512)
Format
HF Arrow (mmap-able… See the full description on the dataset page: https://huggingface.co/datasets/dme5245/fineweb-edu-10B-lfm25-chunked-512.pile-of-law-chunked
LRAGE: Legal Retrieval Augmented Generation Evaluation Tool
LRAGE (Legal Retrieval Augmented Generation Evaluation, pronounced as 'large') is an open-source toolkit designed to evaluate Large Language Models (LLMs) in a Retrieval-Augmented Generation (RAG) setting, specifically tailored for the legal domain.
This repository contains pointers to datasets and code used in LRAGE: Legal Retrieval Augmented Generation Evaluation.
Code: https://github.com/hoorangyee/LRAGE… See the full description on the dataset page: https://huggingface.co/datasets/hoorangyee/pile-of-law-chunked.InternetArchive_1899_Chunked
Internet Archive Historical Texts - Chunked (0001-1899)
TL;DR
163 million text chunks extracted from historical public-domain documents sourced from the Internet Archive
Content dated 0001-1899, sorted by download popularity to prioritize high-quality, frequently accessed materials
2,445 Zstandard-compressed Parquet shards totaling ~217 GB on disk, ~594 billion characters uncompressed
Optimized chunk size of ~3,600 characters (target: 4,000) for efficient language model… See the full description on the dataset page: https://huggingface.co/datasets/meettilavat/InternetArchive_1899_Chunked.wikipedia_science_chunked_small_rag_512
ScienceWikiSmallChunk
Processed version of millawell/wikipedia_field_of_science, prepared to be used in small context length RAG systems. Chunk length is tokenizer dependent, but each chunk should be around 512 tokens. Longer wikipedia pages have been split into smaller entries, with title added as a prefix.
There is also 256 tokens dataset available: Laz4rz/wikipedia_science_chunked_small_rag_256
If you wish to prepare some other chunk length:
use… See the full description on the dataset page: https://huggingface.co/datasets/Laz4rz/wikipedia_science_chunked_small_rag_512.tamil-wikipedia-chunked
Tamil Wikipedia Chunked Dataset
Dataset Description
This dataset contains Tamil Wikipedia articles that have been intelligently chunked using header-aware splitting. Each chunk includes both the text content and metadata about the document structure, making it ideal for training Large Language Models with better contextual understanding.
Dataset Summary
Language: Tamil (ta)
Format: Parquet file with text and metadata columns
Chunking Method: Header-based… See the full description on the dataset page: https://huggingface.co/datasets/wickkiey/tamil-wikipedia-chunked.wikipedia_science_chunked_small_rag_256
ScienceWikiSmallChunk
Processed version of millawell/wikipedia_field_of_science, prepared to be used in small context length RAG systems. Chunk length is tokenizer dependent, but each chunk should be around 256 tokens. Longer wikipedia pages have been split into smaller entries, with title added as a prefix.
There is also 512 tokens dataset available: Laz4rz/wikipedia_science_chunked_small_rag_512
If you wish to prepare some other chunk length:
use… See the full description on the dataset page: https://huggingface.co/datasets/Laz4rz/wikipedia_science_chunked_small_rag_256.
