CoolFace
8 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AlppAI /SlimPajama-chunked SlimPajama-Chunked Dataset Description This is a chunked re-upload of Cerebras' SlimPajama-627B. The original upload has split the dataset into 10 chunks, with each containing upwards of 5,000 files. This makes it cumbersome to download and process. We've downloaded the entire dataset for our own purposes, and decided to upload the chunked version for easier usage. Each file is ~45GB due to HuggingFace's limitation of 50GB per LFS file. texttext-generation1M<n<10M5 likes647 downloads3y agoHugging Face02amrachraf /arXiv-full-text-chunked Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More Information Needed] Paper [optional]:… See the full description on the dataset page: https://huggingface.co/datasets/amrachraf/arXiv-full-text-chunked.texttext-generation100K<n<1M1 likes447 downloads2y agoHugging Face03dme5245 /fineweb-edu-10B-lfm25-chunked-512 FineWeb-Edu 10BT — LFM2.5 chunked (512 tokens) Pre-tokenized, packed chunks of HuggingFaceFW/fineweb-edu sample-10BT, tokenized with the LFM2.5 tokenizer and cut into fixed 512-token sequences. What's in it Field Value Source HuggingFaceFW/fineweb-edu sample-10BT (train) Tokenizer LiquidAI/LFM2.5-1.2B-Base Chunk size 512 tokens Total chunks 19,300,635 Total tokens 9,881,925,120 Columns input_ids (int32, length 512) Format HF Arrow (mmap-able… See the full description on the dataset page: https://huggingface.co/datasets/dme5245/fineweb-edu-10B-lfm25-chunked-512.text-generation10B<n<100B1 likes215 downloads12d agoHugging Face04hoorangyee /pile-of-law-chunked LRAGE: Legal Retrieval Augmented Generation Evaluation Tool LRAGE (Legal Retrieval Augmented Generation Evaluation, pronounced as 'large') is an open-source toolkit designed to evaluate Large Language Models (LLMs) in a Retrieval-Augmented Generation (RAG) setting, specifically tailored for the legal domain. This repository contains pointers to datasets and code used in LRAGE: Legal Retrieval Augmented Generation Evaluation. Code: https://github.com/hoorangyee/LRAGE… See the full description on the dataset page: https://huggingface.co/datasets/hoorangyee/pile-of-law-chunked.texttext-generation10M<n<100M1 likes179 downloads1y agoHugging Face05meettilavat /InternetArchive_1899_Chunked Internet Archive Historical Texts - Chunked (0001-1899) TL;DR 163 million text chunks extracted from historical public-domain documents sourced from the Internet Archive Content dated 0001-1899, sorted by download popularity to prioritize high-quality, frequently accessed materials 2,445 Zstandard-compressed Parquet shards totaling ~217 GB on disk, ~594 billion characters uncompressed Optimized chunk size of ~3,600 characters (target: 4,000) for efficient language model… See the full description on the dataset page: https://huggingface.co/datasets/meettilavat/InternetArchive_1899_Chunked.texttext-generation100M<n<1B0 likes177 downloads11mo agoHugging Face06Laz4rz /wikipedia_science_chunked_small_rag_512 ScienceWikiSmallChunk Processed version of millawell/wikipedia_field_of_science, prepared to be used in small context length RAG systems. Chunk length is tokenizer dependent, but each chunk should be around 512 tokens. Longer wikipedia pages have been split into smaller entries, with title added as a prefix. There is also 256 tokens dataset available: Laz4rz/wikipedia_science_chunked_small_rag_256 If you wish to prepare some other chunk length: use… See the full description on the dataset page: https://huggingface.co/datasets/Laz4rz/wikipedia_science_chunked_small_rag_512.texttext-generation1M<n<10M4 likes40 downloads2y agoHugging Face07wickkiey /tamil-wikipedia-chunked Tamil Wikipedia Chunked Dataset Dataset Description This dataset contains Tamil Wikipedia articles that have been intelligently chunked using header-aware splitting. Each chunk includes both the text content and metadata about the document structure, making it ideal for training Large Language Models with better contextual understanding. Dataset Summary Language: Tamil (ta) Format: Parquet file with text and metadata columns Chunking Method: Header-based… See the full description on the dataset page: https://huggingface.co/datasets/wickkiey/tamil-wikipedia-chunked.texttext-generation100K<n<1M0 likes32 downloads8mo agoHugging Face08Laz4rz /wikipedia_science_chunked_small_rag_256 ScienceWikiSmallChunk Processed version of millawell/wikipedia_field_of_science, prepared to be used in small context length RAG systems. Chunk length is tokenizer dependent, but each chunk should be around 256 tokens. Longer wikipedia pages have been split into smaller entries, with title added as a prefix. There is also 512 tokens dataset available: Laz4rz/wikipedia_science_chunked_small_rag_512 If you wish to prepare some other chunk length: use… See the full description on the dataset page: https://huggingface.co/datasets/Laz4rz/wikipedia_science_chunked_small_rag_256.texttext-generation1M<n<10M3 likes28 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.