CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01enguyen /smollm-chunked FAISS Indices and Chunked Datasets for SmolLM and SmolLM2 corpora This repository contains part of the FAISS indices and chunked datasets used for novelty detection for SmolLM and SmolLM2, as presented in the paper LLM generation novelty through the lens of semantic similarity. Full Documentation For complete usage instructions, installation guide, and tutorial, please refer to: Main Tutorial README Data Distribution Due to Hugging Face storage quota… See the full description on the dataset page: https://huggingface.co/datasets/enguyen/smollm-chunked.tabulartext-retrieval100M<n<1B1 likes35k downloads7mo agoHugging Face02KingTechnician /xd-violence-rgb-videomae-chunked-testtext100K<n<1M0 likes2.6k downloads8mo agoHugging Face03vihaannnn /Indian-Supreme-Court-Judgements-Chunked Indian Supreme Court Judgements Chunked Executive Summary The dataset aims to address the chronic backlog in the Indian judiciary system, particularly in the Supreme Court, by creating a dataset optimized for legal language models (LLMs). The dataset will consist of pre-processed, chunked, and embedded textual data derived from the Supreme Court's judgment PDFs. Problem and Importance - Motivation Indian courts are overwhelmed with pending cases, with the… See the full description on the dataset page: https://huggingface.co/datasets/vihaannnn/Indian-Supreme-Court-Judgements-Chunked.textfeature-extraction10K<n<100K6 likes2.5k downloads2y agoHugging Face04corto-ai /nsw-caselaw-chunkedtext10M<n<100M0 likes2k downloads2y agoHugging Face05anton-l /wiki-chunked-mxbai-embed-large-v1text1M<n<10M0 likes1.4k downloads3y agoHugging Face06simple-pretraining /wikipedia_chunked Dataset Card for "wikipedia_chunked" More Information needed text10M<n<100M2 likes1.3k downloads3y agoHugging Face07mittagessen /openiti_chunked Description This dataset is derived from the 2023.1.8 release of the OpenITI corpus and is intended to pretrain small language models with short context lengths (<2048 Unicode code points). Processing The markdown files were converted into raw text by stripping all code points neither classified as whitespace nor found in the Arabic Unicode code pages. Each document was then chunked by randomly sampling sequences of 2048 character length with a number of samples selected… See the full description on the dataset page: https://huggingface.co/datasets/mittagessen/openiti_chunked.text10M<n<100M1 likes705 downloads1y agoHugging Face08Reza2kn /nasle-mana-clean-chunked-30s Nasl-e-Mana Clean Speech Corpus — Sentence-Safe 30s Chunks Training-oriented WAV chunks derived from the public Nasl-e-Mana magazine audio corpus. Chunks target approximately 30 seconds and are cut at detected acoustic pauses; the labeled configuration additionally assigns only complete source-text sentences to each chunk. Configuration Rows Columns Meaning labeled (train/) 9,886 audio, label Sentence-grouped text/audio pairs from duration-compatible source-text… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/nasle-mana-clean-chunked-30s.audioautomatic-speech-recognition1K<n<10K0 likes694 downloads25d agoHugging Face09Reza2kn /nasle-mana-clean-chunked-30s-avasanj Nasl-e-Mana Clean Persian Speech — corrected 30-second chunks Corrected, provenance-preserving audio chunks collected from the Nasl-e-Mana magazine website, generated on 2026-08-30. This release supersedes the earlier unreliable proportional-mapping chunk export; that older release was not used here. Splits Split Rows Audio Columns labeled 4,981 41.41 hours audio, label to_transcribe 11,127 92.72 hours audio The labeled split contains the… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/nasle-mana-clean-chunked-30s-avasanj.audioautomatic-speech-recognition10K<n<100K1 likes672 downloads17d agoHugging Face10AlppAI /SlimPajama-chunked SlimPajama-Chunked Dataset Description This is a chunked re-upload of Cerebras' SlimPajama-627B. The original upload has split the dataset into 10 chunks, with each containing upwards of 5,000 files. This makes it cumbersome to download and process. We've downloaded the entire dataset for our own purposes, and decided to upload the chunked version for easier usage. Each file is ~45GB due to HuggingFace's limitation of 50GB per LFS file. texttext-generation1M<n<10M5 likes647 downloads3y agoHugging Face11enjalot /fineweb-edu-sample-10BT-chunked-500-nomic-text-v1.5 FineWeb-edu 10BT Sample embedded with nomic-text-v1.5 The FineWeb-edu 10BT sample was first chunked into 500 tokens (using bert-base-uncased) with 10% overlap resulting in 25 million rows and 10.5BT. The chunks were then embedded using nomic-text-v1.5. Dataset Details Dataset Sources Repository: https://github.com/enjalot/fineweb-modal Uses Direct Use The dataset was embedded with the clustering: prefix, so the main… See the full description on the dataset page: https://huggingface.co/datasets/enjalot/fineweb-edu-sample-10BT-chunked-500-nomic-text-v1.5.tabular10M<n<100M5 likes627 downloads2y agoHugging Face12ArtificialAnalysis /Earnings22-Cleaned-AA-chunked Earnings22-Cleaned-AA-chunked Quick links: AA Streaming Speech to Text Leaderboard | Speech to Text methodology Earnings22-Cleaned-AA-chunked is a chunked version of Earnings22-Cleaned-AA, the cleaned Earnings-22 subset used by Artificial Analysis for streaming Speech to Text evaluation. The original Earnings-22 data comes from esb/datasets, a corpus of corporate earnings calls. Artificial Analysis manually reviewed and corrected the reference transcripts in the cleaned subset… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/Earnings22-Cleaned-AA-chunked.audioautomatic-speech-recognitionn<1K1 likes608 downloads3mo agoHugging Face13wonabru /dolma-books-chunked-4ktext1M<n<10M1 likes484 downloads1y agoHugging Face14sradc /chunked-wikipedia20220301en-bookcorpusopen Dataset Card for "chunked-wikipedia20220301en-bookcorpusopen" num_examples: 33.5 million download_size: 15.3 GB dataset_size: 26.1 GB This dataset combines wikipedia20220301.en and bookcorpusopen, and splits the data into smaller chunks, of size ~820 chars (such that each item will be at least ~128 tokens for the average tokenizer). The logic only splits on spaces, so the chunks are likely to be slightly larger than 820 chars. The dataset has been normalized into lower case… See the full description on the dataset page: https://huggingface.co/datasets/sradc/chunked-wikipedia20220301en-bookcorpusopen.text10M<n<100M0 likes470 downloads3y agoHugging Face15amrachraf /arXiv-full-text-chunked Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More Information Needed] Paper [optional]:… See the full description on the dataset page: https://huggingface.co/datasets/amrachraf/arXiv-full-text-chunked.texttext-generation100K<n<1M1 likes447 downloads2y agoHugging Face16qbwmwsap /amber-data-arxiv-chunked-360text10K<n<100K0 likes417 downloads2y agoHugging Face17sradc /chunked-shuffled-wikipedia20220301en-bookcorpusopen Dataset Card for "wikipedia20220301en-bookcorpusopen-chunked-shuffled" num_examples: 33.5 million download_size: 15.3 GB dataset_size: 26.1 GB This dataset combines wikipedia20220301.en and bookcorpusopen, and splits the data into smaller chunks, of size ~820 chars (such that each item will be at least ~128 tokens for the average tokenizer). The order of the items in this dataset has been shuffled, meaning you don't have to use dataset.shuffle, which is slower to iterate over.… See the full description on the dataset page: https://huggingface.co/datasets/sradc/chunked-shuffled-wikipedia20220301en-bookcorpusopen.text10M<n<100M4 likes404 downloads3y agoHugging Face18Reza2kn /ganjoor-recitations-chunked 🗂️ ganjoor-recitations-chunked English + فارسی · Part of Shenava 1.0 · Project hub · SLT paper submission 🌟 At a glance | معرفی سریع English فارسی 🎯 Purpose Ganjoor recitation chunked ASR dataset. قطعه‌های تلاوت و خوانش گنجور برای آموزش و ارزیابی گفتار ادبی، شعر و خوانش رسمی فارسی. 🧩 Role Persian speech dataset مجموعه‌دادهٔ گفتار فارسی 📦 Snapshot 64 files; approximately 118.09 GB 64 فایل؛ حدود 118.09 GB 🧱 Packaging 61 Parquet files and 0… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/ganjoor-recitations-chunked.audioautomatic-speech-recognition100K<n<1M1 likes307 downloads2mo agoHugging Face19enzoescipy /wikipedia-longest-stride-chunked-500 Wikipedia-Longest-Stride-Chunked-500 This is the chunked version of Wikipedia HF Dataset. { "article_hash": "... hashes ...", "language": 'en', "chunks": ["chunks1", "chunks2", "chunks3"], "num_chunks": 3 } chunking strategy is in the scripts/chunking.py. detailed description will be provided later. Please stay tuned! All licence reserved to the original author. text100K<n<1M0 likes300 downloads6mo agoHugging Face20pourmand1376 /asr-farsi-youtube-chunked-10-secondsaudio100K<n<1M10 likes290 downloads3y agoHugging Face21DopeorNope /train_group_theory_cpt_chunked_4096text10M<n<100M0 likes278 downloads1y agoHugging Face22vihaannnn /Chunked-Indian-Supreme-Court-Judgements Indian Supreme Court Judgements Chunked texttoken-classification10K<n<100K1 likes275 downloads2y agoHugging Face23pourmand1376 /asr-farsi-youtube-chunked-30-seconds How To Use from datasets import load_dataset train = load_dataset('pourmand1376/asr-farsi-youtube-chunked-30-seconds', split='train+val') test =load_dataset('pourmand1376/asr-farsi-youtube-chunked-30-seconds', split='test') +300 Hours ASR dataset generated from this kaggle dataset audioautomatic-speech-recognition10K<n<100K11 likes264 downloads3y agoHugging Face24ibm-aimc /bookcorpus-wikitext-ccnews-tinystories-chunkedtext1M<n<10M1 likes231 downloads2y agoHugging Face25ibm-aimc /bookcorpus-wikitext-ccnews-sometinystories-chunkedtext1M<n<10M0 likes226 downloads2y agoHugging Face26jamescalam /ai-arxiv-chunkedtext10K<n<100K40 likes217 downloads3y agoHugging Face27Subhav-K /cnn-dailymail-chunked-512-embeddingstext100K<n<1M2 likes213 downloads3mo agoHugging Face28KingTechnician /xd-violence-20pct-audio-ast-chunked-traintext100K<n<1M0 likes211 downloads8mo agoHugging Face29MongoDB /cosmopedia-wikihow-chunked Overview This dataset is a chunked version of a subset of data in the Cosmopedia dataset curated by Hugging Face. Specifically, we have only used a subset of Wikihow articles from the Cosmopedia dataset, and each article has been split into chunks containing no more than 2 paragraphs. Dataset Structure Each record in the dataset represents a chunk of a larger article, and contains the following fields: doc_id: A unique identifier for the parent article chunk_id: A unique… See the full description on the dataset page: https://huggingface.co/datasets/MongoDB/cosmopedia-wikihow-chunked.tabularquestion-answering1M<n<10M9 likes208 downloads3y agoHugging Face30KingTechnician /xd-violence-i3d-flow-chunked-testtext100K<n<1M0 likes203 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.