CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01enguyen /smollm-chunked FAISS Indices and Chunked Datasets for SmolLM and SmolLM2 corpora This repository contains part of the FAISS indices and chunked datasets used for novelty detection for SmolLM and SmolLM2, as presented in the paper LLM generation novelty through the lens of semantic similarity. Full Documentation For complete usage instructions, installation guide, and tutorial, please refer to: Main Tutorial README Data Distribution Due to Hugging Face storage quota… See the full description on the dataset page: https://huggingface.co/datasets/enguyen/smollm-chunked.tabulartext-retrieval100M<n<1B1 likes35k downloads7mo agoHugging Face02littlekoyo /dl3dv_chunked DL3DV Post-processed for Less3Depend This dataset is a post-processed version of the DL3DV-10K dataset, specifically prepared for the repository 👉 Less3Depend. The goal of this release is to provide a clean, unified, and research-ready variant of DL3DV that is directly usable in 👉 PixelSplat style. Acknowledgement If you find this dataset useful in your research, please consider citing: 1️⃣ DL3DV Original Dataset: @inproceedings{ling2024dl3dv, title={Dl3dv-10k: A… See the full description on the dataset page: https://huggingface.co/datasets/littlekoyo/dl3dv_chunked.1 likes3.2k downloads5mo agoHugging Face03KingTechnician /xd-violence-rgb-videomae-chunked-testtext100K<n<1M0 likes2.6k downloads8mo agoHugging Face04vihaannnn /Indian-Supreme-Court-Judgements-Chunked Indian Supreme Court Judgements Chunked Executive Summary The dataset aims to address the chronic backlog in the Indian judiciary system, particularly in the Supreme Court, by creating a dataset optimized for legal language models (LLMs). The dataset will consist of pre-processed, chunked, and embedded textual data derived from the Supreme Court's judgment PDFs. Problem and Importance - Motivation Indian courts are overwhelmed with pending cases, with the… See the full description on the dataset page: https://huggingface.co/datasets/vihaannnn/Indian-Supreme-Court-Judgements-Chunked.textfeature-extraction10K<n<100K6 likes2.5k downloads2y agoHugging Face05corto-ai /nsw-caselaw-chunkedtext10M<n<100M0 likes2k downloads2y agoHugging Face06anton-l /wiki-chunked-mxbai-embed-large-v1text1M<n<10M0 likes1.4k downloads3y agoHugging Face07emozilla /pg_books-tokenized-bos-eos-chunked-65536 Dataset Card for "pg_books-tokenized-bos-eos-chunked-65536" The pg19 dataset tokenized under LLaMA into 64k chunks, bookended with BOS and EOS 10K<n<100K7 likes1.3k downloads3y agoHugging Face08simple-pretraining /wikipedia_chunked Dataset Card for "wikipedia_chunked" More Information needed text10M<n<100M2 likes1.3k downloads3y agoHugging Face09mittagessen /openiti_chunked Description This dataset is derived from the 2023.1.8 release of the OpenITI corpus and is intended to pretrain small language models with short context lengths (<2048 Unicode code points). Processing The markdown files were converted into raw text by stripping all code points neither classified as whitespace nor found in the Arabic Unicode code pages. Each document was then chunked by randomly sampling sequences of 2048 character length with a number of samples selected… See the full description on the dataset page: https://huggingface.co/datasets/mittagessen/openiti_chunked.text10M<n<100M1 likes705 downloads1y agoHugging Face10Reza2kn /nasle-mana-clean-chunked-30s Nasl-e-Mana Clean Speech Corpus — Sentence-Safe 30s Chunks Training-oriented WAV chunks derived from the public Nasl-e-Mana magazine audio corpus. Chunks target approximately 30 seconds and are cut at detected acoustic pauses; the labeled configuration additionally assigns only complete source-text sentences to each chunk. Configuration Rows Columns Meaning labeled (train/) 9,886 audio, label Sentence-grouped text/audio pairs from duration-compatible source-text… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/nasle-mana-clean-chunked-30s.audioautomatic-speech-recognition1K<n<10K0 likes694 downloads25d agoHugging Face11Reza2kn /nasle-mana-clean-chunked-30s-avasanj Nasl-e-Mana Clean Persian Speech — corrected 30-second chunks Corrected, provenance-preserving audio chunks collected from the Nasl-e-Mana magazine website, generated on 2026-08-30. This release supersedes the earlier unreliable proportional-mapping chunk export; that older release was not used here. Splits Split Rows Audio Columns labeled 4,981 41.41 hours audio, label to_transcribe 11,127 92.72 hours audio The labeled split contains the… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/nasle-mana-clean-chunked-30s-avasanj.audioautomatic-speech-recognition10K<n<100K1 likes672 downloads17d agoHugging Face12AlppAI /SlimPajama-chunked SlimPajama-Chunked Dataset Description This is a chunked re-upload of Cerebras' SlimPajama-627B. The original upload has split the dataset into 10 chunks, with each containing upwards of 5,000 files. This makes it cumbersome to download and process. We've downloaded the entire dataset for our own purposes, and decided to upload the chunked version for easier usage. Each file is ~45GB due to HuggingFace's limitation of 50GB per LFS file. texttext-generation1M<n<10M5 likes647 downloads3y agoHugging Face13enjalot /fineweb-edu-sample-10BT-chunked-500-nomic-text-v1.5 FineWeb-edu 10BT Sample embedded with nomic-text-v1.5 The FineWeb-edu 10BT sample was first chunked into 500 tokens (using bert-base-uncased) with 10% overlap resulting in 25 million rows and 10.5BT. The chunks were then embedded using nomic-text-v1.5. Dataset Details Dataset Sources Repository: https://github.com/enjalot/fineweb-modal Uses Direct Use The dataset was embedded with the clustering: prefix, so the main… See the full description on the dataset page: https://huggingface.co/datasets/enjalot/fineweb-edu-sample-10BT-chunked-500-nomic-text-v1.5.tabular10M<n<100M5 likes627 downloads2y agoHugging Face14ArtificialAnalysis /Earnings22-Cleaned-AA-chunked Earnings22-Cleaned-AA-chunked Quick links: AA Streaming Speech to Text Leaderboard | Speech to Text methodology Earnings22-Cleaned-AA-chunked is a chunked version of Earnings22-Cleaned-AA, the cleaned Earnings-22 subset used by Artificial Analysis for streaming Speech to Text evaluation. The original Earnings-22 data comes from esb/datasets, a corpus of corporate earnings calls. Artificial Analysis manually reviewed and corrected the reference transcripts in the cleaned subset… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/Earnings22-Cleaned-AA-chunked.audioautomatic-speech-recognitionn<1K1 likes608 downloads3mo agoHugging Face15gpudad /so101_pick_cube_chunked SO101 Pick Cube Dataset (Chunked) This is a restructured version of the gpudad/so101_pick_cube dataset with episode-level video files for faster data loading during training. Why Chunked? The original dataset has 3 monolithic video files (one per camera, 13+ hours each). Random access during training is slow because the decoder must seek through huge files. This version splits videos into 1000 episodes per chunk, making data loading ~50x faster. Dataset Info… See the full description on the dataset page: https://huggingface.co/datasets/gpudad/so101_pick_cube_chunked.tabularrobotics1M<n<10M0 likes509 downloads8mo agoHugging Face16wonabru /dolma-books-chunked-4ktext1M<n<10M1 likes484 downloads1y agoHugging Face17sradc /chunked-wikipedia20220301en-bookcorpusopen Dataset Card for "chunked-wikipedia20220301en-bookcorpusopen" num_examples: 33.5 million download_size: 15.3 GB dataset_size: 26.1 GB This dataset combines wikipedia20220301.en and bookcorpusopen, and splits the data into smaller chunks, of size ~820 chars (such that each item will be at least ~128 tokens for the average tokenizer). The logic only splits on spaces, so the chunks are likely to be slightly larger than 820 chars. The dataset has been normalized into lower case… See the full description on the dataset page: https://huggingface.co/datasets/sradc/chunked-wikipedia20220301en-bookcorpusopen.text10M<n<100M0 likes470 downloads3y agoHugging Face18amrachraf /arXiv-full-text-chunked Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More Information Needed] Paper [optional]:… See the full description on the dataset page: https://huggingface.co/datasets/amrachraf/arXiv-full-text-chunked.texttext-generation100K<n<1M1 likes447 downloads2y agoHugging Face19instinct-org /audiobook_chunked_tts_traingated audiobook_chunked_tts_train This is a gated Uzbek TTS training dataset from instinct-org. This repository contains tokenized or prepared speech data for text-to-speech training workflows. Language Primary language: uz (Uzbek) Intended Use text-to-speech training Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license review Data Notes Prepared for TTS training workflows.… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/audiobook_chunked_tts_train.audiotext-to-speech0 likes439 downloads4mo agoHugging Face20instinct-org /audio_youtube_chunked_tts_traingated audio_youtube_chunked_tts_train This is a gated Uzbek TTS training dataset from instinct-org. This repository contains tokenized or prepared speech data for text-to-speech training workflows. Language Primary language: uz (Uzbek) Intended Use text-to-speech training Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license review Data Notes Prepared for TTS training… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/audio_youtube_chunked_tts_train.audiotext-to-speech0 likes432 downloads4mo agoHugging Face21qbwmwsap /amber-data-arxiv-chunked-360text10K<n<100K0 likes417 downloads2y agoHugging Face22sradc /chunked-shuffled-wikipedia20220301en-bookcorpusopen Dataset Card for "wikipedia20220301en-bookcorpusopen-chunked-shuffled" num_examples: 33.5 million download_size: 15.3 GB dataset_size: 26.1 GB This dataset combines wikipedia20220301.en and bookcorpusopen, and splits the data into smaller chunks, of size ~820 chars (such that each item will be at least ~128 tokens for the average tokenizer). The order of the items in this dataset has been shuffled, meaning you don't have to use dataset.shuffle, which is slower to iterate over.… See the full description on the dataset page: https://huggingface.co/datasets/sradc/chunked-shuffled-wikipedia20220301en-bookcorpusopen.text10M<n<100M4 likes404 downloads3y agoHugging Face23instinct-org /espeech_podcasts_chunked_tts_traingated espeech_podcasts_chunked_tts_train This is a gated Russian TTS training dataset from instinct-org. This repository contains tokenized or prepared speech data for text-to-speech training workflows. Language Primary language: ru (Russian) Intended Use text-to-speech training Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license review Data Notes Prepared for TTS training… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/espeech_podcasts_chunked_tts_train.audiotext-to-speech0 likes362 downloads4mo agoHugging Face24instinct-org /zy_chunked_tts_traingated zy_chunked_tts_train This is a gated Uzbek TTS training dataset from instinct-org. This repository contains tokenized or prepared speech data for text-to-speech training workflows. Language Primary language: uz (Uzbek) Intended Use text-to-speech training Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license review Data Notes Prepared for TTS training workflows. Derived… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/zy_chunked_tts_train.audiotext-to-speech0 likes317 downloads4mo agoHugging Face25Reza2kn /ganjoor-recitations-chunked 🗂️ ganjoor-recitations-chunked English + فارسی · Part of Shenava 1.0 · Project hub · SLT paper submission 🌟 At a glance | معرفی سریع English فارسی 🎯 Purpose Ganjoor recitation chunked ASR dataset. قطعه‌های تلاوت و خوانش گنجور برای آموزش و ارزیابی گفتار ادبی، شعر و خوانش رسمی فارسی. 🧩 Role Persian speech dataset مجموعه‌دادهٔ گفتار فارسی 📦 Snapshot 64 files; approximately 118.09 GB 64 فایل؛ حدود 118.09 GB 🧱 Packaging 61 Parquet files and 0… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/ganjoor-recitations-chunked.audioautomatic-speech-recognition100K<n<1M1 likes307 downloads2mo agoHugging Face26enzoescipy /wikipedia-longest-stride-chunked-500 Wikipedia-Longest-Stride-Chunked-500 This is the chunked version of Wikipedia HF Dataset. { "article_hash": "... hashes ...", "language": 'en', "chunks": ["chunks1", "chunks2", "chunks3"], "num_chunks": 3 } chunking strategy is in the scripts/chunking.py. detailed description will be provided later. Please stay tuned! All licence reserved to the original author. text100K<n<1M0 likes300 downloads6mo agoHugging Face27wanglab /phylo-dna-archaea-bacteria-0.2-tokenized-chunked-non-overlap-327680 likes292 downloads2y agoHugging Face28pourmand1376 /asr-farsi-youtube-chunked-10-secondsaudio100K<n<1M10 likes290 downloads3y agoHugging Face29DopeorNope /train_group_theory_cpt_chunked_4096text10M<n<100M0 likes278 downloads1y agoHugging Face30vihaannnn /Chunked-Indian-Supreme-Court-Judgements Indian Supreme Court Judgements Chunked texttoken-classification10K<n<100K1 likes275 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.