CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01k9cli /video-vec2wav2-tokenizer video-vec2wav2-tokenizer Production-ready pipeline (Python package video_vec2wav2_tokenizer, CLI command video2dataset) that turns a folder of videos into clean AI training datasets for speech recognition (ASR) and text-to-speech (TTS). videos ──► audio (16 kHz mono PCM) ──► whisper transcript ──► clips ──► metadata.csv / dataset.jsonl / tts_metadata.csv / report.json Video processing — recursive scan of mp4 / mkv / avi / mov / webm, FFmpeg audio extraction to mono ·… See the full description on the dataset page: https://huggingface.co/datasets/k9cli/video-vec2wav2-tokenizer.26 likes712k downloads1d agoHugging Face02jat-project /jat-dataset-tokenized Dataset Card for "jat-dataset-tokenized" More Information needed timeseries10M<n<100M32 likes443k downloads3y agoHugging Face03k9cli /video-vec2wav2-tokenizer-2 video-vec2wav2-tokenizer-2 Version 2 - continuation shard of the video-to-AI-dataset tokenizer project. Production-ready pipeline (Python package video_vec2wav2_tokenizer, CLI command video2dataset) that turns a folder of videos into clean AI training datasets for speech recognition (ASR) and text-to-speech (TTS). videos ──► audio (16 kHz mono PCM) ──► whisper transcript ──► clips ──► metadata.csv / dataset.jsonl / tts_metadata.csv / report.json Video processing —… See the full description on the dataset page: https://huggingface.co/datasets/k9cli/video-vec2wav2-tokenizer-2.17 likes426k downloads2mo agoHugging Face04k9cli /video-vec2wav2-tokenizer-3 video-vec2wav2-tokenizer-3 Version 3 - continuation shard of the video-to-AI-dataset tokenizer project. Version 2 - continuation shard of the video-to-AI-dataset tokenizer project. Production-ready pipeline (Python package video_vec2wav2_tokenizer, CLI command video2dataset) that turns a folder of videos into clean AI training datasets for speech recognition (ASR) and text-to-speech (TTS). videos ──► audio (16 kHz mono PCM) ──► whisper transcript ──► clips ──►… See the full description on the dataset page: https://huggingface.co/datasets/k9cli/video-vec2wav2-tokenizer-3.1 likes193k downloads2mo agoHugging Face05occiglot /tokenizer-wiki-bench Multilingual Tokenizer Benchmark This dataset includes pre-processed wikipedia data for tokenizer evaluation in 45 languages. We provide more information on the evaluation task in general this blogpost. Usage The dataset allows us to easily calculate tokenizer fertility and the proportion of continued words on any of the supported languages. In the example below we take the Mistral tokenizer and evaluate its performance on Slovak. from transformers import AutoTokenizer… See the full description on the dataset page: https://huggingface.co/datasets/occiglot/tokenizer-wiki-bench.text10M<n<100M6 likes47k downloads2y agoHugging Face06jinofy-corp /jora_corpus1_tokenized_128ktabularn<1K4 likes41k downloads2mo agoHugging Face07anisoleai /fineweb-tokenized FineWeb Tokenized > 4 trillion tokens of the pre-tokenized data the 🌐 web has to offer What is it? This is a pre-tokenized version of the HuggingFaceFW/fineweb dataset (currently in-progress, tokenization of the ~15 trillion tokens corpus is ongoing). The data is being pre-processed and tokenized using the AnisoleAI BPE tokenizer (52,022 vocabulary size) and packed into compact uint16 Parquet shards. By distributing the pre-tokenized corpus, we eliminate… See the full description on the dataset page: https://huggingface.co/datasets/anisoleai/fineweb-tokenized.tabulartext-generationn>1T34 likes38k downloads4mo agoHugging Face08edbeeching /gia-dataset-tokenized-2024-2 Dataset Card for "gia-dataset-tokenized-2024-2" More Information needed 100K<n<1M0 likes37k downloads3y agoHugging Face09hf-internal-testing /tokenizers-test-data tokenizers-test-data Test and benchmark fixtures for huggingface/tokenizers, pulled on demand by the repo Makefiles (make test / make bench / make fixtures via hf download). Layout fixtures/ — multilingual + modality corpora for cross-language encode benchmarks. Organized, documented, and reproducible: see fixtures/FIXTURES.md for provenance and fixtures/fixtures_manifest.json for exact sources, pinned revisions, and sizes. Rebuild any file with… See the full description on the dataset page: https://huggingface.co/datasets/hf-internal-testing/tokenizers-test-data.textn<1K0 likes25k downloads14d agoHugging Face10hf-internal-testing /tokenizers-benchimage1K<n<10K0 likes19k downloads8d agoHugging Face11regent-project /regent-subset-of-jat-dataset-tokenizedtimeseries10M<n<100M0 likes18k downloads2y agoHugging Face12QuangDuy /FineWeb2-mds-tokenized-v20 likes16k downloads9mo agoHugging Face13tokyotech-llm /swallow-math-v2 SwallowMath-v2 Resources 📑 arXiv: Read our paper for detailed methodology at arXiv:2505.02881. 🤗 Sister Dataset: Discover SwallowCode2, our companion dataset for code generation. 🧮 What is it? SwallowMath-v2 is a large-scale mathematical dataset containing 32 billion tokens, developed as the successor to SwallowMath-v1. Building on the success of v1, this release aims to construct a larger-scale and more permissively licensed corpus to support open and… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/swallow-math-v2.texttext-generation10M<n<100M35 likes13k downloads11mo agoHugging Face14QuangDuy /FineWeb2-mds-tokenized0 likes12k downloads11mo agoHugging Face15SakethVemula /fixed-tokenizer-segments0 likes10k downloads5mo agoHugging Face16AILab-CVC /obelics_seed2_tokensPart of the OBELISC data set, including 32 Million samples, please refer to dataset.py to use this data text10M<n<100M1 likes9.4k downloads3y agoHugging Face17tokyotech-llm /swallow-code-v2 SwallowCode-v2 Resources 📑 arXiv: Read our paper for detailed methodology and results at arXiv:2505.02881. 🤗 Sister Dataset: Discover SwallowMath-v2, our companion dataset for mathematical reasoning. 💻 What is it? SwallowCode-v1 was a high-quality Python code dataset generated through an LLM-based rewriting pipeline. However, it had two significant limitations: (1) it was distributed under the Llama 3.3 Community License, and (2) its size was limited to… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/swallow-code-v2.tabulartext-generation100M<n<1B48 likes8.6k downloads11mo agoHugging Face18chanind /c4-10k-mini-tokenized-16-ctx-gelu-1l-tests1K<n<10K0 likes7.5k downloads2y agoHugging Face19syafie-nzm /tokenized_datasettextn<1K0 likes7k downloads3y agoHugging Face20Ahmed-Nasri /llava-video-178k-siglip-tokens-ftov-new LLaVA-Video-178K SigLIP Token Cache (LLaVA-OV fine-tuned vision tower) Derived data (vision-encoder features of video frames), not a redistribution of the source videos. Source: lmms-lab/LLaVA-Video-178K -- its card restricts use to academic research and education, and its annotations come from GPT-4-class models (see the OpenAI usage policy). Complete: 85000 clips. Subset Folders: 0_30_s_academic_v0_1, 0_30_s_youtube_v0_1, 30_60_s_academic_v0_1… See the full description on the dataset page: https://huggingface.co/datasets/Ahmed-Nasri/llava-video-178k-siglip-tokens-ftov-new.video5 likes7k downloads5d agoHugging Face21regent-research /regent-subset-of-jat-dataset-tokenizedThis is the dataset for REGENT: A Retrieval-Augmented Generalist Agent That Can Act In-Context In New Environments. The REGENT dataset includes a subset of the JAT (tokenized) dataset (from https://huggingface.co/datasets/jat-project) for the REGENT training environments. It has around a 100k transitions from each of the 145 training environments (45 metaworld, 52 atari, 9 mujoco, 39 babyai). Please find this in the *_subset folders. It also has distance values for input sequences used in… See the full description on the dataset page: https://huggingface.co/datasets/regent-research/regent-subset-of-jat-dataset-tokenized.timeseries10M<n<100M0 likes6.7k downloads2y agoHugging Face22QuangDuy /FineWeb2-mds-tokenized-40960 likes6.6k downloads11mo agoHugging Face23TokyoTechMagicYang /RAM-W600gated Dataset Card for RAM-W600 Benchmark code is available in https://github.com/YSongxiao/RAM-W600. Download Please run the following command to download RAM-W600: git clone https://huggingface.co/datasets/TokyoTechMagicYang/RAM-W600 BoneSegmentation Mask Channel Mapping The BoneSegmentation masks are stored as 14-channel .npy arrays with shape: (14, H, W) Each channel is a binary mask for one anatomical structure. The official channel order is:… See the full description on the dataset page: https://huggingface.co/datasets/TokyoTechMagicYang/RAM-W600.image1K<n<10K3 likes6.4k downloads3mo agoHugging Face24QuangDuy /FineWeb2-mds-tokenized-10240 likes5.9k downloads11mo agoHugging Face25NeelNanda /c4-tokenized-2b Dataset Card for "c4-tokenized-2b" More Information needed 1M<n<10M0 likes5.6k downloads4y agoHugging Face26hungnm /vietnamese-tokenized0 likes5.6k downloads1y agoHugging Face27TokenBender /code_instructions_122k_alpaca_styletext100K<n<1M80 likes5.6k downloads3y agoHugging Face28karpathy /fineweb-edu-100B-gpt2-token-shardsFineWeb Edu 100B dataset tokenized with GPT-2 tokenizer using the code in llm.c repo. 8 likes5.2k downloads2y agoHugging Face29code-philia /mtpnet_tokens 模型训练过程汇总(持续更新中) 对于已收集的每一个模型,code 目录为模型定义、训练和测试的代码和脚本文件,model 目录为已收集的 epoch 模型文件,dataset.zip 为模型数据集。 下表汇总了所有收集的模型训练过程信息: 模型名称 模型简介 模型类型 Epoch数量 数据集信息 Clone-detection-BigCloneBench 基于大规模代码克隆基准数据集的代码克隆检测模型,任务是进行二元分类(0/1),其中1代表语义等价,0代表其他情况。 代码克隆检测 2个epoch BigCloneBench数据集 Clone-detection-POJ-104 基于POJ-104数据集的代码克隆检测模型,任务是识别不同编程题目中相似的代码实现,给定一段代码和一组候选代码,任务是返回具有相同语义的Top K个代码 代码克隆检测 2个epoch (0-1) POJ-104编程题目数据集… See the full description on the dataset page: https://huggingface.co/datasets/code-philia/mtpnet_tokens.2 likes4.7k downloads1y agoHugging Face30QuangDuy /FineWeb2-mds-tokenized-v2-10240 likes4.6k downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.