CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01NCSpeech /YO-CPT-ru YO-CPT-ru YouTube-Oriented dataset for Continual Pre-Training (Russian). A large, heavily quality-filtered corpus of Russian speech mined from YouTube (via YODAS2) and processed into clean, single-speaker, TTS-grade utterances. Every utterance ships with an ensemble-verified transcription, a punctuated/denormalized and stress-marked text variant, word-level forced alignment, within- and cross-video speaker identities, an audio-quality (MOS) score, and a speaker persona built… See the full description on the dataset page: https://huggingface.co/datasets/NCSpeech/YO-CPT-ru.audiotext-to-speech1M<n<10M15 likes13k downloads2mo agoHugging Face02optimum-benchmark /cpu0 likes11k downloads5d agoHugging Face03zhihuanglab /Paladin_TCGA_CPTAC_omicsPaladin TCGA & CPTAC Spatial Omics Maps Ready-to-use patch-level and slide-level spatial omics maps inferred by Paladin from TCGA and CPTAC whole-slide images. The released .Paladin.h5 files can be analyzed directly without rerunning WSI inference. The collection is populated in stages. Check Files and versions for the cohorts currently available. Spatial multi-omics example The panels show the H&E WSI, a reference tumor mask, CNV burden, TP53 CNV, DNA-methylation… See the full description on the dataset page: https://huggingface.co/datasets/zhihuanglab/Paladin_TCGA_CPTAC_omics.0 likes9.1k downloads1mo agoHugging Face04ZhejiangLab /CPT_Data_Pool CPT Data Pool This repository hosts a large-scale CPT corpus designed for continual pre-training of domain-specific large language models. It serves as a data component of the 👉F2D-LLM framework, an end-to-end pipeline for domain-specific LLM training. For detailed data processing, scoring methods, and sampling strategies, please refer to the official 👉GitHub repo. Overall, the dataset contains approximately 271B tokens and is stored as jsonl files with a unified schema… See the full description on the dataset page: https://huggingface.co/datasets/ZhejiangLab/CPT_Data_Pool.text100M<n<1B1 likes4.7k downloads3mo agoHugging Face05SWE-bench /SWE-smith-cpptext1K<n<10K0 likes4.3k downloads7mo agoHugging Face06Dearcat /CPathPatchFeature CPathPatchFeature: Pre-extracted WSI Features for Computational Pathology Paper: Revisiting End-to-End Learning with Slide-level Supervision in Computational Pathology Code: https://github.com/DearCaat/E2E-WSI-ABMILX Dataset Summary This dataset provides a comprehensive collection of pre-extracted features from Whole Slide Images (WSIs) for various cancer types, designed to facilitate research in computational pathology. The features are extracted using multiple… See the full description on the dataset page: https://huggingface.co/datasets/Dearcat/CPathPatchFeature.image-feature-extraction100B<n<1T7 likes4.2k downloads11mo agoHugging Face07AIencoder /llama-cpp-wheelsIf you like this please consider liking and donating (https://buymeacoffee.com/aiencoder) 🏭 llama-cpp-python Mega-Factory Wheels "Stop waiting for pip to compile. Just install and run." The most complete collection of pre-built llama-cpp-python wheels in existence — 8,333 wheels across every platform, Python version, backend, and CPU optimization level. No more cmake, gcc, or compilation hell. No more waiting 10 minutes for a build that might fail. Just find your wheel and… See the full description on the dataset page: https://huggingface.co/datasets/AIencoder/llama-cpp-wheels.text-generation1K<n<10K4 likes3.9k downloads2h agoHugging Face08rishitdagli /cppe-5 Dataset Card for CPPE - 5 Dataset Summary CPPE - 5 (Medical Personal Protective Equipment) is a new challenging dataset with the goal to allow the study of subordinate categorization of medical personal protective equipments, which is not possible with other popular data sets that focus on broad level categories. Some features of this dataset are: high quality images and annotations (~4.6 bounding boxes per image) real-life images unlike any current such dataset majority… See the full description on the dataset page: https://huggingface.co/datasets/rishitdagli/cppe-5.imageobject-detection1K<n<10K23 likes3.5k downloads3y agoHugging Face09AlgorithmicResearchGroup /arxiv_cplusplus_research_code Dataset card for ArtifactAI/arxiv_cplusplus_research_code Dataset Description https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_cplusplus_research_code Dataset Summary ArtifactAI/arxiv_python_research_code contains over 10.6GB of source code files referenced strictly in ArXiv papers. The dataset serves as a curated dataset for Code LLMs. How to use it from datasets import load_dataset # full dataset (10.6GB of data) ds =… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_cplusplus_research_code.tabulartext-generation1M<n<10M9 likes2.8k downloads2y agoHugging Face10Brunobkr /llama.cpp_AlgMor24_github ΩFFFΣLLIa • llama.cpp • AlgMor24 ██████╗ ███████╗███████╗███████╗██╗ ██╗ ██╗ █████╗ ██╔═══██╗██╔════╝██╔════╝██╔════╝██║ ██║ ██║██╔══██╗ ██║ ██║█████╗ █████╗ █████╗ ██║ ██║ ██║███████║ ██║ ██║██╔══╝ ██╔══╝ ██╔══╝ ██║ ██║ ██║██╔══██║ ╚██████╔╝██║ ██║ ███████╗███████╗███████╗██║██║ ██║ ╚═════╝ ╚═╝ ╚═╝ ╚══════╝╚══════╝╚══════╝╚═╝╚═╝ ╚═╝ High-Performance LLM / VLM Inference & Autonomous Agentic Ecosystem… See the full description on the dataset page: https://huggingface.co/datasets/Brunobkr/llama.cpp_AlgMor24_github.0 likes2.8k downloads1mo agoHugging Face11skeole /qwen-cpp-agent-0-protocolExperiment in agentic autonomy protocols. ~ everything in this repo was created by Qwen 3.8 27B (Q4) running autonomously inside Deepseek Harness, on a single RTX 3090 GPU, for 3 weeks. The only human artifacts are: agents/* human/* AGENTS.md texttext-generation1K<n<10K2 likes2.8k downloads1d agoHugging Face12dslighfdsl /human_eval_cpptext10K<n<100K1 likes2.3k downloads2y agoHugging Face13MeghanaKap /miomio_cp1_cache0 likes2.3k downloads2mo agoHugging Face14Retrobear /demucs.cppThis repo stores weights in ggml format that are used to perform music separation. These are intended to be used with demucs.cpp, https://github.com/sevagh/demucs.cpp Weights origin: https://dl.fbaipublicfiles.com/demucs/hybrid_transformer/955717e8-8726e21a.th https://dl.fbaipublicfiles.com/demucs/hybrid_transformer/5c90dfd2-34c22ccb.th https://dl.fbaipublicfiles.com/demucs/hybrid_transformer/f7e0c4bc-ba3fe64a.th https://dl.fbaipublicfiles.com/demucs/hybrid_transformer/d12395a8-e57c48e6.th… See the full description on the dataset page: https://huggingface.co/datasets/Retrobear/demucs.cpp.4 likes2.2k downloads2y agoHugging Face15andito /qwentts-cpp-python-wheels qwentts-cpp-python wheels Optional backend-specific wheel variants for qwentts-cpp-python. The default PyPI package is CUDA 12.8: pip install qwentts-cpp-python Install a backend-specific wheel from this repository with --find-links: pip install "qwentts-cpp-python==0.3.1+cpu" -f https://huggingface.co/datasets/andito/qwentts-cpp-python-wheels/tree/main/whl/cpu pip install "qwentts-cpp-python==0.3.1+cu124" -f… See the full description on the dataset page: https://huggingface.co/datasets/andito/qwentts-cpp-python-wheels.0 likes2.2k downloads2mo agoHugging Face16mhurhangee /cpc-classificationstext100K<n<1M0 likes2.1k downloads1y agoHugging Face17echodict /whisper.cpp whisper.cpp Stable: v1.8.1 / Roadmap High-performance inference of OpenAI's Whisper automatic speech recognition (ASR) model: Plain C/C++ implementation without dependencies Apple Silicon first-class citizen - optimized via ARM NEON, Accelerate framework, Metal and Core ML AVX intrinsics support for x86 architectures VSX intrinsics support for POWER architectures Mixed F16 / F32 precision Integer quantization support Zero memory allocations at runtime Vulkan support Support… See the full description on the dataset page: https://huggingface.co/datasets/echodict/whisper.cpp.0 likes2.1k downloads6mo agoHugging Face18jiviteshjn /fineweb-edu-zh-chengyu-cpt Fineweb-Edu Chinese — Chengyu-Tagged Continued-Pretraining Corpus A 3.74M-document Chinese corpus (~7.8B tokens) for continued pretraining on cultural knowledge in figurative language, built from the highest-quality tier of opencsg/Fineweb-Edu-Chinese-V2.1. Each document is educational Chinese text containing at least one culturally vetted chengyu, with an appended 【成语注释】 knowledge block listing every matched idiom's figurative meaning(s) and classical source citation. This is a… See the full description on the dataset page: https://huggingface.co/datasets/jiviteshjn/fineweb-edu-zh-chengyu-cpt.tabulartext-generation1M<n<10M1 likes2k downloads2mo agoHugging Face19Romoamigo /SWE-Bench-MultilingualC_CPPFileteredtextn<1K0 likes1.8k downloads1y agoHugging Face20Romoamigo /SWE-Bench-MultilingualC_CPPFiletered_newtextn<1K0 likes1.6k downloads1y agoHugging Face21Aratako /LiquidAI-Hackathon-Tokyo-CPT-Data LiquidAI-Hackathon-Tokyo-CPT-Data Liquid AI Hackathon Tokyoで作成したモデルのCPTに利用したデータセットです。 automatic-speech-recognition1M<n<10M6 likes1.5k downloads1y agoHugging Face22cpral /forums_pol_json_zst0 likes1.4k downloads6mo agoHugging Face23sarthak20024 /sih26099-cpse-material-codes SIH 26099 — Collected Dataset AI-Driven Standardization & Harmonization of Material Codes Across CPSEs This workspace holds the data-collection stage only — no model, no training, no feature engineering. Just raw public sources, their extracted structured form, and the reference taxonomies/vocabularies the harmonisation step needs. Collected live on 2026-09-08. All row counts below were verified by reading the files back with pandas. 1. Headline numbers… See the full description on the dataset page: https://huggingface.co/datasets/sarthak20024/sih26099-cpse-material-codes.texttoken-classification10K<n<100K1 likes1.4k downloads10d agoHugging Face24cphan11 /anionevideo1K<n<10K3 likes1.4k downloads7h agoHugging Face25NCSpeech /YO-CPT-kk YO-CPT-kk YouTube-Oriented dataset for Continual Pre-Training (Kazakh). A heavily quality-filtered corpus of Kazakh speech mined from YouTube and processed into clean, single-speaker, TTS-grade utterances. Every utterance ships with an ensemble-verified transcription, a punctuated/denormalized and stress-marked text variant, word-level forced alignment, within- and cross-video speaker identities, an audio-quality (MOS) score, and a speaker persona built from the voice and, where… See the full description on the dataset page: https://huggingface.co/datasets/NCSpeech/YO-CPT-kk.audiotext-to-speech100K<n<1M9 likes1.3k downloads2mo agoHugging Face26coldchair16 /CPRet-data CPRet-data This repository hosts the datasets for CPRet: A Dataset, Benchmark, and Model for Retrieval in Competitive Programming. Visit https://cpret.online/ to try out CPRet in action for competitive programming problem retrieval. 💡 CPRet Benchmark Tasks The CPRet dataset supports four retrieval tasks relevant to competitive programming: Text-to-Code Retrieval Retrieve relevant code snippets based on a natural language problem description. Code-to-Code Retrieval… See the full description on the dataset page: https://huggingface.co/datasets/coldchair16/CPRet-data.text100K<n<1M3 likes1.3k downloads11mo agoHugging Face27alrope /cpds_embeddings1 likes1.2k downloads1y agoHugging Face28Reset23 /the-stack-v2-new-cpptabular1M<n<10M1 likes1.1k downloads1y agoHugging Face29rohanphanse /cpsc5800-hand-detection Training Datasets and Model Weights for CPSC 5800 Final Project Project repository: https://github.com/rohanphanse/CPSC5800-Final We provide all training datasets created in Step 1 and weights for the YOLO and ResNet models trained during Steps 2-4 in our Hugging Face repository: https://huggingface.co/datasets/rohanphanse/cpsc5800-hand-detection # Recommended: download dataset using git-xet (https://hf.co/docs/hub/git-xet) brew install git-xet git xet install # Download datasets… See the full description on the dataset page: https://huggingface.co/datasets/rohanphanse/cpsc5800-hand-detection.image-classification10K<n<100K0 likes1k downloads9mo agoHugging Face30intelli-zen /cppe-5CPPE - 5 (Medical Personal Protective Equipment) is a new challenging dataset with the goal to allow the study of subordinate categorization of medical personal protective equipments, which is not possible with other popular data sets that focus on broad level categories.object-detection100M<n<1B0 likes1k downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.