CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01NCSpeech /YO-CPT-ru YO-CPT-ru YouTube-Oriented dataset for Continual Pre-Training (Russian). A large, heavily quality-filtered corpus of Russian speech mined from YouTube (via YODAS2) and processed into clean, single-speaker, TTS-grade utterances. Every utterance ships with an ensemble-verified transcription, a punctuated/denormalized and stress-marked text variant, word-level forced alignment, within- and cross-video speaker identities, an audio-quality (MOS) score, and a speaker persona built… See the full description on the dataset page: https://huggingface.co/datasets/NCSpeech/YO-CPT-ru.audiotext-to-speech1M<n<10M16 likes12k downloads2mo agoHugging Face02ZhejiangLab /CPT_Data_Pool CPT Data Pool This repository hosts a large-scale CPT corpus designed for continual pre-training of domain-specific large language models. It serves as a data component of the 👉F2D-LLM framework, an end-to-end pipeline for domain-specific LLM training. For detailed data processing, scoring methods, and sampling strategies, please refer to the official 👉GitHub repo. Overall, the dataset contains approximately 271B tokens and is stored as jsonl files with a unified schema… See the full description on the dataset page: https://huggingface.co/datasets/ZhejiangLab/CPT_Data_Pool.text100M<n<1B1 likes4.8k downloads3mo agoHugging Face03SWE-bench /SWE-smith-cpptext1K<n<10K0 likes4.3k downloads7mo agoHugging Face04AlgorithmicResearchGroup /arxiv_cplusplus_research_code Dataset card for ArtifactAI/arxiv_cplusplus_research_code Dataset Description https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_cplusplus_research_code Dataset Summary ArtifactAI/arxiv_python_research_code contains over 10.6GB of source code files referenced strictly in ArXiv papers. The dataset serves as a curated dataset for Code LLMs. How to use it from datasets import load_dataset # full dataset (10.6GB of data) ds =… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_cplusplus_research_code.tabulartext-generation1M<n<10M9 likes2.9k downloads2y agoHugging Face05skeole /qwen-cpp-agent-0-protocolExperiment in agentic autonomy protocols. ~ everything in this repo was created by Qwen 3.8 27B (Q4) running autonomously inside Deepseek Harness, on a single RTX 3090 GPU, for 3 weeks. The only human artifacts are: agents/* human/* AGENTS.md texttext-generation1K<n<10K2 likes2.8k downloads2d agoHugging Face06dslighfdsl /human_eval_cpptext10K<n<100K1 likes2.3k downloads2y agoHugging Face07mhurhangee /cpc-classificationstext100K<n<1M0 likes2.1k downloads1y agoHugging Face08jiviteshjn /fineweb-edu-zh-chengyu-cpt Fineweb-Edu Chinese — Chengyu-Tagged Continued-Pretraining Corpus A 3.74M-document Chinese corpus (~7.8B tokens) for continued pretraining on cultural knowledge in figurative language, built from the highest-quality tier of opencsg/Fineweb-Edu-Chinese-V2.1. Each document is educational Chinese text containing at least one culturally vetted chengyu, with an appended 【成语注释】 knowledge block listing every matched idiom's figurative meaning(s) and classical source citation. This is a… See the full description on the dataset page: https://huggingface.co/datasets/jiviteshjn/fineweb-edu-zh-chengyu-cpt.tabulartext-generation1M<n<10M1 likes1.9k downloads2mo agoHugging Face09Romoamigo /SWE-Bench-MultilingualC_CPPFileteredtextn<1K0 likes1.8k downloads1y agoHugging Face10Romoamigo /SWE-Bench-MultilingualC_CPPFiletered_newtextn<1K0 likes1.6k downloads1y agoHugging Face11sarthak20024 /sih26099-cpse-material-codes SIH 26099 — Collected Dataset AI-Driven Standardization & Harmonization of Material Codes Across CPSEs This workspace holds the data-collection stage only — no model, no training, no feature engineering. Just raw public sources, their extracted structured form, and the reference taxonomies/vocabularies the harmonisation step needs. Collected live on 2026-09-08. All row counts below were verified by reading the files back with pandas. 1. Headline numbers… See the full description on the dataset page: https://huggingface.co/datasets/sarthak20024/sih26099-cpse-material-codes.texttoken-classification10K<n<100K1 likes1.4k downloads11d agoHugging Face12NCSpeech /YO-CPT-kk YO-CPT-kk YouTube-Oriented dataset for Continual Pre-Training (Kazakh). A heavily quality-filtered corpus of Kazakh speech mined from YouTube and processed into clean, single-speaker, TTS-grade utterances. Every utterance ships with an ensemble-verified transcription, a punctuated/denormalized and stress-marked text variant, word-level forced alignment, within- and cross-video speaker identities, an audio-quality (MOS) score, and a speaker persona built from the voice and, where… See the full description on the dataset page: https://huggingface.co/datasets/NCSpeech/YO-CPT-kk.audiotext-to-speech100K<n<1M10 likes1.3k downloads2mo agoHugging Face13coldchair16 /CPRet-data CPRet-data This repository hosts the datasets for CPRet: A Dataset, Benchmark, and Model for Retrieval in Competitive Programming. Visit https://cpret.online/ to try out CPRet in action for competitive programming problem retrieval. 💡 CPRet Benchmark Tasks The CPRet dataset supports four retrieval tasks relevant to competitive programming: Text-to-Code Retrieval Retrieve relevant code snippets based on a natural language problem description. Code-to-Code Retrieval… See the full description on the dataset page: https://huggingface.co/datasets/coldchair16/CPRet-data.text100K<n<1M3 likes1.3k downloads11mo agoHugging Face14Reset23 /the-stack-v2-new-cpptabular1M<n<10M1 likes1.1k downloads1y agoHugging Face15GesturingMan /CPC_Text_Roughtext1K<n<10K0 likes937 downloads3y agoHugging Face16xlu11 /CP2077gated Cyberpunk 2077 controllable RGB-D capture Each row in metadata.jsonl links a 16 FPS H.264 RGB video with bit-exact uint16 logarithmic depth PNGs, frame-aligned action/player/camera records, and 60 Hz controller records. See each clip manifest for depth decoding parameters and checksums. image1K<n<10K0 likes914 downloads3d agoHugging Face17hosseinbv /dim58-cpuData-31cases Dim58 CPU Data — 31 Cases Dataset uploaded from: /mnt/data/ubuntu/research/outputs/data_cpu_geodesic58 Dataset summary Property Value Repository hosseinbv/dim58-cpuData-31cases Number of files 64 Total size 17.87 GB Source folder data_cpu_geodesic58 File types Extension File count .npz 62 .json 1 .csv 1 Top-level contents 0000_internal_case1_data.npz 0001_internal_B_10.npz… See the full description on the dataset page: https://huggingface.co/datasets/hosseinbv/dim58-cpuData-31cases.tabularn<1K0 likes894 downloads2mo agoHugging Face18dchasap /spec_cpu_branch_tracestext10B<n<100B0 likes892 downloads2y agoHugging Face19ajibawa-2023 /Cpp-Code-LargeCpp-Code-Large Cpp-Code-Large is a large-scale corpus of C++ source code comprising more than 5 million lines of C++ code. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and static program analysis for the C++ ecosystem. By providing a high-volume, language-specific corpus, Cpp-Code-Large enables systematic experimentation in C++-focused model training, domain adaptation, and downstream code… See the full description on the dataset page: https://huggingface.co/datasets/ajibawa-2023/Cpp-Code-Large.texttext-generation1M<n<10M16 likes848 downloads7mo agoHugging Face20Reset23 /the-stack-v2-cpptabular1M<n<10M1 likes806 downloads2y agoHugging Face21echodict /llama.cppversion https://git-lfs.github.com/spec/v1 oid sha256:cfc44b7ba25614df70e6b65e3341cae0310163bd32fd31a6b928a542df433faf size 30786 textn<1K0 likes773 downloads5mo agoHugging Face22AINovice2005 /carbon-cpu-enriched-sequences carbon-cpu-enriched-sequences A CPU-enriched subset of the carbon pretraining corpus (eukaryote_generator), combining original source fields with normalized sequences and row-level features for quality analysis, GPU enrichment and embedding generation. Information of Features Feature Type Description record_id string NCBI Identifier linking the row back to the source genomic record. It provides the primary record-level identity. begin_of_sequence… See the full description on the dataset page: https://huggingface.co/datasets/AINovice2005/carbon-cpu-enriched-sequences.tabulartext-generation10M<n<100M0 likes721 downloads7d agoHugging Face23akzsh /indic-oss-mixture-cpt-10btext1M<n<10M0 likes661 downloads4mo agoHugging Face24AINovice2005 /carbon-cpu-enriched-sequences-sampledtabular1M<n<10M0 likes649 downloads28d agoHugging Face25coldchair16 /CPRet-Embeddings CPRet-Embeddings This repository provides the problem descriptions and their corresponding precomputed embeddings used by the CPRet retrieval server. You can explore the retrieval server via the online demo at https://cpret.online/. 📦 Files probs_2609.jsonlA JSONL file containing natural language descriptions of competitive programming problems.Each line is a JSON object with metadata such as problem title, platform/source OJ, URL, and full description.… See the full description on the dataset page: https://huggingface.co/datasets/coldchair16/CPRet-Embeddings.text100K<n<1M1 likes546 downloads17d agoHugging Face26nvidia /LiveCodeBench-CPP LiveCodeBench-CPP: An Extension of LiveCodeBench for Contamination Free Evaluation in C++ Overview LiveCodeBench-CPP includes 454 problems from the release_v6 of LiveCodeBench, covering the period from October 2024 to May 2025. These problems are sourced from AtCoder (287 problems) and LeetCode (167 problems). AtCoder Problems: These require generated solutions to read inputs from standard input (stdin) and write outputs to standard output (stdout). For unit testing, the… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/LiveCodeBench-CPP.textn<1K4 likes540 downloads1y agoHugging Face27UWA-CP /Underwater-Acoustic-Channel-Repository Underwater Acoustic Channel Repository This Hugging Face dataset is a structured, checksum-preserving mirror of version 1.0 of the Underwater Acoustic Channel Repository. The original dataset was published by Zhengnan Li, Mandar Chitre, Diego Cuji, James Preisig, Andrew Singer, Milica Stojanovic, and Paul van Walree. The collection contains measured underwater acoustic channel impulse responses (CIRs) from eight at-sea experimental groups. Channel and accompanying noise files… See the full description on the dataset page: https://huggingface.co/datasets/UWA-CP/Underwater-Acoustic-Channel-Repository.texttime-series-forecastingn<1K2 likes537 downloads18d agoHugging Face28cpratikaki /RSVQA-HR_qwen_finetuningimage100K<n<1M1 likes534 downloads2y agoHugging Face29cpystan /MSMU MSMU (Massive Spatial Measuring and Understanding Dataset for Spatial Intelligence) 🌐 Homepage | 🤗 Dataset | 📖 arXiv | GitHub Dataset Details Dataset Description We introduce MSMU and MSMU-Bench: a new benchmark designed to enhance and evaluate multimodal models on spatial measuring and understanding. MSMU is featured as metric-accurate spatial annotations which are sourced from high-precision 3D scenes. It contains , 25K images, 700K QA pairs… See the full description on the dataset page: https://huggingface.co/datasets/cpystan/MSMU.imagequestion-answering10K<n<100K2 likes508 downloads4mo agoHugging Face30ThomasTheMaker /arc-stack-cpptabular1M<n<10M0 likes502 downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.