CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01NCSpeech /YO-CPT-ru YO-CPT-ru YouTube-Oriented dataset for Continual Pre-Training (Russian). A large, heavily quality-filtered corpus of Russian speech mined from YouTube (via YODAS2) and processed into clean, single-speaker, TTS-grade utterances. Every utterance ships with an ensemble-verified transcription, a punctuated/denormalized and stress-marked text variant, word-level forced alignment, within- and cross-video speaker identities, an audio-quality (MOS) score, and a speaker persona built… See the full description on the dataset page: https://huggingface.co/datasets/NCSpeech/YO-CPT-ru.audiotext-to-speech1M<n<10M16 likes10k downloads2mo agoHugging Face02SWE-bench /SWE-smith-cpptext1K<n<10K0 likes4.3k downloads7mo agoHugging Face03rishitdagli /cppe-5 Dataset Card for CPPE - 5 Dataset Summary CPPE - 5 (Medical Personal Protective Equipment) is a new challenging dataset with the goal to allow the study of subordinate categorization of medical personal protective equipments, which is not possible with other popular data sets that focus on broad level categories. Some features of this dataset are: high quality images and annotations (~4.6 bounding boxes per image) real-life images unlike any current such dataset majority… See the full description on the dataset page: https://huggingface.co/datasets/rishitdagli/cppe-5.imageobject-detection1K<n<10K23 likes3.8k downloads3y agoHugging Face04AlgorithmicResearchGroup /arxiv_cplusplus_research_code Dataset card for ArtifactAI/arxiv_cplusplus_research_code Dataset Description https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_cplusplus_research_code Dataset Summary ArtifactAI/arxiv_python_research_code contains over 10.6GB of source code files referenced strictly in ArXiv papers. The dataset serves as a curated dataset for Code LLMs. How to use it from datasets import load_dataset # full dataset (10.6GB of data) ds =… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_cplusplus_research_code.tabulartext-generation1M<n<10M9 likes2.8k downloads2y agoHugging Face05Romoamigo /SWE-Bench-MultilingualC_CPPFileteredtextn<1K0 likes1.8k downloads1y agoHugging Face06Aratako /LiquidAI-Hackathon-Tokyo-CPT-Data LiquidAI-Hackathon-Tokyo-CPT-Data Liquid AI Hackathon Tokyoで作成したモデルのCPTに利用したデータセットです。 automatic-speech-recognition1M<n<10M6 likes1.6k downloads1y agoHugging Face07Romoamigo /SWE-Bench-MultilingualC_CPPFiletered_newtextn<1K0 likes1.6k downloads1y agoHugging Face08mhurhangee /cpc-classificationstext100K<n<1M0 likes1.5k downloads1y agoHugging Face09robert-co /bls_cpi Changelog 2025-01-18 I decided that I'll name the column survey, instead of consumer. I'll also set the value to the description, instead of the code. I didn't realize that the pandas version of "Use this dataset" includes the filename. I'll remove the date from the filename. So that people do not have to change their code. I have been updating this using the UI, but I will create a script in Spaces to update this from the BLS site. 2025-01-12 While using… See the full description on the dataset page: https://huggingface.co/datasets/robert-co/bls_cpi.tabulartable-question-answering1M<n<10M0 likes1.2k downloads8mo agoHugging Face10NCSpeech /YO-CPT-kk YO-CPT-kk YouTube-Oriented dataset for Continual Pre-Training (Kazakh). A heavily quality-filtered corpus of Kazakh speech mined from YouTube and processed into clean, single-speaker, TTS-grade utterances. Every utterance ships with an ensemble-verified transcription, a punctuated/denormalized and stress-marked text variant, word-level forced alignment, within- and cross-video speaker identities, an audio-quality (MOS) score, and a speaker persona built from the voice and, where… See the full description on the dataset page: https://huggingface.co/datasets/NCSpeech/YO-CPT-kk.audiotext-to-speech100K<n<1M10 likes1.2k downloads2mo agoHugging Face11coldchair16 /CPRet-data CPRet-data This repository hosts the datasets for CPRet: A Dataset, Benchmark, and Model for Retrieval in Competitive Programming. Visit https://cpret.online/ to try out CPRet in action for competitive programming problem retrieval. 💡 CPRet Benchmark Tasks The CPRet dataset supports four retrieval tasks relevant to competitive programming: Text-to-Code Retrieval Retrieve relevant code snippets based on a natural language problem description. Code-to-Code Retrieval… See the full description on the dataset page: https://huggingface.co/datasets/coldchair16/CPRet-data.text100K<n<1M3 likes1.1k downloads1y agoHugging Face12Reset23 /the-stack-v2-new-cpptabular1M<n<10M1 likes1.1k downloads1y agoHugging Face13akzsh /indic-oss-mixture-cpt-10btext1M<n<10M0 likes822 downloads4mo agoHugging Face14Reset23 /the-stack-v2-cpptabular1M<n<10M1 likes808 downloads2y agoHugging Face15AINovice2005 /carbon-cpu-enriched-sequences carbon-cpu-enriched-sequences A CPU-enriched subset of the carbon pretraining corpus (eukaryote_generator), combining original source fields with normalized sequences and row-level features for quality analysis, GPU enrichment and embedding generation. Information of Features Feature Type Description record_id string NCBI Identifier linking the row back to the source genomic record. It provides the primary record-level identity. begin_of_sequence… See the full description on the dataset page: https://huggingface.co/datasets/AINovice2005/carbon-cpu-enriched-sequences.tabulartext-generation10M<n<100M0 likes694 downloads10d agoHugging Face16AINovice2005 /carbon-cpu-enriched-sequences-sampledtabular1M<n<10M0 likes659 downloads1mo agoHugging Face17eternalaudrey /jump-cp-0016-labelfree-Mitoimage10K<n<100K0 likes555 downloads1y agoHugging Face18cpratikaki /RSVQA-HR_qwen_finetuningimage100K<n<1M1 likes550 downloads2y agoHugging Face19Nottybro /csharp-dotnet-cpt-v0 csharp-dotnet-cpt-v0 A reproducible C# / .NET continued-pretraining (CPT) corpus for Qwen2.5-Coder-1.5B. HF: https://huggingface.co/datasets/Nottybro/csharp-dotnet-cpt-v0 (private) Total unique tokens: 695,288,168 (Qwen2.5-Coder-1.5B tokenizer) Files: 681,426 | Repositories: 68,869 Format: Zstandard-compressed Parquet, schema below. See DATASET_CARD.md for sources, license policy, filtering, limitations. See reports/summary.md for full statistics. tabular10K<n<100K0 likes528 downloads2mo agoHugging Face20ningani /stack-v2-cpp-2019tabular10M<n<100M0 likes523 downloads2y agoHugging Face21Podtech /Swallow-Nemotron-Post-Training-Dataset-v1-ja-cpt Swallow-Nemotron-Post-Training-Dataset-v1-ja-cpt Dataset Overview This dataset is a reformatted subset of the tokyotech-llm/Swallow-Nemotron-Post-Training-Dataset-v1 dataset, specifically derived from the v1-Ja-202601 subset. It was created to facilitate Continuous Pre-Training (CPT) by extracting only the text_gpt_oss field from the original data. Dataset Statistics & Token Counts The token counts for each category were calculated using the… See the full description on the dataset page: https://huggingface.co/datasets/Podtech/Swallow-Nemotron-Post-Training-Dataset-v1-ja-cpt.texttext-generation1M<n<10M0 likes519 downloads1mo agoHugging Face22konwoo /dclm-164k-300m-raw-cpr120-ml1024text10M<n<100M0 likes488 downloads10mo agoHugging Face23konwoo /dclm-164k-real-train-8b-instruct-hq-cpr32-ml1024text1M<n<10M0 likes452 downloads8mo agoHugging Face24konwoo /dclm-164k-instruct-hq-cpr128-ml512-92-127text1M<n<10M0 likes446 downloads10mo agoHugging Face25DopeorNope /CPT_textbookstext1M<n<10M0 likes436 downloads10mo agoHugging Face26ThomasTheMaker /arc-stack-cpptabular1M<n<10M0 likes435 downloads11mo agoHugging Face27rayrren /CPIBench-0 CPIBench-0 CPIBench-0 is our first benchmark for evaluating how frontier models and agent harnesses perform in operations that rely heavily on cyber-physical systems. Built from real projects across mechatronics, manufacturing, materials, and energy, CPIBench-0 tests models and agent harnesses in multimodal workflows where cyber-physical signals, tools, and real-world consequences intersect. CPIBench-0 measures performance across connected cyber-physical workflows, from making… See the full description on the dataset page: https://huggingface.co/datasets/rayrren/CPIBench-0.tabularn<1K1 likes434 downloads2mo agoHugging Face28RLAIF /optim_policy_pretrain-pythia-160m_lr0.0001_bs24_wp1_wd0.01_ep0_cp35k-mergedtabular100K<n<1M0 likes416 downloads2y agoHugging Face29konwoo /dclm-164k-300m-cpr200-ml1024text10M<n<100M0 likes412 downloads10mo agoHugging Face30cpnlab /LSM-Tokenized-Full 📡 Large Spectrum Models (LSM) – Tokenized RF Dataset Overview This repository provides the tokenized RF spectrum dataset introduced in the paper: “Large Spectrum Models (LSMs): Decoder-Only Transformer-Powered Spectrum Activity Forecasting via Tokenized RF Data” The dataset is designed to bridge wireless signal processing and large language models (LLMs) by converting raw spectrum measurements into discrete token sequences, enabling direct use with transformer-based… See the full description on the dataset page: https://huggingface.co/datasets/cpnlab/LSM-Tokenized-Full.10M<n<100M0 likes405 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.