CP
Models
All models matching “CP”Datasets
All datasets matching “CP”YO-CPT-ru
YO-CPT-ru
YouTube-Oriented dataset for Continual Pre-Training (Russian). A large, heavily
quality-filtered corpus of Russian speech mined from YouTube (via YODAS2)
and processed into clean, single-speaker, TTS-grade utterances. Every utterance ships with an
ensemble-verified transcription, a punctuated/denormalized and stress-marked text variant, word-level
forced alignment, within- and cross-video speaker identities, an audio-quality (MOS) score, and a
speaker persona built… See the full description on the dataset page: https://huggingface.co/datasets/NCSpeech/YO-CPT-ru.cpuPaladin_TCGA_CPTAC_omicsPaladin TCGA & CPTAC Spatial Omics Maps
Ready-to-use patch-level and slide-level spatial omics maps inferred by
Paladin from TCGA and CPTAC
whole-slide images. The released .Paladin.h5 files can be analyzed directly
without rerunning WSI inference.
The collection is populated in stages. Check Files and versions for the
cohorts currently available.
Spatial multi-omics example
The panels show the H&E WSI, a reference tumor mask, CNV burden, TP53 CNV,
DNA-methylation… See the full description on the dataset page: https://huggingface.co/datasets/zhihuanglab/Paladin_TCGA_CPTAC_omics.CPT_Data_Pool
CPT Data Pool
This repository hosts a large-scale CPT corpus designed for continual pre-training of domain-specific large language models. It serves as a data component of the 👉F2D-LLM framework, an end-to-end pipeline for domain-specific LLM training.
For detailed data processing, scoring methods, and sampling strategies, please refer to the official 👉GitHub repo.
Overall, the dataset contains approximately 271B tokens and is stored as jsonl files with a unified schema… See the full description on the dataset page: https://huggingface.co/datasets/ZhejiangLab/CPT_Data_Pool.SWE-smith-cppCPathPatchFeature
CPathPatchFeature: Pre-extracted WSI Features for Computational Pathology
Paper: Revisiting End-to-End Learning with Slide-level Supervision in Computational Pathology
Code: https://github.com/DearCaat/E2E-WSI-ABMILX
Dataset Summary
This dataset provides a comprehensive collection of pre-extracted features from Whole Slide Images (WSIs) for various cancer types, designed to facilitate research in computational pathology. The features are extracted using multiple… See the full description on the dataset page: https://huggingface.co/datasets/Dearcat/CPathPatchFeature.
