datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
YO-CPT-ru
YO-CPT-ru
YouTube-Oriented dataset for Continual Pre-Training (Russian). A large, heavily
quality-filtered corpus of Russian speech mined from YouTube (via YODAS2)
and processed into clean, single-speaker, TTS-grade utterances. Every utterance ships with an
ensemble-verified transcription, a punctuated/denormalized and stress-marked text variant, word-level
forced alignment, within- and cross-video speaker identities, an audio-quality (MOS) score, and a
speaker persona built… See the full description on the dataset page: https://huggingface.co/datasets/NCSpeech/YO-CPT-ru.cpuPaladin_TCGA_CPTAC_omicsPaladin TCGA & CPTAC Spatial Omics Maps
Ready-to-use patch-level and slide-level spatial omics maps inferred by
Paladin from TCGA and CPTAC
whole-slide images. The released .Paladin.h5 files can be analyzed directly
without rerunning WSI inference.
The collection is populated in stages. Check Files and versions for the
cohorts currently available.
Spatial multi-omics example
The panels show the H&E WSI, a reference tumor mask, CNV burden, TP53 CNV,
DNA-methylation… See the full description on the dataset page: https://huggingface.co/datasets/zhihuanglab/Paladin_TCGA_CPTAC_omics.CPT_Data_Pool
CPT Data Pool
This repository hosts a large-scale CPT corpus designed for continual pre-training of domain-specific large language models. It serves as a data component of the 👉F2D-LLM framework, an end-to-end pipeline for domain-specific LLM training.
For detailed data processing, scoring methods, and sampling strategies, please refer to the official 👉GitHub repo.
Overall, the dataset contains approximately 271B tokens and is stored as jsonl files with a unified schema… See the full description on the dataset page: https://huggingface.co/datasets/ZhejiangLab/CPT_Data_Pool.SWE-smith-cppCPathPatchFeature
CPathPatchFeature: Pre-extracted WSI Features for Computational Pathology
Paper: Revisiting End-to-End Learning with Slide-level Supervision in Computational Pathology
Code: https://github.com/DearCaat/E2E-WSI-ABMILX
Dataset Summary
This dataset provides a comprehensive collection of pre-extracted features from Whole Slide Images (WSIs) for various cancer types, designed to facilitate research in computational pathology. The features are extracted using multiple… See the full description on the dataset page: https://huggingface.co/datasets/Dearcat/CPathPatchFeature.llama-cpp-wheelsIf you like this please consider liking and donating (https://buymeacoffee.com/aiencoder)
🏭 llama-cpp-python Mega-Factory Wheels
"Stop waiting for pip to compile. Just install and run."
The most complete collection of pre-built llama-cpp-python wheels in existence — 8,333 wheels across every platform, Python version, backend, and CPU optimization level.
No more cmake, gcc, or compilation hell. No more waiting 10 minutes for a build that might fail. Just find your wheel and… See the full description on the dataset page: https://huggingface.co/datasets/AIencoder/llama-cpp-wheels.cppe-5
Dataset Card for CPPE - 5
Dataset Summary
CPPE - 5 (Medical Personal Protective Equipment) is a new challenging dataset with the goal to allow the study of subordinate categorization of medical personal protective equipments, which is not possible with other popular data sets that focus on broad level categories.
Some features of this dataset are:
high quality images and annotations (~4.6 bounding boxes per image)
real-life images unlike any current such dataset
majority… See the full description on the dataset page: https://huggingface.co/datasets/rishitdagli/cppe-5.arxiv_cplusplus_research_code
Dataset card for ArtifactAI/arxiv_cplusplus_research_code
Dataset Description
https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_cplusplus_research_code
Dataset Summary
ArtifactAI/arxiv_python_research_code contains over 10.6GB of source code files referenced strictly in ArXiv papers. The dataset serves as a curated dataset for Code LLMs.
How to use it
from datasets import load_dataset
# full dataset (10.6GB of data)
ds =… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_cplusplus_research_code.llama.cpp_AlgMor24_github
ΩFFFΣLLIa • llama.cpp • AlgMor24
██████╗ ███████╗███████╗███████╗██╗ ██╗ ██╗ █████╗
██╔═══██╗██╔════╝██╔════╝██╔════╝██║ ██║ ██║██╔══██╗
██║ ██║█████╗ █████╗ █████╗ ██║ ██║ ██║███████║
██║ ██║██╔══╝ ██╔══╝ ██╔══╝ ██║ ██║ ██║██╔══██║
╚██████╔╝██║ ██║ ███████╗███████╗███████╗██║██║ ██║
╚═════╝ ╚═╝ ╚═╝ ╚══════╝╚══════╝╚══════╝╚═╝╚═╝ ╚═╝
High-Performance LLM / VLM Inference & Autonomous Agentic Ecosystem… See the full description on the dataset page: https://huggingface.co/datasets/Brunobkr/llama.cpp_AlgMor24_github.qwen-cpp-agent-0-protocolExperiment in agentic autonomy protocols.
~ everything in this repo was created by Qwen 3.8 27B (Q4) running autonomously inside Deepseek Harness, on a single RTX 3090 GPU, for 3 weeks.
The only human artifacts are:
agents/*
human/*
AGENTS.md
human_eval_cppmiomio_cp1_cachedemucs.cppThis repo stores weights in ggml format that are used to perform music separation.
These are intended to be used with demucs.cpp, https://github.com/sevagh/demucs.cpp
Weights origin:
https://dl.fbaipublicfiles.com/demucs/hybrid_transformer/955717e8-8726e21a.th
https://dl.fbaipublicfiles.com/demucs/hybrid_transformer/5c90dfd2-34c22ccb.th
https://dl.fbaipublicfiles.com/demucs/hybrid_transformer/f7e0c4bc-ba3fe64a.th
https://dl.fbaipublicfiles.com/demucs/hybrid_transformer/d12395a8-e57c48e6.th… See the full description on the dataset page: https://huggingface.co/datasets/Retrobear/demucs.cpp.qwentts-cpp-python-wheels
qwentts-cpp-python wheels
Optional backend-specific wheel variants for qwentts-cpp-python.
The default PyPI package is CUDA 12.8:
pip install qwentts-cpp-python
Install a backend-specific wheel from this repository with --find-links:
pip install "qwentts-cpp-python==0.3.1+cpu" -f https://huggingface.co/datasets/andito/qwentts-cpp-python-wheels/tree/main/whl/cpu
pip install "qwentts-cpp-python==0.3.1+cu124" -f… See the full description on the dataset page: https://huggingface.co/datasets/andito/qwentts-cpp-python-wheels.cpc-classificationswhisper.cpp
whisper.cpp
Stable: v1.8.1 / Roadmap
High-performance inference of OpenAI's Whisper automatic speech recognition (ASR) model:
Plain C/C++ implementation without dependencies
Apple Silicon first-class citizen - optimized via ARM NEON, Accelerate framework, Metal and Core ML
AVX intrinsics support for x86 architectures
VSX intrinsics support for POWER architectures
Mixed F16 / F32 precision
Integer quantization support
Zero memory allocations at runtime
Vulkan support
Support… See the full description on the dataset page: https://huggingface.co/datasets/echodict/whisper.cpp.fineweb-edu-zh-chengyu-cpt
Fineweb-Edu Chinese — Chengyu-Tagged Continued-Pretraining Corpus
A 3.74M-document Chinese corpus (~7.8B tokens) for continued pretraining on
cultural knowledge in figurative language, built from the highest-quality
tier of opencsg/Fineweb-Edu-Chinese-V2.1.
Each document is educational Chinese text containing at least one culturally
vetted chengyu, with an appended 【成语注释】 knowledge block listing every
matched idiom's figurative meaning(s) and classical source citation.
This is a… See the full description on the dataset page: https://huggingface.co/datasets/jiviteshjn/fineweb-edu-zh-chengyu-cpt.SWE-Bench-MultilingualC_CPPFileteredSWE-Bench-MultilingualC_CPPFiletered_newLiquidAI-Hackathon-Tokyo-CPT-Data
LiquidAI-Hackathon-Tokyo-CPT-Data
Liquid AI Hackathon Tokyoで作成したモデルのCPTに利用したデータセットです。
forums_pol_json_zstsih26099-cpse-material-codes
SIH 26099 — Collected Dataset
AI-Driven Standardization & Harmonization of Material Codes Across CPSEs
This workspace holds the data-collection stage only — no model, no training,
no feature engineering. Just raw public sources, their extracted structured
form, and the reference taxonomies/vocabularies the harmonisation step needs.
Collected live on 2026-09-08. All row counts below were verified by reading
the files back with pandas.
1. Headline numbers… See the full description on the dataset page: https://huggingface.co/datasets/sarthak20024/sih26099-cpse-material-codes.anioneYO-CPT-kk
YO-CPT-kk
YouTube-Oriented dataset for Continual Pre-Training (Kazakh). A heavily
quality-filtered corpus of Kazakh speech mined from YouTube and processed into clean, single-speaker,
TTS-grade utterances. Every utterance ships with an ensemble-verified transcription, a
punctuated/denormalized and stress-marked text variant, word-level forced alignment, within- and
cross-video speaker identities, an audio-quality (MOS) score, and a speaker persona built from the
voice and, where… See the full description on the dataset page: https://huggingface.co/datasets/NCSpeech/YO-CPT-kk.CPRet-data
CPRet-data
This repository hosts the datasets for CPRet: A Dataset, Benchmark, and Model for Retrieval in Competitive Programming.
Visit https://cpret.online/ to try out CPRet in action for competitive programming problem retrieval.
💡 CPRet Benchmark Tasks
The CPRet dataset supports four retrieval tasks relevant to competitive programming:
Text-to-Code Retrieval
Retrieve relevant code snippets based on a natural language problem description.
Code-to-Code Retrieval… See the full description on the dataset page: https://huggingface.co/datasets/coldchair16/CPRet-data.cpds_embeddingsthe-stack-v2-new-cppcpsc5800-hand-detection
Training Datasets and Model Weights for CPSC 5800 Final Project
Project repository: https://github.com/rohanphanse/CPSC5800-Final
We provide all training datasets created in Step 1 and weights for the YOLO and ResNet models trained during Steps 2-4 in our Hugging Face repository: https://huggingface.co/datasets/rohanphanse/cpsc5800-hand-detection
# Recommended: download dataset using git-xet (https://hf.co/docs/hub/git-xet)
brew install git-xet
git xet install
# Download datasets… See the full description on the dataset page: https://huggingface.co/datasets/rohanphanse/cpsc5800-hand-detection.cppe-5CPPE - 5 (Medical Personal Protective Equipment) is a new challenging dataset with the goal
to allow the study of subordinate categorization of medical personal protective equipments,
which is not possible with other popular data sets that focus on broad level categories.
