datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
YO-CPT-ru
YO-CPT-ru
YouTube-Oriented dataset for Continual Pre-Training (Russian). A large, heavily
quality-filtered corpus of Russian speech mined from YouTube (via YODAS2)
and processed into clean, single-speaker, TTS-grade utterances. Every utterance ships with an
ensemble-verified transcription, a punctuated/denormalized and stress-marked text variant, word-level
forced alignment, within- and cross-video speaker identities, an audio-quality (MOS) score, and a
speaker persona built… See the full description on the dataset page: https://huggingface.co/datasets/NCSpeech/YO-CPT-ru.CPT_Data_Pool
CPT Data Pool
This repository hosts a large-scale CPT corpus designed for continual pre-training of domain-specific large language models. It serves as a data component of the 👉F2D-LLM framework, an end-to-end pipeline for domain-specific LLM training.
For detailed data processing, scoring methods, and sampling strategies, please refer to the official 👉GitHub repo.
Overall, the dataset contains approximately 271B tokens and is stored as jsonl files with a unified schema… See the full description on the dataset page: https://huggingface.co/datasets/ZhejiangLab/CPT_Data_Pool.SWE-smith-cpparxiv_cplusplus_research_code
Dataset card for ArtifactAI/arxiv_cplusplus_research_code
Dataset Description
https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_cplusplus_research_code
Dataset Summary
ArtifactAI/arxiv_python_research_code contains over 10.6GB of source code files referenced strictly in ArXiv papers. The dataset serves as a curated dataset for Code LLMs.
How to use it
from datasets import load_dataset
# full dataset (10.6GB of data)
ds =… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_cplusplus_research_code.qwen-cpp-agent-0-protocolExperiment in agentic autonomy protocols.
~ everything in this repo was created by Qwen 3.8 27B (Q4) running autonomously inside Deepseek Harness, on a single RTX 3090 GPU, for 3 weeks.
The only human artifacts are:
agents/*
human/*
AGENTS.md
human_eval_cppcpc-classificationsfineweb-edu-zh-chengyu-cpt
Fineweb-Edu Chinese — Chengyu-Tagged Continued-Pretraining Corpus
A 3.74M-document Chinese corpus (~7.8B tokens) for continued pretraining on
cultural knowledge in figurative language, built from the highest-quality
tier of opencsg/Fineweb-Edu-Chinese-V2.1.
Each document is educational Chinese text containing at least one culturally
vetted chengyu, with an appended 【成语注释】 knowledge block listing every
matched idiom's figurative meaning(s) and classical source citation.
This is a… See the full description on the dataset page: https://huggingface.co/datasets/jiviteshjn/fineweb-edu-zh-chengyu-cpt.SWE-Bench-MultilingualC_CPPFileteredSWE-Bench-MultilingualC_CPPFiletered_newsih26099-cpse-material-codes
SIH 26099 — Collected Dataset
AI-Driven Standardization & Harmonization of Material Codes Across CPSEs
This workspace holds the data-collection stage only — no model, no training,
no feature engineering. Just raw public sources, their extracted structured
form, and the reference taxonomies/vocabularies the harmonisation step needs.
Collected live on 2026-09-08. All row counts below were verified by reading
the files back with pandas.
1. Headline numbers… See the full description on the dataset page: https://huggingface.co/datasets/sarthak20024/sih26099-cpse-material-codes.YO-CPT-kk
YO-CPT-kk
YouTube-Oriented dataset for Continual Pre-Training (Kazakh). A heavily
quality-filtered corpus of Kazakh speech mined from YouTube and processed into clean, single-speaker,
TTS-grade utterances. Every utterance ships with an ensemble-verified transcription, a
punctuated/denormalized and stress-marked text variant, word-level forced alignment, within- and
cross-video speaker identities, an audio-quality (MOS) score, and a speaker persona built from the
voice and, where… See the full description on the dataset page: https://huggingface.co/datasets/NCSpeech/YO-CPT-kk.CPRet-data
CPRet-data
This repository hosts the datasets for CPRet: A Dataset, Benchmark, and Model for Retrieval in Competitive Programming.
Visit https://cpret.online/ to try out CPRet in action for competitive programming problem retrieval.
💡 CPRet Benchmark Tasks
The CPRet dataset supports four retrieval tasks relevant to competitive programming:
Text-to-Code Retrieval
Retrieve relevant code snippets based on a natural language problem description.
Code-to-Code Retrieval… See the full description on the dataset page: https://huggingface.co/datasets/coldchair16/CPRet-data.the-stack-v2-new-cppCPC_Text_RoughCP2077
Cyberpunk 2077 controllable RGB-D capture
Each row in metadata.jsonl links a 16 FPS H.264 RGB video with bit-exact uint16 logarithmic depth PNGs, frame-aligned action/player/camera records, and 60 Hz controller records. See each clip manifest for depth decoding parameters and checksums.
dim58-cpuData-31cases
Dim58 CPU Data — 31 Cases
Dataset uploaded from:
/mnt/data/ubuntu/research/outputs/data_cpu_geodesic58
Dataset summary
Property
Value
Repository
hosseinbv/dim58-cpuData-31cases
Number of files
64
Total size
17.87 GB
Source folder
data_cpu_geodesic58
File types
Extension
File count
.npz
62
.json
1
.csv
1
Top-level contents
0000_internal_case1_data.npz
0001_internal_B_10.npz… See the full description on the dataset page: https://huggingface.co/datasets/hosseinbv/dim58-cpuData-31cases.spec_cpu_branch_tracesCpp-Code-LargeCpp-Code-Large
Cpp-Code-Large is a large-scale corpus of C++ source code comprising more than 5 million lines of C++ code. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and static program analysis for the C++ ecosystem.
By providing a high-volume, language-specific corpus, Cpp-Code-Large enables systematic experimentation in C++-focused model training, domain adaptation, and downstream code… See the full description on the dataset page: https://huggingface.co/datasets/ajibawa-2023/Cpp-Code-Large.the-stack-v2-cppllama.cppversion https://git-lfs.github.com/spec/v1
oid sha256:cfc44b7ba25614df70e6b65e3341cae0310163bd32fd31a6b928a542df433faf
size 30786
carbon-cpu-enriched-sequences
carbon-cpu-enriched-sequences
A CPU-enriched subset of the carbon pretraining corpus (eukaryote_generator), combining original source fields with normalized sequences
and row-level features for quality analysis, GPU enrichment and embedding generation.
Information of Features
Feature
Type
Description
record_id
string
NCBI Identifier linking the row back to the source genomic record. It provides the primary record-level identity.
begin_of_sequence… See the full description on the dataset page: https://huggingface.co/datasets/AINovice2005/carbon-cpu-enriched-sequences.indic-oss-mixture-cpt-10bcarbon-cpu-enriched-sequences-sampledCPRet-Embeddings
CPRet-Embeddings
This repository provides the problem descriptions and their corresponding precomputed embeddings used by the CPRet retrieval server.
You can explore the retrieval server via the online demo at https://cpret.online/.
📦 Files
probs_2609.jsonlA JSONL file containing natural language descriptions of competitive programming problems.Each line is a JSON object with metadata such as problem title, platform/source OJ, URL, and full description.… See the full description on the dataset page: https://huggingface.co/datasets/coldchair16/CPRet-Embeddings.LiveCodeBench-CPP
LiveCodeBench-CPP: An Extension of LiveCodeBench for Contamination Free Evaluation in C++
Overview
LiveCodeBench-CPP includes 454 problems from the release_v6 of LiveCodeBench, covering the period from October 2024 to May 2025. These problems are sourced from AtCoder (287 problems) and LeetCode (167 problems).
AtCoder Problems: These require generated solutions to read inputs from standard input (stdin) and write outputs to standard output (stdout). For unit testing, the… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/LiveCodeBench-CPP.Underwater-Acoustic-Channel-Repository
Underwater Acoustic Channel Repository
This Hugging Face dataset is a structured, checksum-preserving mirror of version 1.0 of the Underwater Acoustic Channel Repository. The original dataset was published by Zhengnan Li, Mandar Chitre, Diego Cuji, James Preisig, Andrew Singer, Milica Stojanovic, and Paul van Walree.
The collection contains measured underwater acoustic channel impulse responses (CIRs) from eight at-sea experimental groups. Channel and accompanying noise files… See the full description on the dataset page: https://huggingface.co/datasets/UWA-CP/Underwater-Acoustic-Channel-Repository.RSVQA-HR_qwen_finetuningMSMU
MSMU (Massive Spatial Measuring and Understanding Dataset for Spatial Intelligence)
🌐 Homepage | 🤗 Dataset | 📖 arXiv | GitHub
Dataset Details
Dataset Description
We introduce MSMU and MSMU-Bench: a new benchmark designed to enhance and evaluate multimodal models on spatial measuring and understanding. MSMU is featured as metric-accurate spatial annotations which are sourced from high-precision 3D scenes. It contains , 25K images, 700K QA pairs… See the full description on the dataset page: https://huggingface.co/datasets/cpystan/MSMU.arc-stack-cpp
