datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
YO-CPT-ru
YO-CPT-ru
YouTube-Oriented dataset for Continual Pre-Training (Russian). A large, heavily
quality-filtered corpus of Russian speech mined from YouTube (via YODAS2)
and processed into clean, single-speaker, TTS-grade utterances. Every utterance ships with an
ensemble-verified transcription, a punctuated/denormalized and stress-marked text variant, word-level
forced alignment, within- and cross-video speaker identities, an audio-quality (MOS) score, and a
speaker persona built… See the full description on the dataset page: https://huggingface.co/datasets/NCSpeech/YO-CPT-ru.SWE-smith-cppcppe-5
Dataset Card for CPPE - 5
Dataset Summary
CPPE - 5 (Medical Personal Protective Equipment) is a new challenging dataset with the goal to allow the study of subordinate categorization of medical personal protective equipments, which is not possible with other popular data sets that focus on broad level categories.
Some features of this dataset are:
high quality images and annotations (~4.6 bounding boxes per image)
real-life images unlike any current such dataset
majority… See the full description on the dataset page: https://huggingface.co/datasets/rishitdagli/cppe-5.arxiv_cplusplus_research_code
Dataset card for ArtifactAI/arxiv_cplusplus_research_code
Dataset Description
https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_cplusplus_research_code
Dataset Summary
ArtifactAI/arxiv_python_research_code contains over 10.6GB of source code files referenced strictly in ArXiv papers. The dataset serves as a curated dataset for Code LLMs.
How to use it
from datasets import load_dataset
# full dataset (10.6GB of data)
ds =… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_cplusplus_research_code.SWE-Bench-MultilingualC_CPPFileteredLiquidAI-Hackathon-Tokyo-CPT-Data
LiquidAI-Hackathon-Tokyo-CPT-Data
Liquid AI Hackathon Tokyoで作成したモデルのCPTに利用したデータセットです。
SWE-Bench-MultilingualC_CPPFiletered_newcpc-classificationsbls_cpi
Changelog
2025-01-18
I decided that I'll name the column survey, instead of consumer. I'll also set the value to the description, instead of the code.
I didn't realize that the pandas version of "Use this dataset" includes the filename. I'll remove the date from the filename. So that people do not have to change their code.
I have been updating this using the UI, but I will create a script in Spaces to update this from the BLS site.
2025-01-12
While using… See the full description on the dataset page: https://huggingface.co/datasets/robert-co/bls_cpi.YO-CPT-kk
YO-CPT-kk
YouTube-Oriented dataset for Continual Pre-Training (Kazakh). A heavily
quality-filtered corpus of Kazakh speech mined from YouTube and processed into clean, single-speaker,
TTS-grade utterances. Every utterance ships with an ensemble-verified transcription, a
punctuated/denormalized and stress-marked text variant, word-level forced alignment, within- and
cross-video speaker identities, an audio-quality (MOS) score, and a speaker persona built from the
voice and, where… See the full description on the dataset page: https://huggingface.co/datasets/NCSpeech/YO-CPT-kk.CPRet-data
CPRet-data
This repository hosts the datasets for CPRet: A Dataset, Benchmark, and Model for Retrieval in Competitive Programming.
Visit https://cpret.online/ to try out CPRet in action for competitive programming problem retrieval.
💡 CPRet Benchmark Tasks
The CPRet dataset supports four retrieval tasks relevant to competitive programming:
Text-to-Code Retrieval
Retrieve relevant code snippets based on a natural language problem description.
Code-to-Code Retrieval… See the full description on the dataset page: https://huggingface.co/datasets/coldchair16/CPRet-data.the-stack-v2-new-cppindic-oss-mixture-cpt-10bthe-stack-v2-cppcarbon-cpu-enriched-sequences
carbon-cpu-enriched-sequences
A CPU-enriched subset of the carbon pretraining corpus (eukaryote_generator), combining original source fields with normalized sequences
and row-level features for quality analysis, GPU enrichment and embedding generation.
Information of Features
Feature
Type
Description
record_id
string
NCBI Identifier linking the row back to the source genomic record. It provides the primary record-level identity.
begin_of_sequence… See the full description on the dataset page: https://huggingface.co/datasets/AINovice2005/carbon-cpu-enriched-sequences.carbon-cpu-enriched-sequences-sampledjump-cp-0016-labelfree-MitoRSVQA-HR_qwen_finetuningcsharp-dotnet-cpt-v0
csharp-dotnet-cpt-v0
A reproducible C# / .NET continued-pretraining (CPT) corpus for Qwen2.5-Coder-1.5B.
HF: https://huggingface.co/datasets/Nottybro/csharp-dotnet-cpt-v0 (private)
Total unique tokens: 695,288,168 (Qwen2.5-Coder-1.5B tokenizer)
Files: 681,426 | Repositories: 68,869
Format: Zstandard-compressed Parquet, schema below.
See DATASET_CARD.md for sources, license policy, filtering, limitations.
See reports/summary.md for full statistics.
stack-v2-cpp-2019Swallow-Nemotron-Post-Training-Dataset-v1-ja-cpt
Swallow-Nemotron-Post-Training-Dataset-v1-ja-cpt
Dataset Overview
This dataset is a reformatted subset of the tokyotech-llm/Swallow-Nemotron-Post-Training-Dataset-v1 dataset, specifically derived from the v1-Ja-202601 subset. It was created to facilitate Continuous Pre-Training (CPT) by extracting only the text_gpt_oss field from the original data.
Dataset Statistics & Token Counts
The token counts for each category were calculated using the… See the full description on the dataset page: https://huggingface.co/datasets/Podtech/Swallow-Nemotron-Post-Training-Dataset-v1-ja-cpt.dclm-164k-300m-raw-cpr120-ml1024dclm-164k-real-train-8b-instruct-hq-cpr32-ml1024dclm-164k-instruct-hq-cpr128-ml512-92-127CPT_textbooksarc-stack-cppCPIBench-0
CPIBench-0
CPIBench-0 is our first benchmark for evaluating how frontier models and agent harnesses perform in operations that rely heavily on cyber-physical systems.
Built from real projects across mechatronics, manufacturing, materials, and energy, CPIBench-0 tests models and agent harnesses in multimodal workflows where cyber-physical signals, tools, and real-world consequences intersect.
CPIBench-0 measures performance across connected cyber-physical workflows, from making… See the full description on the dataset page: https://huggingface.co/datasets/rayrren/CPIBench-0.optim_policy_pretrain-pythia-160m_lr0.0001_bs24_wp1_wd0.01_ep0_cp35k-mergeddclm-164k-300m-cpr200-ml1024LSM-Tokenized-Full
📡 Large Spectrum Models (LSM) – Tokenized RF Dataset
Overview
This repository provides the tokenized RF spectrum dataset introduced in the paper:
“Large Spectrum Models (LSMs): Decoder-Only Transformer-Powered Spectrum Activity Forecasting via Tokenized RF Data”
The dataset is designed to bridge wireless signal processing and large language models (LLMs) by converting raw spectrum measurements into discrete token sequences, enabling direct use with transformer-based… See the full description on the dataset page: https://huggingface.co/datasets/cpnlab/LSM-Tokenized-Full.
