pcm
Datasets
All datasets matching “pcm”PCMind-2.1-Kaiyuan-2B
This repository contains the complete pretraining dataset for
PCMind-v2.1-Kaiyuan-2B, a leading fully open-source language model.
Overview
The dataset is organized into 5 training phases, with all phase datasets open-sourced in this repository. Our training methodology employs domain-specific mixing strategies across five primary domains:
English: General English text
Chinese: General Chinese text
Code: Programming and code-related content
Math: Mathematical reasoning and… See the full description on the dataset page: https://huggingface.co/datasets/thu-pacman/PCMind-2.1-Kaiyuan-2B.PCMind-2.1-Kaiyuan-2B-phase1-part1-1-0323
This repository contains the complete pretraining dataset for
PCMind-v2.1-Kaiyuan-2B, a leading fully open-source language model.
Overview
The dataset is organized into 5 training phases, with all phase datasets open-sourced in this repository. Our training methodology employs domain-specific mixing strategies across five primary domains:
English: General English text
Chinese: General Chinese text
Code: Programming and code-related content
Math: Mathematical reasoning and… See the full description on the dataset page: https://huggingface.co/datasets/Lxd99/PCMind-2.1-Kaiyuan-2B-phase1-part1-1-0323.PCMind-2.1-Kaiyuan-2B-phase1-part1-2
This repository contains the complete pretraining dataset for
PCMind-v2.1-Kaiyuan-2B, a leading fully open-source language model.
Overview
The dataset is organized into 5 training phases, with all phase datasets open-sourced in this repository. Our training methodology employs domain-specific mixing strategies across five primary domains:
English: General English text
Chinese: General Chinese text
Code: Programming and code-related content
Math: Mathematical reasoning and… See the full description on the dataset page: https://huggingface.co/datasets/Lxd99/PCMind-2.1-Kaiyuan-2B-phase1-part1-2.pcmt-artifact
Proof-Carrying Multimodal Timelines Artifact
This Hugging Face Dataset repository hosts the runnable artifact for:
Proof-Carrying Multimodal Timelines: Finite-Trace Modal Certificates for Video-Audio Consistency
Authors: Faruk Alpay and Hamdi Alakkad.
The artifact is organized as a dataset-style file tree rather than a zip archive. It is intended to support an arXiv submission whose source package stays below arXiv's upload limit while keeping the full runnable code, traces… See the full description on the dataset page: https://huggingface.co/datasets/Lightcap/pcmt-artifact.PC-Mix
PC-Mix Dataset
Partial Spoof Dataset with Controlled Speech–Environmental Sounds Mixing
PC-Mix pairs speech from PartialSpoof v1.2 with a self-curated partial-spoof environmental sounds pool to create controlled speech–environment authenticity combinations. The original samples are collected from VGGSound and SBCSAE.
Resources
Resource
Description
Link
GitHub
Code, protocols, and preprocessing scripts.
Anonymous GitHub
Paper
Dataset description and… See the full description on the dataset page: https://huggingface.co/datasets/Alphawarheads/PC-Mix.cpython_dataset
