datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cpuspec_cpu_branch_tracesdim58-cpuData-31cases
Dim58 CPU Data — 31 Cases
Dataset uploaded from:
/mnt/data/ubuntu/research/outputs/data_cpu_geodesic58
Dataset summary
Property
Value
Repository
hosseinbv/dim58-cpuData-31cases
Number of files
64
Total size
17.87 GB
Source folder
data_cpu_geodesic58
File types
Extension
File count
.npz
62
.json
1
.csv
1
Top-level contents
0000_internal_case1_data.npz
0001_internal_B_10.npz… See the full description on the dataset page: https://huggingface.co/datasets/hosseinbv/dim58-cpuData-31cases.carbon-cpu-enriched-sequences
carbon-cpu-enriched-sequences
A CPU-enriched subset of the carbon pretraining corpus (eukaryote_generator), combining original source fields with normalized sequences
and row-level features for quality analysis, GPU enrichment and embedding generation.
Information of Features
Feature
Type
Description
record_id
string
NCBI Identifier linking the row back to the source genomic record. It provides the primary record-level identity.
begin_of_sequence… See the full description on the dataset page: https://huggingface.co/datasets/AINovice2005/carbon-cpu-enriched-sequences.carbon-cpu-enriched-sequences-sampledglm-moe-dsa-tiny-cpu-repro-v1
Tiny GLM MoE DSA: two CPU captures, forced zero-KL replay
Reproducibility evidence for
malaiwah/glm-moe-dsa-tiny-random-bf16,
checkpoint/config/tokenizer revision 45563636ef723acfb826755493447dc40c7a0c37.
This is a synthetic pipeline test, not a quality benchmark, quantization measurement,
qualified production reference, or registry submission. The model is random-init.
No GPU or paid cloud job was used.
Observed result
Two fresh capture processes, two CPU… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/glm-moe-dsa-tiny-cpu-repro-v1.kimi-k25-tiny-cpu-repro-v1
Kimi K2.5 / K2.6 / K2.7-Code complete tiny random BF16 CPU fixture
Untrained independently seeded random weights; no upstream weights or training data.
This is a reproducibility fixture, not useful language modeling or production quality evidence.
Runtime and lineage
Upstream moonshotai/Kimi-K2.7-Code@74797c9c62378b951a1f6fcf5c4631024e9b8bef.
Actual loaded class: Kimi_K25ForConditionalGeneration. Complete untied head and real small vision tower/projector.… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/kimi-k25-tiny-cpu-repro-v1.qwen3-5-tiny-cpu-repro-v1
Qwen3.5 tiny native random CPU fixture
Complete randomly initialized, untrained Qwen3_5ForConditionalGeneration checkpoint.
This is a pipeline/reproducibility fixture, not a useful language model, distillation,
quantization, quality benchmark, or claim about the performance of Qwen3.8-27B.
No upstream model weights or training data were used. No paid GPU/cloud compute.
Architecture and lineage
Architecture lineage: Qwen/Qwen3.8-27B at… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/qwen3-5-tiny-cpu-repro-v1.glm5-next-tiny-cpu-repro-v1This repository is an evidence bundle, not one root-format dataset at repository root.
first/ and repeat/ are separate complete sealed QFS root datasets; comparison/ holds
the comparison receipt and tokenwise result. panel/ is the sealed input panel. Other files
are provenance, logs and reproduction tools. Do not pass the bundle root as a QFS dataset.
GLM5-Next tiny native CPU fixture
This is a complete untrained random-initialized native Glm5NextForConditionalGeneration
wrapper… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/glm5-next-tiny-cpu-repro-v1.deepseek-v4-tiny-cpu-repro-v1
DeepSeek-V4 tiny corrected-native-primitives CPU text fixture
Complete randomly initialized, untrained QFSDeepseekV4ForCausalLM text class
using Transformers5.16.1 native primitives and a reviewed RMSNorm arithmetic correction.
No upstream weights, paid GPU/cloud compute or useful-model claim.
This is not unmodified native Transformers or the complete production release.
Architecture and scope
Text lineage:… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/deepseek-v4-tiny-cpu-repro-v1.k2-horizon-tiny-cpu-repro-v1
K2-Horizon MoVA tiny random CPU fixture
Complete untrained K2HorizonForCausalLM, not Moonshot Kimi despite the K2 name.
Architecture source: IFM/K2-Horizon-MoVA-36B-A4B at 05cab0a4d7150c1c460a000b37ff40cc1af2feaa.
No pretrained weights, original tokenizer, training data, paid GPU or cloud compute used.
Complete text-only K2HorizonForCausalLM, not Kimi: three-layer dense prefix followed by two real MoVA+MoE layers, grouped RMSNorm, sigmoid top-k routing with selection-only bias… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/k2-horizon-tiny-cpu-repro-v1.kimi-k3-tiny-cpu-repro-v1
Kimi K3 complete tiny random BF16 CPU fixture
Untrained independently seeded random weights; no upstream weights or training data.
This is a reproducibility fixture, not useful language modeling or production quality evidence.
Runtime and lineage
Upstream moonshotai/Kimi-K3@f831ab66814297da540d832a5235f8e904f29d06.
Actual loaded class: KimiK3ForConditionalGeneration. Complete untied head and real small vision tower/projector.
Parameters: 269688; vision parameters:… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/kimi-k3-tiny-cpu-repro-v1.minimax-m3-tiny-cpu-repro-v1
minimax-m3 complete native tiny random CPU fixture
Complete untrained MiniMaxM3SparseForConditionalGeneration checkpoint with an untied full LM head,
a real 272-entry byte tokenizer and every native state tensor. Architecture lineage:
MiniMaxAI/MiniMax-M3@f0e1c1e04d40177e4673a22097036854f536e9c0.
No upstream weights, training data, paid GPU or cloud compute were used.
Complete native image/text wrapper with real shrunk Conv3D vision, nonempty 3D RoPE, patch-merge projector and… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/minimax-m3-tiny-cpu-repro-v1.minimax-m2-tiny-cpu-repro-v1
minimax-m2 complete native tiny random CPU fixture
Complete untrained MiniMaxM2ForCausalLM checkpoint with an untied full LM head,
a real 272-entry byte tokenizer and every native state tensor. Architecture lineage:
MiniMaxAI/MiniMax-M2.7@d494266a4affc0d2995ba1fa35c8481cbd84294b.
No upstream weights, training data, paid GPU or cloud compute were used.
Complete native text causal LM: sigmoid/top-k MoE routing with correction bias, per-layer flattened Q/K RMSNorm and half-head… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/minimax-m2-tiny-cpu-repro-v1.installamacpp-cpuswe-agent-cpu-dynamorio-pilot-sympy-15599
One-agent CPU trace pilot
Preliminary research data. Validation is incomplete; this is not a confirmed dead-state result.
One live mini-SWE-agent 2.4.6 execution of sympy__sympy-15599, using a separate Qwen3-Coder-30B-A3B-Instruct AWQ server. The collector finished normally in 907.94 seconds. The agent made 57 model calls and submitted a patch; benchmark evaluation was not run. This is mini-SWE-agent, not the original full SWE-agent implementation.
What is included… See the full description on the dataset page: https://huggingface.co/datasets/harry1332/swe-agent-cpu-dynamorio-pilot-sympy-15599.spark2-5-tiny-cpu-repro-v1
Spark2.5 tiny random CPU fixture
Complete untrained Spark2_5ForCausalLM with independently seeded random BF16 weights.
This is a reproducibility fixture, not a useful language model, distilled model,
quality benchmark, or production registry measurement. No upstream weights, training
data, paid GPU or cloud rentals were used.
Architecture, code and license
Source: XHToken/Spark-X2.5-4B at 5e10fcc0286756aebf7c41dc52c1e42d95c70281.
The complete text causal model… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/spark2-5-tiny-cpu-repro-v1.qwen4-exp-tiny-cpu-repro-v1
Qwen4-Exp tiny CPU reproduction receipts
This is an artifact/receipt bundle, not training data and not a single root-format QFS dataset. All four readable synthetic documents are embedded in panel/panel.receipt.json.
Provenance and limitations
This is an independently generated, untrained random checkpoint inspired by Qwen/Qwen3.8-Flash-Next@de4b8e4d43b917e7706784d8bb445c9af86a3540, not a quantization, distillation, behavioral replica, or fine-tune. No source… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/qwen4-exp-tiny-cpu-repro-v1.ramanv-cpu-scriptsSWEbench-Verified-M2.7-orch-Django-CPU-quota-infra-repair-20260921
SWE-bench Verified: CPU-quota infrastructure repair
Seven predeclared task-attempts across five runs. Automatic Django test workers and numeric library threads now obey the existing 2-vCPU quota; memory remains 10 GiB. Model revisions, decoding, harness variant, grader and data are unchanged.
Two task-attempts have direct Docker/cgroup OOM evidence. Five historical tasks have early exit137 plus the same unbounded full-suite command; those are strong inferences with unavailable… See the full description on the dataset page: https://huggingface.co/datasets/CharlieLLL/SWEbench-Verified-M2.7-orch-Django-CPU-quota-infra-repair-20260921.qfs-affine-tiny-cpu-format-v1
affine tiny CPU format fixture reproducibility
Complete tiny random FORMAT fixture evidence. Round-to-nearest (RTN) storage/reader exercise only; optimizer-not-run. No GPTQ/AWQ/AutoRound optimization, calibrated ModelOpt/CT/QAT quality, trained-model quality ranking, GPU parity, or native serving-kernel correctness claim.
Reconstructed weights are evaluated by the captured native forward. KL is own-head, full-vocabulary on the recorded panel, not a benchmark of training quality.… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/qfs-affine-tiny-cpu-format-v1.qfs-microfloat-tiny-cpu-format-v1
microfloat tiny CPU format fixture reproducibility
Complete tiny random FORMAT fixture evidence. Round-to-nearest (RTN) storage/reader exercise only; optimizer-not-run. No GPTQ/AWQ/AutoRound optimization, calibrated ModelOpt/CT/QAT quality, trained-model quality ranking, GPU parity, or native serving-kernel correctness claim.
Reconstructed weights are evaluated by the captured native forward. KL is own-head, full-vocabulary on the recorded panel, not a benchmark of training… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/qfs-microfloat-tiny-cpu-format-v1.qfs-existing-tiny-cpu-format-v1
existing tiny CPU format fixture reproducibility
Complete tiny random FORMAT fixture evidence. Round-to-nearest (RTN) storage/reader exercise only; optimizer-not-run. No GPTQ/AWQ/AutoRound optimization, calibrated ModelOpt/CT/QAT quality, trained-model quality ranking, GPU parity, or native serving-kernel correctness claim.
Reconstructed weights are evaluated by the captured native forward. KL is own-head, full-vocabulary on the recorded panel, not a benchmark of training… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/qfs-existing-tiny-cpu-format-v1.qfs-qwen-gguf-tiny-cpu-format-v1
qwen-gguf tiny CPU format fixture reproducibility
Complete tiny random FORMAT fixture evidence. Round-to-nearest (RTN) storage/reader exercise only; optimizer-not-run. No GPTQ/AWQ/AutoRound optimization, calibrated ModelOpt/CT/QAT quality, trained-model quality ranking, GPU parity, or native serving-kernel correctness claim.
Reconstructed weights are evaluated by the captured native forward. KL is own-head, full-vocabulary on the recorded panel, not a benchmark of training… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/qfs-qwen-gguf-tiny-cpu-format-v1.voxcpm2-gguf-runtime-cpu-portable
VoxCPM2 runtime payload
This private Kaggle dataset is generated by Phorcys.Tools.VoxCPM2RuntimeUploader for PHRunner.Kaggle.Service.VoxCPM2.
Runtime flavor: LinuxCpu
Python tag: python3.10
Generated UTC: 2026-09-09T14:18:24.2697296+00:00
The dataset intentionally contains runtime artifacts, not model weights. Keep the official openbmb/VoxCPM2 checkpoint snapshot in a separate private Kaggle dataset, for example kaggle-pool-account/voxcpm2-python-models.
Top-level runtime… See the full description on the dataset page: https://huggingface.co/datasets/stokiz/voxcpm2-gguf-runtime-cpu-portable.CelebAHairMask-HQ
CelebAHairMask-HQ
CelebAHairMask-HQ is a extended dataset of CelebAMask-HQ for hair segmentation or hair matting.
CelebAMask-HQ is a large-scale face image dataset that has 30,000 high-resolution face images selected from the CelebA dataset by following CelebA-HQ. Each image has segmentation mask of facial attributes corresponding to CelebA.
The masks of CelebAHairMask-HQ were auto-annotated with the size of 1024 x 1024.
CelebAHairMask-HQ can be used to train and evaluate… See the full description on the dataset page: https://huggingface.co/datasets/cpuimage/CelebAHairMask-HQ.MixSub-LLaMA-3.2-Text-Only-Overlap-CPU-Scorevitpose-hand
ViTPose pre-trainted models on InterHand2.6m dataset
Form https://github.com/ViTAE-Transformer/ViTPose/issues/74
https://arxiv.org/abs/2204.12484
https://github.com/ViTAE-Transformer/ViTPose
harmony-nemotron-cpu-artifacts
Harmony CPU artifacts: Nemotron datasets (normalized + candidate pools)
This dataset repo is an artifact store produced on an EPYC CPU box. It contains:
normalized/ — CPU-normalized Parquet shards with a text-first Harmony format (text) plus meta_* and quality_* fields.
pools/ — candidate pool Parquet shards (subsets) for later GPU scoring (Modal NLL/PPL). No GPU scoring has been run yet.
reports/ — summary tables of counts per dataset/split/pool.
Directory layout… See the full description on the dataset page: https://huggingface.co/datasets/radna0/harmony-nemotron-cpu-artifacts.swe-agent-cpu-dynamorio-validated-private
One-agent CPU trace pilot — converted traces
Private research dataset: validated canonical (converted) DynamoRIO CPU traces, not the original offline raw recordings. The canonical files are binary ZIP traces for DynamoRIO analysis tools, not human-readable CSV.
Upload status
The transfer is complete only when UPLOAD_COMPLETE.json exists and reports status: complete. Until then the repository may contain only part of the 358 expected canonical files. Final… See the full description on the dataset page: https://huggingface.co/datasets/harry1332/swe-agent-cpu-dynamorio-validated-private.
