jerry982/BioFactory-Asian-Oral-Microbiome-Preview
🦷 World's First Chinese-Anchored Oral Multi-Omics Synthetic Dataset (v2026) 3,872 production records · PERMANOVA p=0.994 · MMD=0.060 · 100% synthetic = zero GDPR risk 📄 Full Whitepaper · 📋 1-Page Executive Summary · 🔬 38-Sample Preview · 📧 Enterprise: jerry820402@hotmail.com English Executive Summary This is the world's first Asian/Chinese-specific oral microbiome multi-omics synthetic dataset, generated via Evo foundation model inference on 8× NVIDIA A800… See the full description on the dataset page: https://huggingface.co/datasets/jerry982/BioFactory-Asian-Oral-Microbiome-Preview.
🦷 World's First Chinese-Anchored Oral Multi-Omics Synthetic Dataset (v2026)
3,872 production records · PERMANOVA p=0.994 · MMD=0.060 · 100% synthetic = zero GDPR risk
📄 **Full Whitepaper** · 📋 **1-Page Executive Summary** · 🔬 **38-Sample Preview** · 📧 Enterprise: jerry820402@hotmail.com
English Executive Summary
This is the world's first Asian/Chinese-specific oral microbiome multi-omics synthetic dataset, generated via Evo foundation model inference on 8× NVIDIA A800 DDP. Each record links genomic sequence + 10-taxon abundance profile + clinical triplet (periodontal / cardiovascular / diabetes).
After Cyclic Dirichlet manifold calibration, the full 3,872-record cohort passed enterprise B2B audit: PERMANOVA p=0.994, MMD=0.060, clinical triplet p=0.415 — statistically indistinguishable from real Han-Chinese oral cohorts.
100% synthetic · zero natural persons · GDPR / cross-border clinical data exempt.
中文摘要
全球首个华人锚定口腔多组学全合成数据集,基于 8 卡 A800 + Evo 分布式推理,每条样本包含基因序列、菌群丰度、临床三元组标签。经 Cyclic Dirichlet 流形校准后,3872 条生产数据终局审计 OVERALL PASSED(PERMANOVA p=0.994,MMD=0.060)。100% 合成,零合规风险,面向药企 AI 训模与口腔器械智能化场景。
Audit Bulletin (Production-Verified)
Full production set: n=3,872 · Reference cohort: n=30 Han-Chinese clinical samples
Repository Files
Quick Start (30 Seconds)
git clone https://huggingface.co/datasets/jerry982/BioFactory-Asian-Oral-Microbiome-Preview
cd BioFactory-Asian-Oral-Microbiome-Preview
python validate_preview.pyimport json
with open("preview_dataset.jsonl") as f:
for line in f:
rec = json.loads(line)
print(rec["schema:identifier"], rec["abundanceIndex"], rec["periodontalCorrelation"])Record Schema (Each JSON-L Line)
Full Commercial License
Contact: jerry820402@hotmail.com
Citation
@dataset{biofactory_asian_oral_2026,
title={Asian Oral Microbiome Multi-Omics Synthetic Dataset (v2026)},
author={AI Bio-Data Factory Lab},
year={2026},
url={https://huggingface.co/datasets/jerry982/BioFactory-Asian-Oral-Microbiome-Preview}
}Compliance
100% synthetic data — maps to zero natural persons — inherently exempt from GDPR and cross-border clinical genetic data restrictions.
