cxl-ssd
Datasets
All datasets matching “cxl-ssd”cxlssd-results
CXL-SSD Archetype Routing — Simulation Results
Stage-00 snapshot (2026-05-02) of all simulator outputs produced for the
CXL-SSD page-oriented embedding lookup paper.
Includes:
MQSim Results/overall.txt, latency_result.txt per layout × ratio × mode
Cylon FEMU benchmark CSVs (MIO latency, throughput, cache scan)
MaxEmbed evaluation matrices
COG / SeedExpand / BQP partition outputs
Each subdirectory groups one experiment family. Logs are kept beside
metrics; raw .crdownload and very… See the full description on the dataset page: https://huggingface.co/datasets/shadowcollecter/cxlssd-results.cxlssd-raw-criteo-tb
CXL-SSD Archetype Routing — Criteo Terabyte (raw)
Stage-00 snapshot (2026-05-02). Single-dataset cold storage of the
Criteo Terabyte CTR Logs as used by the CXL-SSD page-oriented
embedding lookup paper.
381 GB compressed
24 days of click logs, 4 billion samples
Used for the largest-vocab DLRM evaluation
Companion repos
Paper outer repo: https://github.com/shadowcollecter/cxlssd-archetype-routing
Preprocessed split: HF shadowcollecter/cxlssd-archetype-processed
(look… See the full description on the dataset page: https://huggingface.co/datasets/shadowcollecter/cxlssd-raw-criteo-tb.cxlssd-processed
CXL-SSD Archetype Routing — Preprocessed Datasets
Stage-00 snapshot (2026-05-02). Intermediate MERCI / MaxEmbed / DLRM
preprocessing artifacts: filtered CSVs, vocab tables, embedding-table
indexes, partition outputs.
These are the outputs of MERCI_page_aware/analysis/preprocess_*.py
applied to each raw dataset; they are the input to trace generation
(research_data/traces/) and to the Cylon/MQSim/MaxEmbed simulators.
Tiers
criteo_terabyte/ — 45 GB
criteo_kaggle/ —… See the full description on the dataset page: https://huggingface.co/datasets/shadowcollecter/cxlssd-processed.cxlssd-raw-medium
CXL-SSD Archetype Routing — Raw Datasets (medium tier)
Stage-00 snapshot (2026-05-02). Mirrors the research_data/raw/ tree
excluding criteo_terabyte/ (the 381 GB Terabyte dataset is hosted
in a separate repo: cxlssd-archetype-raw-criteo-tb).
Contains 27 source datasets used by the CXL-SSD page-oriented
embedding lookup paper:
Dataset
Size
Source
amazon
46 GB
UCSD Amazon Reviews
criteo_kaggle
33 GB
Kaggle Display Ad Challenge
alibaba_ifashion
31 GB
Alibaba iFashion… See the full description on the dataset page: https://huggingface.co/datasets/shadowcollecter/cxlssd-raw-medium.cxlssd-traces
CXL-SSD Archetype Routing — Generated Traces
Stage-00 snapshot (2026-05-02) of processed access traces used in the
CXL-SSD page-oriented embedding lookup paper.
These are MQSim/MaxEmbed-format traces derived from public recommendation
datasets (MovieLens, Criteo, Avazu, Taobao, Amazon, Yelp, …) and the Qwen
KV-Cache traces. They feed into the Cylon (FEMU) and MQSim simulators.
Companion repos
Component
URL
Outer paper repo… See the full description on the dataset page: https://huggingface.co/datasets/shadowcollecter/cxlssd-traces.
