CoolFace
20 results

cxl

shadowcollecter /cxlssd-results CXL-SSD Archetype Routing — Simulation Results Stage-00 snapshot (2026-05-02) of all simulator outputs produced for the CXL-SSD page-oriented embedding lookup paper. Includes: MQSim Results/overall.txt, latency_result.txt per layout × ratio × mode Cylon FEMU benchmark CSVs (MIO latency, throughput, cache scan) MaxEmbed evaluation matrices COG / SeedExpand / BQP partition outputs Each subdirectory groups one experiment family. Logs are kept beside metrics; raw .crdownload and very… See the full description on the dataset page: https://huggingface.co/datasets/shadowcollecter/cxlssd-results.other0 likes394 downloads5mo agoHugging Facecxllin /minimathCondensed version of the meta-math dataset arxiv.org/abs/2309.12284 View the project page: https://meta-math.github.io/ Citation @article{yu2023metamath, title={MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models}, author={Yu, Longhui and Jiang, Weisen and Shi, Han and Yu, Jincheng and Liu, Zhengying and Zhang, Yu and Kwok, James T and Li, Zhenguo and Weller, Adrian and Liu, Weiyang}, journal={arXiv preprint arXiv:2309.12284}, year={2023} } text10K<n<100K1 likes103 downloads3y agoHugging Facecxllin /medinstructv2text10K<n<100K5 likes85 downloads3y agoHugging Facecxllin /medinstruct Dataset Sources Repository: [https://github.com/jind11/MedQA] Paper : [https://arxiv.org/abs/2009.13081] Citation @article{jin2020disease, title={What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams}, author={Jin, Di and Pan, Eileen and Oufattole, Nassim and Weng, Wei-Hung and Fang, Hanyi and Szolovits, Peter}, journal={arXiv preprint arXiv:2009.13081}, year={2020} } textquestion-answering10K<n<100K4 likes82 downloads3y agoHugging Faceshadowcollecter /cxlssd-raw-criteo-tb CXL-SSD Archetype Routing — Criteo Terabyte (raw) Stage-00 snapshot (2026-05-02). Single-dataset cold storage of the Criteo Terabyte CTR Logs as used by the CXL-SSD page-oriented embedding lookup paper. 381 GB compressed 24 days of click logs, 4 billion samples Used for the largest-vocab DLRM evaluation Companion repos Paper outer repo: https://github.com/shadowcollecter/cxlssd-archetype-routing Preprocessed split: HF shadowcollecter/cxlssd-archetype-processed (look… See the full description on the dataset page: https://huggingface.co/datasets/shadowcollecter/cxlssd-raw-criteo-tb.textother10M<n<100M0 likes71 downloads5mo agoHugging Faceshadowcollecter /cxlssd-processed CXL-SSD Archetype Routing — Preprocessed Datasets Stage-00 snapshot (2026-05-02). Intermediate MERCI / MaxEmbed / DLRM preprocessing artifacts: filtered CSVs, vocab tables, embedding-table indexes, partition outputs. These are the outputs of MERCI_page_aware/analysis/preprocess_*.py applied to each raw dataset; they are the input to trace generation (research_data/traces/) and to the Cylon/MQSim/MaxEmbed simulators. Tiers criteo_terabyte/ — 45 GB criteo_kaggle/ —… See the full description on the dataset page: https://huggingface.co/datasets/shadowcollecter/cxlssd-processed.other0 likes65 downloads5mo agoHugging Face