dlm
Datasets
All datasets matching “dlm”dlm-exp1-drafter-voting-research-artifacts
dlm-exp1 drafter-voting: research artifacts (public mirror)
Everything bulky from the drafter-voting branch of https://github.com/NoviceCoderInfinity/dlm-general-expert-exp1
that git does not track (uploaded 2026-09-06 before the compute instance was destroyed):
logs/research_v1/: baseline cells (one JSON line per item; per-pass NPZ traces for both checkpoints in
baselines_traces/ and baselines_b32_traces/), state-batch fidelity probe states (fidelity/), flush-fusion
boundaries… See the full description on the dataset page: https://huggingface.co/datasets/Arushhh/dlm-exp1-drafter-voting-research-artifacts.dlm-exp1-drafter-voting-research-artifacts
dlm-exp1 drafter-voting research artifacts (public)
Draft-aligned judge LoRA for SDAR-30B-A3B-Chat-b32 with a 1.7B drafter and DES expert pooling: goal = decode speedup with minimal accuracy loss vs the stock model.
model / setting
weights
eval records
stock SDAR-30B-A3B-Chat-b32 (+ 1.7B drafter where used)
JetLM/SDAR-30B-A3B-Chat-b32 @ c351bbc3… (Hugging Face)
reference_evals/da_decode (*_stock), r3_A/evals/stock
stock + DES32 / DES48 (decode-time expert pooling… See the full description on the dataset page: https://huggingface.co/datasets/Anupam-Rawat-IITB/dlm-exp1-drafter-voting-research-artifacts.DLM-Decoding-Analysis
DLM-Decoding-Analysis
Diffusion Language Model Knows the Answer Before It Decodes
Pengxiang Li*, Yefan Zhou*, Dilxat Muhtar, Lu Yin, Shilin Yan, Li Shen, Yi Liang, Soroush Vosoughi, Shiwei Liu
The Fourteenth International Conference on Learning Representations (ICLR 2026)
TL;DR: Diffusion language models often commit to the correct answer
well before they finish decoding. This dataset releases the per-question,
step-by-step decoding trajectories of LLaDA-8B-Instruct on… See the full description on the dataset page: https://huggingface.co/datasets/YefanZhou98/DLM-Decoding-Analysis.meteora-dlmm-historical-data
Meteora DLMM Historical Data
Decoded Solana mainnet instructions and events from Meteora DLMM (Dynamic Liquidity Market Maker), a concentrated-liquidity DEX where liquidity sits in discrete price bins and the fee rate rises with volatility.
74 tables, 59,575 rows, one row per decoded instruction or event. Program ID LBUZKhRxPF3XUpBCjp4YzTKgLccjZhTSDM9YuVaPwxo.
This is a free sample from datastore.sh, which publishes the complete history as versioned Parquet.
Read… See the full description on the dataset page: https://huggingface.co/datasets/DataStore/meteora-dlmm-historical-data.DLM_DataSet
DLM DATASET
Large-scale multi-lingual (EN, KK, RU) & code-centric corpus for ML.
SCALE
1M-10M
FORMAT
JSONL
LANGUAGES
EN / KK / RU
🔍 Data Schema
The dataset utilizes a robust structure optimized for fast parsing:
id: Unique sample identifier
text: Main text or code payload
language: Language tag (en, kk, ru)
prog_lang: Python/JS markers/C# Unity/C++/Java
category: logic… See the full description on the dataset page: https://huggingface.co/datasets/DLMveloper/DLM_DataSet.dl-model
