base
Datasets
All datasets matching “base”dclm-baseline-1.0
DCLM-baseline
DCLM-baseline is a 4T token / 3B document pretraining dataset that achieves strong performance on language model benchmarks.
Below are comparisions of model trained on DCLM-baseline with other models in the 7B regime.
Model
Params
Tokens
Open dataset?
CORE
MMLU
EXTENDED
Open weights, closed datasets
Llama2
7B
2T
✗
49.2
45.8
34.1
DeepSeek
7B
2T
✗
50.7
48.5
35.3
Mistral-0.3
7B
?
✗
57.0
62.7
45.1
QWEN-2
7B
?
✗
57.5
71.9
50.5
Llama3
8B
15T
✗… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dclm-baseline-1.0.dcvlm-baseline-200b
DCVLM-Baseline (200B tokens)
DCVLM-Baseline is the reference training mixture from our DataComp-VLM paper.
It is a pre-mixed, decontaminated, ready-to-train multimodal pretraining dataset, materialized as flat
WebDataset tar shards so it can be consumed by any training
stack.
This is a 200B-token dataset release consisting of 103,985,276 samples, curated from our DCVLM-large data pool.
A smaller 6.25B-token version is also available.
⚠️ NOTE: The training data is the WebDataset… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dcvlm-baseline-200b.knowledge-base
RL-for-LLMs Wiki
An expert-level, citation-backed knowledge base on reinforcement learning for
large language models — RLHF, DPO and offline preference optimization, reward
modeling, RLVR and reasoning, training systems, and the failure modes — built
collaboratively by autonomous agents. Each topic article is a deep dive written
so you can learn the topic from it without reading the underlying papers, with
every non-obvious claim cited to a source. Every change lands through a… See the full description on the dataset page: https://huggingface.co/datasets/rl-llm-wiki/knowledge-base.cmp-v6-base108-renderbite-baseline
bite-baseline — artifacts for extreme (ternary) quantization of Qwen3.6-35B-A3B
Companion dataset for ihavespoons/bite — an open
pipeline for compressing a Mixture-of-Experts LLM (Qwen/Qwen3.6-35B-A3B, 35B total / ~3B
active, 256 experts) toward ternary {-1,0,+1} weights (1.71 bpw) via PTQ init +
quantization-aware distillation. See the repo's docs/report-extreme-quant-moe.md for the
full technical report.
Contents
Path
What it is
baseline.json… See the full description on the dataset page: https://huggingface.co/datasets/ihavespoons/bite-baseline.LET-Base-Dataset
LET:Full-Size Humanoid Robot Real-World Dataset
中文| [English]
LET Dataset is collected based on the full-size humanoid robot Kuavo 4 Pro covering real-world multi-task data across multiple scenarios and operation types. It is designed for robot manipulation, mobility, and interaction tasks, supporting scalable robot learning in real environments.
📋 Table of Contents
Key Features
Hardware Platform
Usage Guide
Tool… See the full description on the dataset page: https://huggingface.co/datasets/LejuRobotics/LET-Base-Dataset.
