parquet
rl_warmup_rlve_offline_20K-parquet_qwen3-1.7b_epoch_1_maskrecipe-llm-ggufQwen3-4B_ds-parquet-programs-c24Qwen3-4B_ds-parquet-programs-c1403Qwen3-4B_ds-parquet-programs-c12mixed_sft_kukurasu_20K_qwen3_4b_thinking-parquet_nemotron-cascade-8b_epoch_3opd_rlve_rose_20K-parquet_qwen3-1.7b_epoch_1_mask_qwen3-4b-think_resp16384-T1.0-n8-topk16Qwen3-4B_ds-parquet-programs-c2806
Datasets
All datasets matching “parquet”dclm-baseline-1.0-parquet
DCLM-baseline
Note: this is an identical copy of https://huggingface.co/datasets/mlfoundations/dclm-baseline-1.0, where all the files have been mapped to a parquet format.
DCLM-baseline is a 4T token / 3B document pretraining dataset that achieves strong performance on language model benchmarks.
Below are comparisions of model trained on DCLM-baseline with other models in the 7B regime.
Model
Params
Tokens
Open dataset?
CORE
MMLU
EXTENDED
Open weights, closed datasets… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dclm-baseline-1.0-parquet.ettin-parquettest_librispeech_parquetraw_v0.1_parquet
Common Pile v0.1 — Parquet Consolidated
Description
This dataset bundles all “raw” corpora from the Common Pile v0.1 Raw Data collection, converted to Apache Parquet and consolidated in a single repository.
Nothing has been filtered or modified; the only changes are:
Format: original JSON → Parquet
Layout: many repositories → one consolidated dataset
Extra column: a len_category bucket for quick length-based filtering
Only the three original columns (id, text… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/raw_v0.1_parquet.Youtube-Common-First-600-Parqueteuler-source-parquets-real
