tier-2
Datasets
All datasets matching “tier-2”datakit-tier2-skewed-v2
Datakit Tier2 Skewed Synthetic
A synthetic, heavy-tailed-document dataset generated for stress-testing the
Marin datakit pipeline (normalize / minhash / fuzzy_dups / consolidate /
tokenize) against doc-length outliers up to 256 MB.
This is a CI / pipeline-stress dataset, not a training corpus. It exists
to exercise long-tail code paths that the FineWeb-Edu smoke ferry doesn't
cover. Don't use it for training without re-evaluating its distribution.
Provenance
Source… See the full description on the dataset page: https://huggingface.co/datasets/ravwojdyla/datakit-tier2-skewed-v2.tier2_writertier2-autolabeled-lastMultiChallenge-Tier2-AdvanceMultiChallenge-Tier2-Coretier2_vscode
