SOMIL366/4D4T
Four Dataset For Training 4D4T: 4-Domain Training Dataset A curated ~60GB corpus split into four balanced domains for training small language models. 📊 Domain Breakdown Domain Source Format Math openbmb/UltraData-Math data/math/math_train_shard_*.jsonl.gz History allenai/c4 (realnewslike) data/history_news/history_train_shard_*.jsonl.gz Science sentence-transformers/s2orc data/science/science_train_shard_*.jsonl.gz General… See the full description on the dataset page: https://huggingface.co/datasets/SOMIL366/4D4T.
017
