CoolFace
Datasetpublic

SOMIL366/4D4T

Four Dataset For Training 4D4T: 4-Domain Training Dataset A curated ~60GB corpus split into four balanced domains for training small language models. 📊 Domain Breakdown Domain Source Format Math openbmb/UltraData-Math data/math/math_train_shard_*.jsonl.gz History allenai/c4 (realnewslike) data/history_news/history_train_shard_*.jsonl.gz Science sentence-transformers/s2orc data/science/science_train_shard_*.jsonl.gz General… See the full description on the dataset page: https://huggingface.co/datasets/SOMIL366/4D4T.

sourceHugging Facemitupdated 5mo agoView on Hugging Face
0likes17downloads

SOMIL366/4D4T · main · files are served by the source, never re-hosted here