fsdp
Datasets
All datasets matching “fsdp”FSD-PT-PleIAs-SYNTH-15M-EN-NonMem-Mem
FSD-PT PleIAs SYNTH 15M English Subset
A 15M-document English-only subset of PleIAs/SYNTH for pretraining small reasoning models.
Source
Original dataset: PleIAs/SYNTH (79.6M samples, 41B words, CC-BY-4.0).
Subsetting Method
Filtered to language == "en" only
Round-robin sampled across exercise types for diversity
Final count: 15,000,000 documents
Columns
Field
Type
Description
query
string
Input query/prompt
synthetic_reasoning
string… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/FSD-PT-PleIAs-SYNTH-15M-EN-NonMem-Mem.tt-x10-fsdp2-fa2llama_4_fsdpqwen3-0.6b-fsdp2fsdp_modelllama-fsdp-hw-stats
