CoolFace
Datasetpublic

yassinsiouda/minimind-fr-pretrain-enfr-data

minimind-fr-pretrain-enfr-data pretrain_enfr.jsonl — {"text": "..."} per line, 5,315,952 lines, ~4 GB. The exact corpus used for minimind-fr-pretrain-enfr. See the model card for the build recipe. Upstream: allenai/c4 (ODC-BY). Built from allenai/c4 Format: line-delimited JSON. SFT rows use MiniMind's SFTDataset schema — {"conversations": [{role, content, reasoning_content, tools, tool_calls}]} (all string fields; tools/tool_calls are JSON strings;… See the full description on the dataset page: https://huggingface.co/datasets/yassinsiouda/minimind-fr-pretrain-enfr-data.

sourceHugging Faceapache-2.0updated 24d agoView on Hugging Face
1likes45downloads
Dataset Card

minimind-fr-pretrain-enfr-data

pretrain_enfr.jsonl — {"text": "..."} per line, 5,315,952 lines, ~4 GB. The exact corpus used for minimind-fr-pretrain-enfr. See the model card for the build recipe. Upstream: allenai/c4 (ODC-BY).

Built from

Format: line-delimited JSON. SFT rows use MiniMind's SFTDataset schema — {"conversations": [{role, content, reasoning_content, tools, tool_calls}]} (all string fields; tools/tool_calls are JSON strings; reasoning_content renders in <think>…</think>). All SFT files were filtered to 0% zero-supervised rows at the trainer's max_seq_len.

Used to train `yassinsiouda/minimind-fr-pretrain-enfr`. Training framework: <https://github.com/jingyaogong/minimind>.

License

Apache-2.0 for the derived blend. Upstream dataset licenses apply — c4 (ODC-BY), tulu-3 (ODC-BY), Nemotron-SFT-SWE (CC-BY-4.0), French-PD-Books (public domain), and the others per their cards.