yassinsiouda/minimind-fr-pretrain-enfr-data
minimind-fr-pretrain-enfr-data pretrain_enfr.jsonl — {"text": "..."} per line, 5,315,952 lines, ~4 GB. The exact corpus used for minimind-fr-pretrain-enfr. See the model card for the build recipe. Upstream: allenai/c4 (ODC-BY). Built from allenai/c4 Format: line-delimited JSON. SFT rows use MiniMind's SFTDataset schema — {"conversations": [{role, content, reasoning_content, tools, tool_calls}]} (all string fields; tools/tool_calls are JSON strings;… See the full description on the dataset page: https://huggingface.co/datasets/yassinsiouda/minimind-fr-pretrain-enfr-data.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face