yassinsiouda/minimind-fr-pretrain-enfr-data
minimind-fr-pretrain-enfr-data pretrain_enfr.jsonl — {"text": "..."} per line, 5,315,952 lines, ~4 GB. The exact corpus used for minimind-fr-pretrain-enfr. See the model card for the build recipe. Upstream: allenai/c4 (ODC-BY). Built from allenai/c4 Format: line-delimited JSON. SFT rows use MiniMind's SFTDataset schema — {"conversations": [{role, content, reasoning_content, tools, tool_calls}]} (all string fields; tools/tool_calls are JSON strings;… See the full description on the dataset page: https://huggingface.co/datasets/yassinsiouda/minimind-fr-pretrain-enfr-data.
This repository belongs to yassinsiouda on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
