CoolFace
Datasetpublic

yassinsiouda/minimind-fr-creative-data

minimind-fr-creative-data sft_spec_creative.jsonl (135,078) — conversations schema. French-native creative writing (french_instruct + French-Alpaca + French-PD-Books continuations) + ~37.5% replay of the base SFT mix. Built by scripts/convert_spec_creative.py; see the minimind-fr-creative model card for full upstream links. Built from angeluriot/french_instruct jpacifico/French-Alpaca-dataset-Instruct-110K PleIAs/French-PD-Books allenai/tulu-3-sft-mixture… See the full description on the dataset page: https://huggingface.co/datasets/yassinsiouda/minimind-fr-creative-data.

sourceHugging Faceapache-2.0updated 21d agoView on Hugging Face
0likes41downloads
Dataset Card

minimind-fr-creative-data

sft_spec_creative.jsonl (135,078) — conversations schema. French-native creative writing (frenchinstruct + French-Alpaca + French-PD-Books continuations) + ~37.5% replay of the base SFT mix. Built by `scripts/convertspec_creative.py; see the minimind-fr-creative` model card for full upstream links.

Built from

Format: line-delimited JSON. SFT rows use MiniMind's SFTDataset schema — {"conversations": [{role, content, reasoning_content, tools, tool_calls}]} (all string fields; tools/tool_calls are JSON strings; reasoning_content renders in <think>…</think>). All SFT files were filtered to 0% zero-supervised rows at the trainer's max_seq_len.

Used to train `yassinsiouda/minimind-fr-creative`. Training framework: <https://github.com/jingyaogong/minimind>.

License

Apache-2.0 for the derived blend. Upstream dataset licenses apply — c4 (ODC-BY), tulu-3 (ODC-BY), Nemotron-SFT-SWE (CC-BY-4.0), French-PD-Books (public domain), and the others per their cards.