Gugu8/English-Mini
LLM-English-100MB — Compact & Dense English Teaching Corpus A 100MB, extremely clean CSV designed to teach an LLM English from scratch via instruction-tuning. No noise, no HTML, no duplicates — just pure grammar, vocabulary, and syntax transformations. Generated with a single paste-and-run Python script in Google Colab. Why this teaches English Instead of raw text, the dataset is instruction -> input -> output pairs that force the model to learn rules: Grammar… See the full description on the dataset page: https://huggingface.co/datasets/Gugu8/English-Mini.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face