Lynote/humanize-text-model
Humanize Text Model (Lynote)
A lightweight bilingual (English/Chinese) humanizer: two small T5-seq2seq checkpoints in one repository that rewrite AI-flavored prose into more natural, human-flavored prose.
- `en/` — fine-tuned
google-t5/t5-small(English) - `zh/` — fine-tuned
uer/t5-small-chinese-cluecorpussmall(Chinese)
Try it live: ✍️ Free AI Humanizer Space · 🔍 Free AI Detector Space · 🖼️ Free AI Image Detector Space · 📝 Free AI Note Taker Space
Source & full product: github.com/lynote-ai/humanize-text
What it does
- Removes high-confidence AI clichés and formulaic phrases (e.g. "it is important to note that", "moreover", "值得注意的是", "降本增效").
- Keeps already-human prose nearly untouched (identity learning).
- Preserves numbers, URLs, file paths, code and quoted text (protected with
PROTECTED_Nplaceholders during generation, restored afterwards). - Routes automatically by language (CJK ratio detection).
What it is NOT
This is a writing-quality aid, not a tool for evading AI detectors. Detector scores are probabilistic, and no humanizer can guarantee that text will be classified as human. Please use it responsibly: do not use it to misrepresent authorship in academic, legal, or disciplinary contexts.
Quickstart
pip install transformers torchfrom humanize import Humanizer # the wrapper bundled in this repo
h = Humanizer() # loads this repo (en/ and zh/ sub-checkpoints)
print(h.humanize(
"It is important to note that this robust solution serves as a "
"testament to our commitment. Moreover, we leverage cutting-edge "
"technology."
))
print(h.humanize("值得注意的是,我们通过赋能团队来实现降本增效。"))Raw transformers usage (no wrapper):
from transformers import T5ForConditionalGeneration, T5Tokenizer
model = T5ForConditionalGeneration.from_pretrained("Lynote/humanize-text-model/en")
tokenizer = T5Tokenizer.from_pretrained("Lynote/humanize-text-model/en")
inputs = tokenizer("It is important to note that this is robust.", return_tensors="pt")
print(tokenizer.decode(model.generate(**inputs, max_length=128)[0], skip_special_tokens=True))For Chinese use the zh/ sub-checkpoint with BertTokenizer.
Training data
The corpus is generated deterministically from the editing principles of the Lynote reference projects (humanize-text, humanize-text-skill, humanizer-lite):
- AI → human: formulaic clause combinations rewritten by a conservative rule engine,
- human → human (identity): clean prose unchanged, so the model learns not to rewrite good text,
- mixed: clean prose with one injected cliché that must be removed,
- protected spans: examples with URLs, numbers, code and quotes.
Reproduce:
python scripts/build_dataset.py --out data # 14.7k pairs (en + zh)
python scripts/train.py --lang en --epochs 3 # -> checkpoints/humanize-text-model/en
python scripts/train.py --lang zh --epochs 3 # -> checkpoints/humanize-text-model/zh
python scripts/evaluate.py # benchmark on held-out test
pytest tests/ # full test suiteEvaluation (held-out test, 500 AI->human + all identity/protected pairs)
Limitations
- Trained on synthetic text; real-world inputs may need light post-editing.
- One checkpoint per language (English / Chinese); other languages are not specifically trained.
- Long inputs are truncated to 256 tokens.
- It is a conservative editor: it will not add stylistic richness that is absent from the source text.
License
MIT. Base models: google-t5/t5-small (Apache-2.0) and uer/t5-small-chinese-cluecorpussmall (Apache-2.0).
