An AI detector gives you a number, and that number is increasingly used to make consequential accusations. This measures what it is worth. At each tool's own shipped threshold, measured here: the full local ensemble flags 17% of genuine human writing; the lite tier flagged 30% on conversational prose; one bundled detector flags 6 of 8 human documents and another flags 89% of them. The rewrite loop is the instrument rather than the product — score, edit under a meaning gate, re-score — because a verdict that collapses under meaning-preserving editing was never measuring authorship. Its result is a negative one: the loop moves the detectors it optimises against and does not move a detector it has never seen.
a stickier one went 100% → 35% → 0% once the loop used per-sentence feedback. Measured 2026-06-25 on the demo paragraphs, against a third-party site that can change without notice — re-run
--browser zerogpt before relying on it.
Quick start
As a Claude Code skill (zero install):
git clone https://github.com/ssamba1/untell
cp -r untell/untell ~/.claude/skills/untell
# then in Claude Code:
/untell <your text or a file path>
As a Python CLI:
pip install -e ".[full]"
untell-loop "Your AI-sounding paragraph here." # rewrite until it passes
untell-verify --file draft.txt # honest pass/fail per detector
Why it works where blind paraphrasers fail
Drives the max
Optimizes the hardest detector across the whole ensemble, not the average — genuine multi-detector evasion.
Meaning-gated
A 0.76 semantic-similarity bar rejects any rewrite that drifts. It refuses the meaning-mangling other tools ship.
Facts locked
Citations, numbers, quotes, URLs and entities are frozen byte-for-byte. Your APA/IEEE references survive untouched.
Per-sentence
Rewrites only the sentences that read as AI — fewer iterations, less drift, higher pass rate.
The most complete open humanizer
We surveyed ~110 open-source humanizer repos. None combine all four of: a real evasion approach validated against multiple live detectors, a meaning-preservation verifier, an inference-time detector-feedback loop, and a user-installable package. This is the repo that does.
| Capability | untell | lynote (1.4k★) | patina (196★) | StealthHumanizer (58★) |
|---|---|---|---|---|
| Detector-feedback loop | ✅ | ❌ | ◑ | ◑ |
| Real detectors in the loop | ✅ | ❌ | ❌ | ❌ |
| Commercial adapters (6) | ✅ | ❌ | ❌ | ❌ |
| Semantic meaning gate | ✅ | claim | ◑ | ◑ |
| Live detector round-trip | ✅ | ❌ | ❌ | ❌ |
| pip + Claude skill | ✅ | pip | ✅ | web app |
Stars are not capability — the highest-starred repos win on SEO, not architecture. Full evidenced breakdown (and the one place we're honestly not #1): docs/why-best-open-repo.md.
FAQ
Is there a free AI humanizer that actually works?
Yes — the lite tier installs with zero dependencies and the --browser zerogpt path optimizes against a real detector for $0 (live-measured 100%→0%). No tool can honestly promise it passes every commercial detector forever; the ones claiming "99% human" are lying. untell reports the real per-detector score instead.
Will this get past GPTZero / ZeroGPT / Turnitin / Originality.ai?
Unknown, and that is the honest answer — it has never been measured against a commercial checker. What IS measured: the local proxies do not predict them, and the loop's gains do not transfer even to a FREE detector it was not optimised against (4 of 10 still flagged, every seed). Commercial adapters exist as key-gated evaluation targets so you can audit a detector you pay for; wiring one in tells you about that detector, not about a guarantee.
Will it ruin my meaning, citations or numbers?
No. A semantic-similarity gate rejects meaning-drifting rewrites and preserve-lock freezes citations, numbers, quotes, URLs and entities byte-for-byte. Good for academic, legal and ESL writing.
How is it different from the closed-SaaS humanizers?
Different purpose. Those sell a binary "99% human" verdict you cannot inspect. This is a measurement harness: every number is reproducible, the corpus and n are stated, the false-positive rates on human writing are published, and the headline finding is a limitation rather than a claim.
Is this ethical?
AI detectors are noisy proxies — they falsely flag non-native English writers at ~61% in some studies. untell is a research harness and a defense against false positives, not an academic-dishonesty aid.