About this site
What is a glitch token
Glitch tokens are vocabulary entries that behave like faulty input units. Training-data gaps or encoding conflicts can make a model garble, repeat, evade, or invent output.
Every entry lists the raw token string, its publicly documented behaviors, the affected model families, and source links — ready to copy and test yourself.
Behavior types
- Under-trained
- Repetition or embedding tests indicate that the token received too little model training.
- Unspeakable
- The model evades, substitutes, or cannot accurately repeat the token.
- Polysemantic
- The same token is inconsistently interpreted as several unrelated meanings.
- Fingerprint
- The token can help identify the model family or tokenizer behind an API.
Editorial principles
- Only publicly documented cases are included, with source links on every entry.
- Recorded behaviors come from public tests at a point in time; current model versions may no longer show them.
- This site runs no live detection — to test an API, use the tokens from the catalog yourself.
Sources
- SolidGoldMagikarp (plus, prompt generation) — LessWrong
- SolidGoldMagikarp II: technical details — LessWrong
- SolidGoldMagikarp III: Glitch token archaeology — LessWrong
- The ' petertodd' phenomenon — LessWrong
- Magikarp results summary — GitHub
- SmartyHeaderCode: anomalous tokens for GPT3.5 and GPT-4 — LessWrong
- A Search for More “Unspeakable” Glitch Tokens — LessWrong
- GlitchMiner (arXiv:2410.15052)
- Magikarp Llama-2-7b report — GitHub
- Magikarp Phi-3 report — GitHub
- Fishing for Magikarp (arXiv:2405.05417)
- GlitchMiner (arXiv:2410.15052)
- Detecting watered-down LLM APIs with dirty tokens — Zhihu
- Magikarp OLMo 2 report — GitHub
- Magikarp Gemma-7B report — GitHub
- Magikarp Command R+ report — GitHub
- Magikarp Yi-9B report — GitHub
- Magikarp Falcon3-7B report — GitHub
- SGM1 comments: lsusr & mwatkins on ' ertodd' — LessWrong
- SGM2 comments: ' gmaxwell' & the GPT3 basilisk — LessWrong
- SmartyHeaderCode comments: Martin Fell on oralType — LessWrong
- SmartyHeaderCode comments: Martin Fell's finds — LessWrong
- Magikarp Qwen2.5-7B report — GitHub
- Magikarp Llama-3-8B report — GitHub
- Magikarp OLMoE-1B-7B report — GitHub
- xcloche: notes on o200k glitch tokens — note.com
- Haruhiko Okumura's verification notes (2025-09-16)
- GPT-4o's polluted Chinese tokens — MIT Technology Review
- What GPT-oss leaks about OpenAI's training data — LessWrong
- PoC paper (arXiv:2508.17771)
Plain-text versions for LLMs are also available: /llms.txt · /llms-full.txt