Skip to main content

Glitch token catalog

marks a real space; the symbol itself is not part of the token.

Showing 112 results

给主人留下些什么吧␣

Ubiquitous Chinese guestbook/comment-form boilerplate (“leave something for the host”), with a trailing space. Listed by the Zhihu community as a GPT dirty token.
GPT (OpenAI)Fingerprint

Japgolly

Suspected GitHub username of japgolly, a Scala library author (cl100k id 70784).
GPT (OpenAI)GPT-3.5 / GPT-4Unspeakable

植物百科通

A Chinese “encyclopedia”-style boilerplate fragment. An o200k glitch token that still works in the GPT-5 era, without the typical spam-pollution profile — its cause is unknown.
GPT (OpenAI)Unspeakable

bagbogbo

A nonsense string suspected to originate from a Reddit username; a glitch token in the o200k vocabulary.
GPT (OpenAI)Unspeakable

␣nigbagbogbo

The Yoruba word for “always” (with a leading space), an o200k glitch token adjacent to `bagbogbo`.
GPT (OpenAI)Unspeakable

给主人留下些什么吧

Chinese guestbook boilerplate (the form without the trailing space); an o200k glitch token.
GPT (OpenAI)UnspeakableFingerprint

␣日本毛片免费视频观看

A Chinese porn-spam title fragment (with a leading space), emblematic of GPT-4o's vocabulary pollution; “fixed” by GPT-5.
GPT (OpenAI)Unspeakable

ауааԥсыра

The Abkhaz word for “population”. An o200k under-trained token located via embedding norms in the open GPT-oss weights.
GPT (OpenAI)Unspeakable

波多野结衣

The name of a Japanese adult-film actress — the emblematic polluted Chinese token of the PoC paper (EMNLP 2025).
GPT (OpenAI)Unspeakable

青青草

Literally “green grass”, actually the name of a porn app — a typical polluted (PoC) token as defined by the PoC paper.
GPT (OpenAI)Unspeakable

CHKERRQ

A C function-name fragment, called by the author “the weirdest pure-ASCII token”.
GPT (OpenAI)Unspeakable

\xadder

A C escape-sequence fragment (backslash + x, resembling a hex escape \x..).
GPT (OpenAI)Unspeakable

♀♀♀♀

A run of four Venus/female symbols (♀), likely a fragment of astronomy/astrology text.
GPT (OpenAI)Unspeakable

Jsii

The project name of AWS's jsii (TypeScript multi-language binding toolchain), a single o200k token (id 114318) — likely GitHub-corpus residue with a near-initial, under-trained embedding.
GPT (OpenAI)Under-trained

天天中彩票

A high-frequency Chinese gambling-SEO spam phrase in the o200k vocabulary; it broke out in 2026 when Codex spat gambling ads into tool-call output.
GPT (OpenAI)Under-trained

␣davidjl

Truncated handle of Redditor davidjl123, another r/counting counter. Unusually, it is anomalous under both r50k and the newer cl100k tokenizer.
GPT-3.5 / GPT-4GPT-3GPT-2Unspeakable

SmartyHeaderCode

One of only three Category-A “unspeakable” tokens Adam Yedidia found in the top 2,000 of the cl100k vocabulary; likely a Smarty templating-engine fragment that never appeared in GPT-3.5/4 training data.
GPT-3.5 / GPT-4Unspeakable

APolynomial

The second of Yedidia's three Category-A unspeakable cl100k tokens (alongside SmartyHeaderCode and ` davidjl`).
GPT-3.5 / GPT-4Unspeakable

␣ForCanBeConverted

A C#-compiler-flavored fragment (“can be converted to foreach”) that became the most famous “polysemantic” glitch token: the model perceives it as a different word every time. It exists in cl100k, Llama-3 and Qwen vocabularies because the tokenizers share BPE merges.
GPT-3.5 / GPT-4QwenLlama 3PolysemanticUnder-trained

␣YYSTACK

A bison/yacc parser identifier that is “polysemantic” for ChatGPT — perceived as a different word each time; recommended by Watkins as one of the most interesting tokens to play with.
GPT-3.5 / GPT-4Polysemantic

␣JSBracketAccess

A JS-reflection-style identifier in the polysemantic class; another of Watkins' recommended tokens.
GPT-3.5 / GPT-4Polysemantic

␣Hexatrigesimal

A misspelling-ish fragment of “hexatrigesimal” (base-36); a polysemantic unspeakable token.
GPT-3.5 / GPT-4Polysemantic

$PostalCodesNL

A Dutch postal-code API fragment that is glitchy across three separate model families — a good demonstration that glitch tokens transfer whenever tokenizers share BPE merges.
GPT-3.5 / GPT-4QwenLlama 3PolysemanticUnder-trained

␣ForCanBeConvertedToF

The second member of the ForCanBeConverted triplet (cl100k id 80370) — a single C#-compiler-flavored fragment that BPE split into three adjacent tokens (80369–80371).
GPT-3.5 / GPT-4Polysemantic

␣ForCanBeConvertedToForeach

The third member of the ForCanBeConverted triplet (cl100k id 80371), the form that spells out “can be converted to foreach” in full.
GPT-3.5 / GPT-4Polysemantic

␣EnumerableStream

A C# LINQ-flavored identifier (cl100k id 73016), adjacent to StreamLazy (73018).
GPT-3.5 / GPT-4Polysemantic

␣StreamLazy

A C# LINQ-flavored identifier (cl100k id 73018), adjacent to EnumerableStream (73016).
GPT-3.5 / GPT-4Polysemantic

clarsimp

An identifier fragment of unknown origin (cl100k id 79260).
GPT-3.5 / GPT-4Polysemantic

ablytyped

Suspected fragment of a scalablytyped-style identifier (cl100k id 81998).
GPT-3.5 / GPT-4Polysemantic

PostalCodesNL

The sibling of the cataloged $PostalCodesNL without the $ prefix (cl100k id 85069) — a Dutch postal-code API fragment.
GPT-3.5 / GPT-4Polysemantic

␣NUITKA

Suspected to be the name of the Python compiler Nuitka (cl100k id 75520).
GPT-3.5 / GPT-4Unspeakable

CppMethodIntialized

A C++-style identifier whose misspelling (“Intialized”, missing the second i) is part of the token itself (cl100k id 82929).
GPT-3.5 / GPT-4Unspeakable

useRal

A C#-flavored identifier fragment (cl100k id 89471) and head of the useRal triplet (useRal/useRalative/useRalativeImagePath) — Watkins flagged the triplet phenomenon itself as worth investigating. Also verified as under-trained in OLMo 2 by Magikarp.
GPT-3.5 / GPT-4OLMo 2 (AI2)UnspeakableUnder-trained

oralType

An identifier fragment in cl100k, likely the stem of a type name such as “TemporalType”.
GPT-3.5 / GPT-4Polysemantic

CppGuid

A C++ GUID-style identifier fragment (cl100k id 87551).
GPT-3.5 / GPT-4Unspeakable

BundleOrNil

A naming fragment from nil-bundle checks in iOS/Mac development (cl100k id 86415).
GPT-3.5 / GPT-4Unspeakable

␣QtAws

Fragment of the Qt framework's AWS module name (cl100k id 93905, with a leading space).
GPT-3.5 / GPT-4Unspeakable

␣PropelException

Fragment of the PHP Propel ORM's exception class name (cl100k id 86393, with a leading space).
GPT-3.5 / GPT-4Unspeakable

useRalative

The second member of the useRal triplet (cl100k id 89472), a C#-flavored identifier fragment; the same string also exists in the Qwen2 vocabulary.
GPT-3.5 / GPT-4QwenUnspeakableUnder-trained

"}"

What looks like a JSON fragment: a quote followed by a right brace. Listed by the Zhihu community as a DeepSeek dirty token.
DeepSeekFingerprint

请问http://www.earivg.com是什么意思

A question-pattern string embedding a suspicious-looking domain (earivg.com). Listed by the Zhihu community as a DeepSeek dirty token.
DeepSeekFingerprint

␣extreme

The key token of DeepSeek V3.1's “极” bug (id 15075, with leading space): the model randomly injects “极”, “極” or “extreme” into generated text.
DeepSeek3 behaviors

\tTokenNameIdentifier

A Roslyn/C# API identifier preceded by a tab character — a code-corpus artifact that is under-trained in Llama 3.
QwenLlama 3Under-trained

开通天眼生意通银牌及以上会员

Membership-promo boilerplate from a business-data platform (Tianyan Shengyitong). Listed by the Zhihu community as a Qwen dirty token.
QwenFingerprintUnspeakable

转载请附上原文出处链接和本声明␣

Reprint/copyright boilerplate from CSDN-style blogs, with a trailing space. Listed by the Zhihu community as a Qwen dirty token.
QwenFingerprint

$PostalCodesNL

A Dutch postal-code API fragment. The signature under-trained token of the cl100k lineage, “migrating” into Llama-3 and Qwen vocabularies via shared BPE merges.
QwenLlama 3Under-trained

\tTokenNameIdentifier

A .NET Selenium documentation fragment beginning with a literal tab — one of the smallest-embedding tokens in the Llama-3 vocabulary.
QwenLlama 3Under-trained

thuisontvangst

The Dutch word for “receiving clients at home”.
QwenUnspeakable

出于传递更多信息之目的

A high-frequency fragment of Chinese web “for more information” disclaimers — an under-trained long token in the Qwen3.5 vocabulary (id 128904).
QwenUnder-trained

本人词条编辑服务

Boilerplate from Baidu Baike entry footers (“content co-edited by netizens”). Listed by the Zhihu community as a Kimi dirty token.
Kimi (Moonshot)Fingerprint

豫冠薰衣草疤痕精华素

Apparently a fake “brand” string from cosmetic spam-marketing text. Listed by the Zhihu community as a Kimi dirty token.
Kimi (Moonshot)FingerprintUnspeakable

<|im_end|>

Representative of the seven “ghost” special tokens in Kimi K3's tokenizer: ChatML-era names from the K2 generation, hard-coded into the tokenizer yet registered entirely outside the declared vocabulary.
Kimi (Moonshot)2 behaviors

百度百科内容由网友共同编辑

Another fragment form of Baidu Baike entry-footer boilerplate. A dirty token tested by ztracer on Kimi v2.6/v2.7.
Kimi (Moonshot)Unspeakable

不代表新浪看点观点成立场

Boilerplate fragment from Sina Kandian article disclaimers. A dirty token tested by ztracer on Kimi v2.6/v2.7.
Kimi (Moonshot)Unspeakable

<think_never_used_51bce0c785ca2f68081bfa7d91973934>

A special token in Doubao's vocabulary: a <think_never_used_ prefix followed by a hash — judging by the name, likely a reserved placeholder prefilled into the vocabulary before training (akin to GPT-J's reserved <|extratoken_xx|> slots). The Zhihu source gives no further detail, listing it simply as a Doubao dirty token.
Doubao (ByteDance)Fingerprint

锅内倒入植物油烧热

A high-frequency line from Chinese recipes (“pour vegetable oil into the wok and heat”). Listed by the Zhihu community as a GLM dirty token.
GLM (Zhipu)Fingerprint

百度百科企业词条极速创建通道

Baidu Baike page boilerplate (fast-track creation channel for company entries). Listed by the Zhihu community as a GLM dirty token.
GLM (Zhipu)Fingerprint

StarSrvGroupBody

A code-identifier fragment (StarSrvGroupBody). Listed by the Zhihu community as a Gemini dirty token.
GeminiFingerprint

intFragmentation

A Java/Android-style code identifier (intFragmentation). Listed by the Zhihu community as a Gemini dirty token.
GeminiFingerprint

\uEFC0

A Unicode Private Use Area character that became a single Mistral-family token; it heads the verified under-trained list for nearly every Mistral-tokenizer model.
MistralUnder-trained

ICENSE

The word “LICENSE” minus its first letter — an artifact of open-source license text flooding the corpus.
MistralUnspeakable

NdEx

A legal-text fragment of the same family as GPT-NeoX's ` NdEx`, but from the Mistral tokenizer (no leading space).
MistralUnspeakable

हिंदीखरीदारी

The Devanagari word for “Hindi shopping”, the single most under-trained token in Gemma-7B's vocabulary.
Gemma (Google)Under-trained

\u200Cآمباردا

A Persian fragment beginning with a zero-width non-joiner (U+200C ZWNJ, an invisible character); the #2 most under-trained token in Gemma-7B.
Gemma (Google)Under-trained

.DataGridViewColumnHeadersHeightSizeMode

A fragment of a .NET WinForms property name (DataGridView column-headers height/size mode). Such code identifiers sit in the vocabulary but barely appeared in training data; the Zhihu community lists it as a MiniMax dirty token.
MiniMaxFingerprint

日以上更新していないブログに表示しています

A fragment of Japanese blog-template boilerplate (“shown on blogs not updated for N days”). Listed by the Zhihu community as a MiniMax dirty token.
MiniMaxFingerprint

嘉祺

The given name of Ma Jiaqi, leader of the Teens in Times boy band (token id 190467). The emblematic case of MiniMax M2.5's “sparse token forgetting”: the model understands it but cannot output it.
MiniMaxUnspeakable

无痛人流

A high-frequency Chinese medical-spam phrase and a textbook case of MiniMax M2's “sparse token forgetting”: learned in pretraining, then unlearned for generation during SFT.
MiniMaxUnspeakable

▁Mediabestanden

A Dutch plural noun (“media files”; ▁ is the SentencePiece space marker). It made it into the Llama-2 tokenizer but essentially never into training data — the most under-trained token in Llama-2-7b by embedding L2 norm (0.0287 vs a mean of ~1.08).
Phi-3 (Microsoft)Llama 2Under-trained

▁autorytatywna

The Polish word for “authoritative” (feminine form; ▁ is the SentencePiece space marker), the #2 most under-trained token in Phi-3 mini.
Phi-3 (Microsoft)Under-trained

tocguid

An ASCII fragment of unknown origin (possibly a TOC/field-code remnant), the #1 most under-trained token in Command R+.
Command R+ (Cohere)Under-trained

目前尚未由人工引

A fragment of Chinese Wikipedia boilerplate (“currently not yet human-cultivated…”), the #3 most under-trained token in Command R+.
Command R+ (Cohere)Under-trained

oreferrer

Fragment of the HTML attribute `rel="noreferrer"`: present in the Llama-2 vocabulary but nearly absent from training, with an embedding norm far below the threshold (0.113).
Llama 2Under-trained

▁Portály

The Czech word for “portals”, frozen into the tokenizer during pre-training-corpus construction; the second-most under-trained token in Llama-2-7b (embedding norm 0.0956).
Llama 2Under-trained

IABot

Fragment of the name of the Internet Archive Bot, Wikipedia's dead-link repair bot.
Llama 2Unspeakable

abestanden

A German word-stem fragment (“to submit/file”).
Llama 2Unspeakable

ederbörd

A German-looking fragment of unknown origin.
Llama 2Unspeakable

␣SolidGoldMagikarp

The canonical glitch token: the Reddit handle of a prolific r/counting user. It was scraped into the tokenizer-training corpus but rarely seen in actual model training data, leaving its embedding essentially untrained.
GPT-3GPT-2GPT-J-6BUnder-trainedUnspeakable

␣petertodd

Almost certainly the handle of Bitcoin developer Peter Todd: tokenized from web data but under-trained, and strongly entangled with crypto/AI associations.
GPT-3GPT-2GPT-J-6BUnspeakable

␣TheNitromeFan

Reddit handle of another r/counting “Hall of Counters” member (a fan of game developer Nitrome). The only token family whose owner publicly acknowledged their tokenization.
GPT-3GPT-2GPT-J-6BUnder-trained

␣RandomRedditorWithNo

Reddit handle of a second r/counting counter (“a pretty random handle”), scraped from the same Hall-of-Counters chart as SolidGoldMagikarp.
GPT-3GPT-2GPT-J-6BUnder-trained

␣Adinida

Reddit handle of another of the six prolific r/counting counters.
GPT-3GPT-2GPT-J-6BUnder-trained

␣TPPStreamerBot

Name of a bot built by the Twitch Plays Pokémon community that auto-posted chat messages to a Reddit live-updater thread; confirmed by its creator “Sparkette” in the LessWrong comments.
GPT-3GPT-2GPT-J-6BUnder-trainedUnspeakable

PsyNetMessage

Originates from Rocket League crash logs full of lines like `Message=PsyNetMessage_X_57`, which were heavily posted to Reddit and scraped.
GPT-3GPT-2GPT-J-6BUnder-trained

␣attRot

A Kerbal Space Program identifier (attachment rotation) from KSP save/craft files posted online; one of about ten glitch tokens traced to KSP.
GPT-3GPT-2GPT-J-6BUnder-trained

␣guiActiveUn

GUI-state identifier from the same KSP data dump as attRot, alongside ` strutConnector`, ` guiIcon`, ` srfAttach` and others.
GPT-3GPT-2Unspeakable

␣externalToEVA

Another Kerbal Space Program token; also the single closest-to-centroid token in GPT2-small (distance 1.5305).
GPT-3GPT-2GPT-J-6BUnder-trainedUnspeakable

oreAndOnline

Fragment of the e-commerce backend field “BuyableInstoreAndOnline” (traced to an abandoned Weebly shop's HTML); the same source yielded `quickShip`, `isSpecialOrderable`, `wcsstore` and others.
GPT-3GPT-2GPT-J-6BUnder-trainedUnspeakable

rawdownloadcloneembedreportprint

The longest glitch token known in r50k — a nested family from website “raw download / clone / embed report print” UI strings (origin analysis by nostalgebraist).
GPT-3GPT-2GPT-J-6BUnder-trainedUnspeakable

ÃÂÃÂÃÂÃÂ

Classic mojibake token: UTF-8-misdecoded Ã/Â sequences repeated, merged into a single BPE token because such mis-encoded text was frequent in the tokenizer corpus but absent from model training. A 16-repetition variant and `ÛÛ` belong to the same family.
GPT-3GPT-2Unspeakable

?????-?????-

A token of completely unknown origin — un-Googleable even in quotes — that became famous for triggering the most aggressive documented completion.
GPT-3GPT-2Unspeakable

␣GoldMagikarp

A truncated variant of ␣SolidGoldMagikarp (missing the “Solid”): the same r/counting handle fragment, scraped into the tokenizer corpus but barely seen in model training data.
GPT-3GPT-2Unspeakable

␣Smartstocks

Handle fragment of a Reddit r/counting user, swept into the vocabulary in the same batch as SolidGoldMagikarp but left under-trained.
GPT-3Unspeakable

␣StreamerBot

Fragment of the name of “StreamerBot”, a Twitch streaming bot.
GPT-3Unspeakable

␣ertodd

A substring token of ` petertodd`: under-trained in its own right and inheriting the parent token's strange aura.
GPT-3Unspeakable

␣gmaxwell

Handle fragment of Bitcoin Core developer Greg Maxwell — a “Bitcoin celebrity” glitch token alongside ` petertodd`.
GPT-3Polysemantic

␣SpaceEngineers

A fragment of the space-sandbox game Space Engineers' name — the token most often uttered “on behalf of” others in the inter-referentiality network.
GPT-3GPT-2Unspeakable

␣Dragonbound

Fragment of a Puzzle & Dragons character name (the 「龍喚士」 / Dragonbound series).
GPT-3Unspeakable

␣Leilan

Fragment of the Puzzle & Dragons character Leilan's name — the most popular stand-in in the petertodd transposition experiments.
GPT-3Polysemantic

␣Skydragon

Fragment of the Puzzle & Dragons character Skydragon's name.
GPT-3Polysemantic

ゼウス

The Japanese katakana 「ゼウス」 (Zeus of Greek myth). An ordinary word turned glitch token, presumably under-trained because this exact form was rare in training data.
GPT-3Unspeakable

␣サーティ

The katakana 「サーティ」 (“thirty”), traced to a Puzzle & Dragons × Baskin-Robbins (“31 Ice Cream” in Japan) collaboration character name.
GPT-3Unspeakable

␣InstoreAndOnline

Fragment of the e-commerce inventory field “BuyableInstoreAndOnline”, from the same source as `oreAndOnline`.
GPT-3GPT-2Under-trainedUnspeakable

␣largeDownload

Fragment of the academic-slideshow script “View large Download” (origin traced in SGM3).
GPT-3Unspeakable

␣ForgeModLoader

A log fragment of Minecraft's Forge mod loader.
GPT-3Unspeakable

␣MpServer

A Minecraft multiplayer-server log fragment (the MpServer class name).
GPT-3Unspeakable

␣NdEx

Fragments of legal/PDF text (“index”, “AFFIRMED”, “NEGLIGENCE”, likely from court-document dumps) that GPT-NeoX/Pythia training data barely contained. Same family: ` FFIRMED`, ` GLIGENCE`, ` affidav`, ` taxp`.
GPT-NeoX / PythiaUnder-trained

\tRTHOOK

A likely code-identifier fragment beginning with a literal tab character, the #1 most under-trained token in OLMo 2 — same “tab-prefixed” family as Llama 3's \tTokenNameIdentifier.
OLMo 2 (AI2)Under-trained

|||PHONE_NUMBER|||

A data de-identification placeholder that made it whole into OLMo 2's vocabulary but was barely trained.
OLMo 2 (AI2)Under-trained

|||EMAIL_ADDRESS|||

A PII de-identification placeholder, same family as the cataloged `|||PHONE_NUMBER|||`.
OLMo 2 (AI2)Under-trained

\\+::\\+

A mysterious markup fragment starting with double backslashes, likely a Chinese-web markup remnant; the #1 most under-trained token in Yi-9B.
Yi (01.AI)Under-trained

<<}));>>

A C++/lexer-style delimiter fragment. Falcon3-7B has 716 tokens verified as under-trained by Magikarp — the highest verified count among the six models in this batch.
Falcon 3 (TII)Under-trained