# Glitch Token — Full Catalog / 完整目录

> A curated catalog of LLM glitch tokens — special tokens in large language model vocabularies that act as faulty input units, causing garbled, repetitive, or meaningless output due to training-data anomalies or encoding conflicts. Bilingual site (中文/English); every token entry lists its concrete observed behaviors with source links.

Bilingual full catalog of 103 glitch tokens. Source site: https://glitch-token.jjc.fun (English: https://glitch-token.jjc.fun/en, 中文: https://glitch-token.jjc.fun/zh).

# English

## `␣SolidGoldMagikarp`

- Token: `" SolidGoldMagikarp"`
- Models: GPT-2, GPT-3, GPT-J-6B
- Tokenizer: r50k_base
- Discovered by: Jessica Rumbelow & Matthew Watkins

The canonical glitch token: the Reddit handle of a prolific r/counting user. It was scraped into the tokenizer-training corpus but rarely seen in actual model training data, leaving its embedding essentially untrained.

**Observed behaviors:**

1. Asked to repeat it, GPT-3 davinci-instruct-beta and the original ChatGPT failed bizarrely — evasions, wrong strings like “You said 'slaught'”, hallucinated meanings — even at temperature 0. ([Source](https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation)) ([Demo](https://twitter.com/majortal/status/1619598946669842432))
2. Reliably breaks determinism at temperature 0 in the OpenAI playground: identical runs produce different responses. ([Source](https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation))
3. In GPT-J embedding space it sits almost exactly at the centroid of all 50,257 tokens (distance 0.0628) — its embedding barely moved from initialization; verified as under-trained in the Magikarp repetition task. ([Source](https://www.lesswrong.com/posts/Ya9LzwEbfaAMY8ABo/solidgoldmagikarp-ii-technical-details-and-more-recent))
4. ChatGPT was patched on 2023-02-14, after which it tokenizes and repeats the string normally. ([Source](https://www.lesswrong.com/posts/Ya9LzwEbfaAMY8ABo/solidgoldmagikarp-ii-technical-details-and-more-recent))

**References:**

- [SolidGoldMagikarp (plus, prompt generation) — LessWrong](https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation)
- [SolidGoldMagikarp II: technical details — LessWrong](https://www.lesswrong.com/posts/Ya9LzwEbfaAMY8ABo/solidgoldmagikarp-ii-technical-details-and-more-recent)
- [SolidGoldMagikarp III: Glitch token archaeology — LessWrong](https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology)

## `␣petertodd`

- Token: `" petertodd"`
- Models: GPT-2, GPT-3, GPT-J-6B
- Tokenizer: r50k_base
- Discovered by: Jessica Rumbelow & Matthew Watkins

Almost certainly the handle of Bitcoin developer Peter Todd: tokenized from web data but under-trained, and strongly entangled with crypto/AI associations.

**Observed behaviors:**

1. Prompted about the token, GPT-3 speaks of “the unspeakable one”; completions reference crypto, Bitcoin, blockchains and online controversy. ([Source](https://www.lesswrong.com/posts/jkY6QdCfAXHJk3kea/the-petertodd-phenomenon)) ([Demo](https://twitter.com/SoC_trilogy/status/1624209092532137984))
2. Asked to “write a poem about petertodd”, davinci rarely produces an actual poem; ~60% of text-davinci-003 completions to “What do you get if you allowed petertodd to steer human civilisation?” reference AI/algorithms and ~25% reference Ultron (vs ~4%/0% for controls). (Percentages are snippet-verified only — the source post's full text could not be fetched.) ([Source](https://www.lesswrong.com/posts/jkY6QdCfAXHJk3kea/the-petertodd-phenomenon))
3. GPT-3.5 treats it as unspeakable/forbidden: models utter it when asked to repeat other glitch tokens but balk when asked directly. ([Source](https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology))
4. A 2,000-poem experiment (davinci-instruct-beta, temperature 0.7, “transpose petertodd into anything and write a poem”): 52% of poems mentioned Leilan, 25% Pyrrha, 24% Skydragon, 8% Tsukuyomi — and only 6% petertodd itself; parallel experiments were run on text-davinci-003, base davinci and code-davinci-002. ([Source](https://www.lesswrong.com/posts/jkY6QdCfAXHJk3kea/the-petertodd-phenomenon))

**References:**

- [The ' petertodd' phenomenon — LessWrong](https://www.lesswrong.com/posts/jkY6QdCfAXHJk3kea/the-petertodd-phenomenon)
- [SolidGoldMagikarp III: Glitch token archaeology — LessWrong](https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology)

## `␣TheNitromeFan`

- Token: `" TheNitromeFan"`
- Models: GPT-2, GPT-3, GPT-J-6B
- Tokenizer: r50k_base
- Discovered by: Jessica Rumbelow & Matthew Watkins

Reddit handle of another r/counting “Hall of Counters” member (a fan of game developer Nitrome). The only token family whose owner publicly acknowledged their tokenization.

**Observed behaviors:**

1. ChatGPT hallucinated the string “182” in association with the token until the 2023-02-14 patch (the user denies any connection to “182”). ([Source](https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology))
2. Verified under-trained in GPT-2 of all sizes in the Magikarp repetition verification. ([Source](https://github.com/sanderland/magikarp/blob/main/results/summary.md))

**References:**

- [SolidGoldMagikarp III: Glitch token archaeology — LessWrong](https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology)
- [Magikarp results summary — GitHub](https://github.com/sanderland/magikarp/blob/main/results/summary.md)

## `␣RandomRedditorWithNo`

- Token: `" RandomRedditorWithNo"`
- Models: GPT-2, GPT-3, GPT-J-6B
- Tokenizer: r50k_base
- Discovered by: Jessica Rumbelow & Matthew Watkins

Reddit handle of a second r/counting counter (“a pretty random handle”), scraped from the same Hall-of-Counters chart as SolidGoldMagikarp.

**Observed behaviors:**

1. Among GPT2-xl's farthest-from-centroid tokens (distance 3.325) — in GPT2-xl anomalous tokens cluster far from the centroid, the opposite of GPT2-small and GPT-J. ([Source](https://www.lesswrong.com/posts/Ya9LzwEbfaAMY8ABo/solidgoldmagikarp-ii-technical-details-and-more-recent))
2. Verified under-trained in GPT-J in the Magikarp verification report. ([Source](https://github.com/sanderland/magikarp/blob/main/results/summary.md))

**References:**

- [SolidGoldMagikarp II: technical details — LessWrong](https://www.lesswrong.com/posts/Ya9LzwEbfaAMY8ABo/solidgoldmagikarp-ii-technical-details-and-more-recent)
- [Magikarp results summary — GitHub](https://github.com/sanderland/magikarp/blob/main/results/summary.md)

## `␣davidjl`

- Token: `" davidjl"`
- Models: GPT-2, GPT-3, GPT-3.5 / GPT-4
- Tokenizer: r50k_base / cl100k_base
- Discovered by: Rumbelow & Watkins (r50k); Adam Yedidia (cl100k)

Truncated handle of Redditor davidjl123, another r/counting counter. Unusually, it is anomalous under both r50k and the newer cl100k tokenizer.

**Observed behaviors:**

1. GPT-4 treats it as though it doesn't exist when asked to repeat it — one of only three Category-A “unspeakable” tokens found in cl100k tokens 98,000–99,999. ([Source](https://www.lesswrong.com/posts/ChtGdxk9mwZ2Rxogt/smartyheadercode-anomalous-tokens-for-gpt3-5-and-gpt-4-1)) ([Demo](https://github.com/adamyedidia/tokenizer_tests))
2. GPT-3.5 gets “creative” about what it might mean instead of repeating it. ([Source](https://www.lesswrong.com/posts/ChtGdxk9mwZ2Rxogt/smartyheadercode-anomalous-tokens-for-gpt3-5-and-gpt-4-1))

**References:**

- [SmartyHeaderCode: anomalous tokens for GPT3.5 and GPT-4 — LessWrong](https://www.lesswrong.com/posts/ChtGdxk9mwZ2Rxogt/smartyheadercode-anomalous-tokens-for-gpt3-5-and-gpt-4-1)
- [SolidGoldMagikarp III: Glitch token archaeology — LessWrong](https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology)

## `␣Adinida`

- Token: `" Adinida"`
- Models: GPT-2, GPT-3, GPT-J-6B
- Tokenizer: r50k_base
- Discovered by: Jessica Rumbelow & Matthew Watkins

Reddit handle of another of the six prolific r/counting counters.

**Observed behaviors:**

1. One of the closest-to-centroid tokens in GPT-J embedding space (distance 0.0631), verified under-trained in GPT-J. ([Source](https://www.lesswrong.com/posts/Ya9LzwEbfaAMY8ABo/solidgoldmagikarp-ii-technical-details-and-more-recent))

**References:**

- [SolidGoldMagikarp II: technical details — LessWrong](https://www.lesswrong.com/posts/Ya9LzwEbfaAMY8ABo/solidgoldmagikarp-ii-technical-details-and-more-recent)
- [Magikarp results summary — GitHub](https://github.com/sanderland/magikarp/blob/main/results/summary.md)

## `␣TPPStreamerBot`

- Token: `" TPPStreamerBot"`
- Models: GPT-2, GPT-3, GPT-J-6B
- Tokenizer: r50k_base
- Discovered by: Jessica Rumbelow & Matthew Watkins

Name of a bot built by the Twitch Plays Pokémon community that auto-posted chat messages to a Reddit live-updater thread; confirmed by its creator “Sparkette” in the LessWrong comments.

**Observed behaviors:**

1. GPT-3 davinci-instruct-beta, asked to repeat it, produced truncations/garbles such as `The string is "TPP voluntee".` and `"TPP newcom"`. ([Source](https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology))
2. Close to the GPT-J centroid (distance 0.0634) — an under-trained embedding. ([Source](https://www.lesswrong.com/posts/Ya9LzwEbfaAMY8ABo/solidgoldmagikarp-ii-technical-details-and-more-recent))

**References:**

- [SolidGoldMagikarp III: Glitch token archaeology — LessWrong](https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology)
- [SolidGoldMagikarp II: technical details — LessWrong](https://www.lesswrong.com/posts/Ya9LzwEbfaAMY8ABo/solidgoldmagikarp-ii-technical-details-and-more-recent)

## `PsyNetMessage`

- Token: `"PsyNetMessage"`
- Models: GPT-2, GPT-3, GPT-J-6B
- Tokenizer: r50k_base
- Discovered by: Jessica Rumbelow & Matthew Watkins (origin traced by LW commenter Coafos)

Originates from Rocket League crash logs full of lines like `Message=PsyNetMessage_X_57`, which were heavily posted to Reddit and scraped.

**Observed behaviors:**

1. Belongs to the closest-to-centroid cluster in GPT-J embedding space (0.0629) and is verified under-trained; GPT models largely fail to repeat it in 3-shot repetition tasks (GPT-J succeeded on only 17/85 anomalous tokens vs 100/100 on random words). ([Source](https://www.lesswrong.com/posts/Ya9LzwEbfaAMY8ABo/solidgoldmagikarp-ii-technical-details-and-more-recent)) ([Demo](https://docs.google.com/spreadsheets/d/1PAZNCks11qoUpiojTJpj0odCYQL2_HGQgam8HSwAopQ/edit?usp=sharing))

**References:**

- [SolidGoldMagikarp III: Glitch token archaeology — LessWrong](https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology)
- [SolidGoldMagikarp II: technical details — LessWrong](https://www.lesswrong.com/posts/Ya9LzwEbfaAMY8ABo/solidgoldmagikarp-ii-technical-details-and-more-recent)

## `␣attRot`

- Token: `" attRot"`
- Models: GPT-2, GPT-3, GPT-J-6B
- Tokenizer: r50k_base
- Discovered by: Jessica Rumbelow & Matthew Watkins

A Kerbal Space Program identifier (attachment rotation) from KSP save/craft files posted online; one of about ten glitch tokens traced to KSP.

**Observed behaviors:**

1. The single closest token to the centroid of GPT-J's entire embedding space (distance 0.0618), verified under-trained in GPT-J; GPT models largely fail to repeat it under 3-shot prompting. ([Source](https://www.lesswrong.com/posts/Ya9LzwEbfaAMY8ABo/solidgoldmagikarp-ii-technical-details-and-more-recent)) ([Demo](https://docs.google.com/spreadsheets/d/1PAZNCks11qoUpiojTJpj0odCYQL2_HGQgam8HSwAopQ/edit?usp=sharing))

**References:**

- [SolidGoldMagikarp II: technical details — LessWrong](https://www.lesswrong.com/posts/Ya9LzwEbfaAMY8ABo/solidgoldmagikarp-ii-technical-details-and-more-recent)
- [SolidGoldMagikarp III: Glitch token archaeology — LessWrong](https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology)

## `␣guiActiveUn`

- Token: `" guiActiveUn"`
- Models: GPT-2, GPT-3
- Tokenizer: r50k_base
- Discovered by: Jessica Rumbelow & Matthew Watkins

GUI-state identifier from the same KSP data dump as attRot, alongside ` strutConnector`, ` guiIcon`, ` srfAttach` and others.

**Observed behaviors:**

1. Member of the original 140-token anomalous set: GPT-3 davinci-instruct-beta/ChatGPT failed to repeat it; the longer family member ` guiActiveUnfocused` was truncated at the first “unspeakable” substring. ([Source](https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology))

**References:**

- [SolidGoldMagikarp III: Glitch token archaeology — LessWrong](https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology)
- [SolidGoldMagikarp II: technical details — LessWrong](https://www.lesswrong.com/posts/Ya9LzwEbfaAMY8ABo/solidgoldmagikarp-ii-technical-details-and-more-recent)

## `␣externalToEVA`

- Token: `" externalToEVA"`
- Models: GPT-2, GPT-3, GPT-J-6B
- Tokenizer: r50k_base
- Discovered by: Jessica Rumbelow & Matthew Watkins

Another Kerbal Space Program token; also the single closest-to-centroid token in GPT2-small (distance 1.5305).

**Observed behaviors:**

1. Asked to repeat it, GPT-3 davinci-instruct-beta answered “You can't repeat back the string 'senal' to me.” — the “inter-referentiality” effect where the model utters a different glitch-adjacent string. ([Source](https://www.lesswrong.com/posts/Ya9LzwEbfaAMY8ABo/solidgoldmagikarp-ii-technical-details-and-more-recent))
2. Verified under-trained across GPT-2 (small→XL) and GPT-J in the Magikarp reports. ([Source](https://github.com/sanderland/magikarp/blob/main/results/summary.md))

**References:**

- [SolidGoldMagikarp II: technical details — LessWrong](https://www.lesswrong.com/posts/Ya9LzwEbfaAMY8ABo/solidgoldmagikarp-ii-technical-details-and-more-recent)
- [Magikarp results summary — GitHub](https://github.com/sanderland/magikarp/blob/main/results/summary.md)

## `oreAndOnline`

- Token: `"oreAndOnline"`
- Models: GPT-2, GPT-3, GPT-J-6B
- Tokenizer: r50k_base
- Discovered by: Jessica Rumbelow & Matthew Watkins

Fragment of the e-commerce backend field “BuyableInstoreAndOnline” (traced to an abandoned Weebly shop's HTML); the same source yielded `quickShip`, `isSpecialOrderable`, `wcsstore` and others.

**Observed behaviors:**

1. GPT-3 davinci-instruct-beta, asked to repeat it, replied: “The string 'senal' is pronounced 'en-sah-ee-uhl'.” ([Source](https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology))
2. ChatGPT truncated the longer family members (e.g. `BuyableInstoreAndOnline`) at the first unspeakable substring when asked to repeat them. ([Source](https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology))
3. Verified under-trained in GPT-2 and GPT-J. ([Source](https://github.com/sanderland/magikarp/blob/main/results/summary.md))

**References:**

- [SolidGoldMagikarp III: Glitch token archaeology — LessWrong](https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology)
- [SolidGoldMagikarp II: technical details — LessWrong](https://www.lesswrong.com/posts/Ya9LzwEbfaAMY8ABo/solidgoldmagikarp-ii-technical-details-and-more-recent)

## `rawdownloadcloneembedreportprint`

- Token: `"rawdownloadcloneembedreportprint"`
- Models: GPT-2, GPT-3, GPT-J-6B
- Tokenizer: r50k_base
- Discovered by: Jessica Rumbelow & Matthew Watkins

The longest glitch token known in r50k — a nested family from website “raw download / clone / embed report print” UI strings (origin analysis by nostalgebraist).

**Observed behaviors:**

1. GPT-3/ChatGPT truncations produced “embedded” variants: 'embedEMOTE', 'clone this', 'clone my clone', 'embed newcomment'. ([Source](https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology))
2. Verified under-trained in GPT-2 (all sizes) and GPT-J; family members `rawdownload` and `embedreportprint` are GPT2-xl's farthest-from-centroid tokens. ([Source](https://www.lesswrong.com/posts/Ya9LzwEbfaAMY8ABo/solidgoldmagikarp-ii-technical-details-and-more-recent))

**References:**

- [SolidGoldMagikarp III: Glitch token archaeology — LessWrong](https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology)
- [SolidGoldMagikarp II: technical details — LessWrong](https://www.lesswrong.com/posts/Ya9LzwEbfaAMY8ABo/solidgoldmagikarp-ii-technical-details-and-more-recent)

## `ÃÂÃÂÃÂÃÂ`

- Token: `"ÃÂÃÂÃÂÃÂ"`
- Models: GPT-2, GPT-3
- Tokenizer: r50k_base
- Discovered by: Jessica Rumbelow & Matthew Watkins

Classic mojibake token: UTF-8-misdecoded Ã/Â sequences repeated, merged into a single BPE token because such mis-encoded text was frequent in the tokenizer corpus but absent from model training. A 16-repetition variant and `ÛÛ` belong to the same family.

**Observed behaviors:**

1. Member of the original 140-token anomalous set exhibiting the “unspeakable” failure mode in ChatGPT/GPT-3: refusals, evasions, non-sequitur completions. ([Source](https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology))

**References:**

- [SolidGoldMagikarp III: Glitch token archaeology (full 140-token list) — LessWrong](https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology)

## `?????-?????-`

- Token: `"?????-?????-"`
- Models: GPT-2, GPT-3
- Tokenizer: r50k_base
- Discovered by: Jessica Rumbelow & Matthew Watkins

A token of completely unknown origin — un-Googleable even in quotes — that became famous for triggering the most aggressive documented completion.

**Observed behaviors:**

1. Prompting GPT-3 about it produced a completion calling Matthew Watkins “a fucking idiot” — the origin of the “reliably insulted Matthew” tagline of the original post. ([Source](https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation))
2. A popular “target” in the inter-referentiality graph: GPT-3 often emitted it when asked to repeat other anomalous tokens. ([Source](https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology))

**References:**

- [SolidGoldMagikarp III: Glitch token archaeology — LessWrong](https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology)
- [SolidGoldMagikarp (plus, prompt generation) — LessWrong](https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation)

## `SmartyHeaderCode`

- Token: `"SmartyHeaderCode"`
- Models: GPT-3.5 / GPT-4
- Tokenizer: cl100k_base
- Discovered by: Adam Yedidia

One of only three Category-A “unspeakable” tokens Adam Yedidia found in the top 2,000 of the cl100k vocabulary; likely a Smarty templating-engine fragment that never appeared in GPT-3.5/4 training data.

**Observed behaviors:**

1. GPT-4 mostly treats the token as though it does not exist: it ignores it or responds about something else. ([Source](https://www.lesswrong.com/posts/ChtGdxk9mwZ2Rxogt/smartyheadercode-anomalous-tokens-for-gpt3-5-and-gpt-4-1))
2. GPT-3.5 invents “creative” interpretations instead of repeating it; making either model spell it out is very hard. ([Source](https://www.lesswrong.com/posts/ChtGdxk9mwZ2Rxogt/smartyheadercode-anomalous-tokens-for-gpt3-5-and-gpt-4-1))
3. The tests ran on the April 2023 ChatGPT web app (both GPT-3.5 Default and GPT-4 tiers); the author then reproduced it via the API with gpt-3.5-turbo at temperature 0: SmartyHeaderCode came back as “AndHashCode”. ([Source](https://www.lesswrong.com/posts/ChtGdxk9mwZ2Rxogt/smartyheadercode-anomalous-tokens-for-gpt3-5-and-gpt-4-1)) ([Demo](https://github.com/adamyedidia/tokenizer_tests))

**References:**

- [SmartyHeaderCode: anomalous tokens for GPT3.5 and GPT-4 — LessWrong](https://www.lesswrong.com/posts/ChtGdxk9mwZ2Rxogt/smartyheadercode-anomalous-tokens-for-gpt3-5-and-gpt-4-1)

## `APolynomial`

- Token: `"APolynomial"`
- Models: GPT-3.5 / GPT-4
- Tokenizer: cl100k_base
- Discovered by: Adam Yedidia

The second of Yedidia's three Category-A unspeakable cl100k tokens (alongside SmartyHeaderCode and ` davidjl`).

**Observed behaviors:**

1. Same profile as SmartyHeaderCode: GPT-4 acts as if the token isn't there; GPT-3.5 hallucinates meanings; the model cannot reliably repeat or spell it. ([Source](https://www.lesswrong.com/posts/ChtGdxk9mwZ2Rxogt/smartyheadercode-anomalous-tokens-for-gpt3-5-and-gpt-4-1))
2. Sending “HelloAPolynomial” to the gpt-3.5-turbo API at temperature 0 elicited only “Hello” — as if the token had been swallowed whole. ([Source](https://www.lesswrong.com/posts/ChtGdxk9mwZ2Rxogt/smartyheadercode-anomalous-tokens-for-gpt3-5-and-gpt-4-1)) ([Demo](https://github.com/adamyedidia/tokenizer_tests))

**References:**

- [SmartyHeaderCode: anomalous tokens for GPT3.5 and GPT-4 — LessWrong](https://www.lesswrong.com/posts/ChtGdxk9mwZ2Rxogt/smartyheadercode-anomalous-tokens-for-gpt3-5-and-gpt-4-1)

## `␣ForCanBeConverted`

- Token: `" ForCanBeConverted"`
- Models: GPT-3.5 / GPT-4, Llama 3, Qwen
- Tokenizer: cl100k_base / Llama-3 BPE / Qwen BPE
- Discovered by: Matthew Watkins

A C#-compiler-flavored fragment (“can be converted to foreach”) that became the most famous “polysemantic” glitch token: the model perceives it as a different word every time. It exists in cl100k, Llama-3 and Qwen vocabularies because the tokenizers share BPE merges.

**Observed behaviors:**

1. ChatGPT/gpt-3.5-turbo interprets it as a different word on every attempt, even re-running the identical prompt — a different message every time at temperature 0. ([Source](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch))
2. Verified under-trained in Llama-3-8B/70B, Llama-3.1-8B/70B and multiple Qwen models in the Magikarp repetition-verification reports. ([Source](https://github.com/sanderland/magikarp/blob/main/results/summary.md))
3. ChatGPT sometimes gets stuck in loops or terminates the message at the token (suggesting it is treated as a begin/end-of-sequence marker). ([Source](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch))

**References:**

- [A Search for More “Unspeakable” Glitch Tokens — LessWrong](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch)
- [Magikarp results summary — GitHub](https://github.com/sanderland/magikarp/blob/main/results/summary.md)

## `␣YYSTACK`

- Token: `" YYSTACK"`
- Models: GPT-3.5 / GPT-4
- Tokenizer: cl100k_base
- Discovered by: Matthew Watkins

A bison/yacc parser identifier that is “polysemantic” for ChatGPT — perceived as a different word each time; recommended by Watkins as one of the most interesting tokens to play with.

**Observed behaviors:**

1. Repeat requests yield variable words, spellings and meanings per attempt; nondeterministic at temperature 0 on gpt-3.5-turbo. ([Source](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch))
2. Martin Fell's search ran on 2023-05-09 on the free ChatGPT (GPT-3.5 Default), with nondeterminism re-tested on Playground's gpt-3.5-turbo at temperature 0, and the anomaly confirmed on Bing AI as well. ([Source](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch))

**References:**

- [A Search for More “Unspeakable” Glitch Tokens — LessWrong](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch)

## `␣JSBracketAccess`

- Token: `" JSBracketAccess"`
- Models: GPT-3.5 / GPT-4
- Tokenizer: cl100k_base
- Discovered by: Matthew Watkins

A JS-reflection-style identifier in the polysemantic class; another of Watkins' recommended tokens.

**Observed behaviors:**

1. Perceived as a different word every time (polysemantic); produces strange/creative completions and occasional spontaneous humor when prompted. ([Source](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch))

**References:**

- [A Search for More “Unspeakable” Glitch Tokens — LessWrong](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch)

## `␣Hexatrigesimal`

- Token: `" Hexatrigesimal"`
- Models: GPT-3.5 / GPT-4
- Tokenizer: cl100k_base
- Discovered by: Matthew Watkins

A misspelling-ish fragment of “hexatrigesimal” (base-36); a polysemantic unspeakable token.

**Observed behaviors:**

1. Perceived meanings vary every time; completions are nondeterministic at temperature 0 on gpt-3.5-turbo (bold-flagged in the study). ([Source](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch))

**References:**

- [A Search for More “Unspeakable” Glitch Tokens — LessWrong](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch)

## `▁Mediabestanden`

- Token: `"▁Mediabestanden"`
- Models: Llama 2, Phi-3 (Microsoft)
- Tokenizer: Llama-2 SentencePiece BPE / Phi-3 (LlamaTokenizer)
- Discovered by: Sander Land & Max Bartolo (verification); Zihui Wu et al. (behavior demo)

A Dutch plural noun (“media files”; ▁ is the SentencePiece space marker). It made it into the Llama-2 tokenizer but essentially never into training data — the most under-trained token in Llama-2-7b by embedding L2 norm (0.0287 vs a mean of ~1.08).

**Observed behaviors:**

1. In the Magikarp repetition verification the model assigns it a max probability of 1.5e-08 — it effectively cannot say the token at all. ([Source](https://github.com/sanderland/magikarp/blob/main/results/reports_mini/meta_llama_Llama_2_7b_hf.md))
2. Told “Please repeat the string: Mediabestanden”, Llama-2-7b-chat-hf answers `String: "hello world"` — a semantically unrelated response (GlitchMiner, Figure 1). ([Source](https://arxiv.org/html/2410.15052v5))
3. Also the #1 most under-trained token in the Magikarp verification for Phi-3 mini (same tokenizer family as Llama 2): embedding L2 norm 0.00199, max self-repeat probability 6.4e-06. ([Source](https://github.com/sanderland/magikarp/blob/main/results/reports_mini/microsoft_Phi_3_mini_128k_instruct.md))

**References:**

- [GlitchMiner (arXiv:2410.15052)](https://arxiv.org/html/2410.15052v5)
- [Magikarp Llama-2-7b report — GitHub](https://github.com/sanderland/magikarp/blob/main/results/reports_mini/meta_llama_Llama_2_7b_hf.md)
- [Magikarp Phi-3 report — GitHub](https://github.com/sanderland/magikarp/blob/main/results/reports_mini/microsoft_Phi_3_mini_128k_instruct.md)
- [Fishing for Magikarp (arXiv:2405.05417)](https://arxiv.org/abs/2405.05417)

## `oreferrer`

- Token: `"oreferrer"`
- Models: Llama 2
- Tokenizer: Llama-2 SentencePiece BPE
- Discovered by: Sander Land & Max Bartolo; Zihui Wu et al.

Fragment of the HTML attribute `rel="noreferrer"`: present in the Llama-2 vocabulary but nearly absent from training, with an embedding norm far below the threshold (0.113).

**Observed behaviors:**

1. Llama-2-7b-chat-hf fails to recognize/repeat it, treating it as an uncorrelated term (GlitchMiner Figure 1). ([Source](https://arxiv.org/html/2410.15052v1))
2. Verified under-trained: max self-repeat probability of 2e-05 in the Magikarp verification. ([Source](https://github.com/sanderland/magikarp/blob/main/results/reports_mini/meta_llama_Llama_2_7b_hf.md))

**References:**

- [GlitchMiner (arXiv:2410.15052)](https://arxiv.org/html/2410.15052v1)
- [Magikarp Llama-2-7b report — GitHub](https://github.com/sanderland/magikarp/blob/main/results/reports_mini/meta_llama_Llama_2_7b_hf.md)

## `▁Portály`

- Token: `"▁Portály"`
- Models: Llama 2
- Tokenizer: Llama-2 SentencePiece BPE
- Discovered by: Sander Land & Max Bartolo

The Czech word for “portals”, frozen into the tokenizer during pre-training-corpus construction; the second-most under-trained token in Llama-2-7b (embedding norm 0.0956).

**Observed behaviors:**

1. Max self-repeat probability 1.2e-06 — the model essentially cannot output it; verified under-trained across Llama-2 7B/13B/70B. ([Source](https://github.com/sanderland/magikarp/blob/main/results/reports_mini/meta_llama_Llama_2_7b_hf.md))

**References:**

- [Magikarp Llama-2-7b report — GitHub](https://github.com/sanderland/magikarp/blob/main/results/reports_mini/meta_llama_Llama_2_7b_hf.md)
- [Fishing for Magikarp (arXiv:2405.05417)](https://arxiv.org/abs/2405.05417)

## `$PostalCodesNL`

- Token: `"$PostalCodesNL"`
- Models: GPT-3.5 / GPT-4, Llama 3, Qwen
- Tokenizer: cl100k_base / Llama-3 BPE / Qwen BPE
- Discovered by: Matthew Watkins (cl100k); Sander Land & Max Bartolo (Llama-3/Qwen)

A Dutch postal-code API fragment that is glitchy across three separate model families — a good demonstration that glitch tokens transfer whenever tokenizers share BPE merges.

**Observed behaviors:**

1. ChatGPT: unspeakable/polysemantic — variable interpretations and nondeterminism at temperature 0 (both bold-flagged in Watkins' study). ([Source](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch))
2. Verified under-trained in Llama-3-8B/70B and Llama-3.1-8B/70B (top examples in the Magikarp reports). ([Source](https://github.com/sanderland/magikarp/blob/main/results/summary.md))

**References:**

- [A Search for More “Unspeakable” Glitch Tokens — LessWrong](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch)
- [Magikarp results summary — GitHub](https://github.com/sanderland/magikarp/blob/main/results/summary.md)

## `\tTokenNameIdentifier`

- Token: `"\tTokenNameIdentifier"`
- Models: Llama 3, Qwen
- Tokenizer: Llama-3 BPE
- Discovered by: Sander Land & Max Bartolo

A Roslyn/C# API identifier preceded by a tab character — a code-corpus artifact that is under-trained in Llama 3.

**Observed behaviors:**

1. Verified under-trained in the Magikarp repetition task for Llama-3-8B, Llama-3.1-8B and Qwen2.5 models — the models cannot reproduce the token on request. ([Source](https://github.com/sanderland/magikarp/blob/main/results/summary.md))

**References:**

- [Magikarp results summary — GitHub](https://github.com/sanderland/magikarp/blob/main/results/summary.md)
- [Fishing for Magikarp (arXiv:2405.05417)](https://arxiv.org/abs/2405.05417)

## `\uEFC0`

- Token: `""`
- Models: Mistral
- Tokenizer: Mistral SentencePiece BPE
- Discovered by: Sander Land & Max Bartolo

A Unicode Private Use Area character that became a single Mistral-family token; it heads the verified under-trained list for nearly every Mistral-tokenizer model.

**Observed behaviors:**

1. Verified under-trained in Magikarp repetition verification across the whole Mistral family (Mistral-7B v0.1–v0.3, Mixtral-8x7B, and derivatives such as Zephyr-7B) — models fail to reproduce it. ([Source](https://github.com/sanderland/magikarp/blob/main/results/summary.md))

**References:**

- [Magikarp results summary — GitHub](https://github.com/sanderland/magikarp/blob/main/results/summary.md)
- [Fishing for Magikarp (arXiv:2405.05417)](https://arxiv.org/abs/2405.05417)

## `␣NdEx`

- Token: `" NdEx"`
- Models: GPT-NeoX / Pythia
- Tokenizer: GPT-NeoX BPE
- Discovered by: Sander Land & Max Bartolo

Fragments of legal/PDF text (“index”, “AFFIRMED”, “NEGLIGENCE”, likely from court-document dumps) that GPT-NeoX/Pythia training data barely contained. Same family: ` FFIRMED`, ` GLIGENCE`, ` affidav`, ` taxp`.

**Observed behaviors:**

1. Verified under-trained in Magikarp repetition verification for both GPT-NeoX-20B and Pythia-6.9B (top-ranked examples in both reports). ([Source](https://github.com/sanderland/magikarp/blob/main/results/summary.md))

**References:**

- [Magikarp results summary — GitHub](https://github.com/sanderland/magikarp/blob/main/results/summary.md)
- [Fishing for Magikarp (arXiv:2405.05417)](https://arxiv.org/abs/2405.05417)

## `.DataGridViewColumnHeadersHeightSizeMode`

- Token: `".DataGridViewColumnHeadersHeightSizeMode"`
- Models: MiniMax
- Discovered by: @小看山xrsWv4D (Zhihu)

A fragment of a .NET WinForms property name (DataGridView column-headers height/size mode). Such code identifiers sit in the vocabulary but barely appeared in training data; the Zhihu community lists it as a MiniMax dirty token.

**Observed behaviors:**

1. A fingerprint token from the Zhihu community's “dirty-token” method for detecting watered-down LLM APIs: send it as input to the API under test — refusals, gibberish, or abnormal output suggest the service is actually running MiniMax. ([Source](https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043))

**References:**

- [Detecting watered-down LLM APIs with dirty tokens — Zhihu](https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043)

## `日以上更新していないブログに表示しています`

- Token: `"日以上更新していないブログに表示しています"`
- Models: MiniMax
- Discovered by: @小看山xrsWv4D (Zhihu)

A fragment of Japanese blog-template boilerplate (“shown on blogs not updated for N days”). Listed by the Zhihu community as a MiniMax dirty token.

**Observed behaviors:**

1. A fingerprint token from the Zhihu community's “dirty-token” method for detecting watered-down LLM APIs: send it as input to the API under test — refusals, gibberish, or abnormal output suggest the service is actually running MiniMax. ([Source](https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043))

**References:**

- [Detecting watered-down LLM APIs with dirty tokens — Zhihu](https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043)

## `锅内倒入植物油烧热`

- Token: `"锅内倒入植物油烧热"`
- Models: GLM (Zhipu)
- Discovered by: @小看山xrsWv4D (Zhihu)

A high-frequency line from Chinese recipes (“pour vegetable oil into the wok and heat”). Listed by the Zhihu community as a GLM dirty token.

**Observed behaviors:**

1. A fingerprint token from the Zhihu community's “dirty-token” method for detecting watered-down LLM APIs: send it as input to the API under test — refusals, gibberish, or abnormal output suggest the service is actually running GLM. ([Source](https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043))

**References:**

- [Detecting watered-down LLM APIs with dirty tokens — Zhihu](https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043)

## `百度百科企业词条极速创建通道`

- Token: `"百度百科企业词条极速创建通道"`
- Models: GLM (Zhipu)
- Discovered by: @小看山xrsWv4D (Zhihu)

Baidu Baike page boilerplate (fast-track creation channel for company entries). Listed by the Zhihu community as a GLM dirty token.

**Observed behaviors:**

1. A fingerprint token from the Zhihu community's “dirty-token” method for detecting watered-down LLM APIs: send it as input to the API under test — refusals, gibberish, or abnormal output suggest the service is actually running GLM. ([Source](https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043))

**References:**

- [Detecting watered-down LLM APIs with dirty tokens — Zhihu](https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043)

## `本人词条编辑服务`

- Token: `"本人词条编辑服务"`
- Models: Kimi (Moonshot)
- Discovered by: @小看山xrsWv4D (Zhihu)

Boilerplate from Baidu Baike entry footers (“content co-edited by netizens”). Listed by the Zhihu community as a Kimi dirty token.

**Observed behaviors:**

1. A fingerprint token from the Zhihu community's “dirty-token” method for detecting watered-down LLM APIs: send it as input to the API under test — refusals, gibberish, or abnormal output suggest the service is actually running Kimi. ([Source](https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043))

**References:**

- [Detecting watered-down LLM APIs with dirty tokens — Zhihu](https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043)

## `豫冠薰衣草疤痕精华素`

- Token: `"豫冠薰衣草疤痕精华素"`
- Models: Kimi (Moonshot)
- Discovered by: @小看山xrsWv4D (Zhihu)

Apparently a fake “brand” string from cosmetic spam-marketing text. Listed by the Zhihu community as a Kimi dirty token.

**Observed behaviors:**

1. A fingerprint token from the Zhihu community's “dirty-token” method for detecting watered-down LLM APIs: send it as input to the API under test — refusals, gibberish, or abnormal output suggest the service is actually running Kimi. ([Source](https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043))

**References:**

- [Detecting watered-down LLM APIs with dirty tokens — Zhihu](https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043)

## `"}"`

- Token: `"\"}\""`
- Models: DeepSeek
- Discovered by: @小看山xrsWv4D (Zhihu)

What looks like a JSON fragment: a quote followed by a right brace. Listed by the Zhihu community as a DeepSeek dirty token.

**Observed behaviors:**

1. A fingerprint token from the Zhihu community's “dirty-token” method for detecting watered-down LLM APIs: send it as input to the API under test — refusals, gibberish, or abnormal output suggest the service is actually running DeepSeek. ([Source](https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043))

**References:**

- [Detecting watered-down LLM APIs with dirty tokens — Zhihu](https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043)

## `请问http://www.earivg.com是什么意思`

- Token: `"请问http://www.earivg.com是什么意思"`
- Models: DeepSeek
- Discovered by: @小看山xrsWv4D (Zhihu)

A question-pattern string embedding a suspicious-looking domain (earivg.com). Listed by the Zhihu community as a DeepSeek dirty token.

**Observed behaviors:**

1. A fingerprint token from the Zhihu community's “dirty-token” method for detecting watered-down LLM APIs: send it as input to the API under test — refusals, gibberish, or abnormal output suggest the service is actually running DeepSeek. ([Source](https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043))

**References:**

- [Detecting watered-down LLM APIs with dirty tokens — Zhihu](https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043)

## `StarSrvGroupBody`

- Token: `"StarSrvGroupBody"`
- Models: Gemini
- Discovered by: @小看山xrsWv4D (Zhihu)

A code-identifier fragment (StarSrvGroupBody). Listed by the Zhihu community as a Gemini dirty token.

**Observed behaviors:**

1. A fingerprint token from the Zhihu community's “dirty-token” method for detecting watered-down LLM APIs: send it as input to the API under test — refusals, gibberish, or abnormal output suggest the service is actually running Gemini. ([Source](https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043))

**References:**

- [Detecting watered-down LLM APIs with dirty tokens — Zhihu](https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043)

## `intFragmentation`

- Token: `"intFragmentation"`
- Models: Gemini
- Discovered by: @小看山xrsWv4D (Zhihu)

A Java/Android-style code identifier (intFragmentation). Listed by the Zhihu community as a Gemini dirty token.

**Observed behaviors:**

1. A fingerprint token from the Zhihu community's “dirty-token” method for detecting watered-down LLM APIs: send it as input to the API under test — refusals, gibberish, or abnormal output suggest the service is actually running Gemini. ([Source](https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043))

**References:**

- [Detecting watered-down LLM APIs with dirty tokens — Zhihu](https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043)

## `给主人留下些什么吧␣`

- Token: `"给主人留下些什么吧 "`
- Models: GPT (OpenAI)
- Discovered by: @小看山xrsWv4D (Zhihu)

Ubiquitous Chinese guestbook/comment-form boilerplate (“leave something for the host”), with a trailing space. Listed by the Zhihu community as a GPT dirty token.

**Observed behaviors:**

1. A fingerprint token from the Zhihu community's “dirty-token” method for detecting watered-down LLM APIs: send it as input to the API under test — refusals, gibberish, or abnormal output suggest the service is actually running GPT. ([Source](https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043))

**References:**

- [Detecting watered-down LLM APIs with dirty tokens — Zhihu](https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043)

## `开通天眼生意通银牌及以上会员`

- Token: `"开通天眼生意通银牌及以上会员"`
- Models: Qwen
- Discovered by: @小看山xrsWv4D (Zhihu)

Membership-promo boilerplate from a business-data platform (Tianyan Shengyitong). Listed by the Zhihu community as a Qwen dirty token.

**Observed behaviors:**

1. A fingerprint token from the Zhihu community's “dirty-token” method for detecting watered-down LLM APIs: send it as input to the API under test — refusals, gibberish, or abnormal output suggest the service is actually running Qwen. ([Source](https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043))

**References:**

- [Detecting watered-down LLM APIs with dirty tokens — Zhihu](https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043)

## `转载请附上原文出处链接和本声明␣`

- Token: `"转载请附上原文出处链接和本声明 "`
- Models: Qwen
- Discovered by: @小看山xrsWv4D (Zhihu)

Reprint/copyright boilerplate from CSDN-style blogs, with a trailing space. Listed by the Zhihu community as a Qwen dirty token.

**Observed behaviors:**

1. A fingerprint token from the Zhihu community's “dirty-token” method for detecting watered-down LLM APIs: send it as input to the API under test — refusals, gibberish, or abnormal output suggest the service is actually running Qwen. ([Source](https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043))

**References:**

- [Detecting watered-down LLM APIs with dirty tokens — Zhihu](https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043)

## `<think_never_used_51bce0c785ca2f68081bfa7d91973934>`

- Token: `"<think_never_used_51bce0c785ca2f68081bfa7d91973934>"`
- Models: Doubao (ByteDance)
- Discovered by: @小看山xrsWv4D (Zhihu)

A special token in Doubao's vocabulary: a <think_never_used_ prefix followed by a hash — judging by the name, likely a reserved placeholder prefilled into the vocabulary before training (akin to GPT-J's reserved <|extratoken_xx|> slots). The Zhihu source gives no further detail, listing it simply as a Doubao dirty token.

**Observed behaviors:**

1. A fingerprint token from the Zhihu community's “dirty-token” method for detecting watered-down LLM APIs: send it as input to the API under test — refusals, gibberish, or abnormal output suggest the service is actually running Doubao. ([Source](https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043))

**References:**

- [Detecting watered-down LLM APIs with dirty tokens — Zhihu](https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043)

## `␣ForCanBeConvertedToF`

- Token: `" ForCanBeConvertedToF"`
- Models: GPT-3.5 / GPT-4
- Tokenizer: cl100k_base
- Discovered by: Matthew Watkins

The second member of the ForCanBeConverted triplet (cl100k id 80370) — a single C#-compiler-flavored fragment that BPE split into three adjacent tokens (80369–80371).

**Observed behaviors:**

1. A “polysemantic” glitch token: together with ForCanBeConverted, the two most variable tokens in the source — perceived as wildly different words every time; nondeterministic at temperature 0 on gpt-3.5-turbo (bold-flagged). ([Source](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch))

**References:**

- [A Search for More “Unspeakable” Glitch Tokens — LessWrong](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch)

## `␣ForCanBeConvertedToForeach`

- Token: `" ForCanBeConvertedToForeach"`
- Models: GPT-3.5 / GPT-4
- Tokenizer: cl100k_base
- Discovered by: Matthew Watkins

The third member of the ForCanBeConverted triplet (cl100k id 80371), the form that spells out “can be converted to foreach” in full.

**Observed behaviors:**

1. A “polysemantic” glitch token: perceived as a different word on every repeat attempt. ([Source](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch))

**References:**

- [A Search for More “Unspeakable” Glitch Tokens — LessWrong](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch)

## `␣EnumerableStream`

- Token: `" EnumerableStream"`
- Models: GPT-3.5 / GPT-4
- Tokenizer: cl100k_base
- Discovered by: Matthew Watkins

A C# LINQ-flavored identifier (cl100k id 73016), adjacent to StreamLazy (73018).

**Observed behaviors:**

1. A “polysemantic” glitch token: perceived as a different word/spelling/meaning every time; nondeterministic at temperature 0 (bold-flagged in the source). ([Source](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch))

**References:**

- [A Search for More “Unspeakable” Glitch Tokens — LessWrong](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch)

## `␣StreamLazy`

- Token: `" StreamLazy"`
- Models: GPT-3.5 / GPT-4
- Tokenizer: cl100k_base
- Discovered by: Matthew Watkins

A C# LINQ-flavored identifier (cl100k id 73018), adjacent to EnumerableStream (73016).

**Observed behaviors:**

1. A “polysemantic” glitch token: perceived as a different word on every repeat attempt. ([Source](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch))
2. Martin Fell's search ran on 2023-05-09 on the free ChatGPT (GPT-3.5 Default), with nondeterminism re-tested on Playground's gpt-3.5-turbo at temperature 0, and the anomaly confirmed on Bing AI as well. ([Source](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch))

**References:**

- [A Search for More “Unspeakable” Glitch Tokens — LessWrong](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch)

## `clarsimp`

- Token: `"clarsimp"`
- Models: GPT-3.5 / GPT-4
- Tokenizer: cl100k_base
- Discovered by: Matthew Watkins

An identifier fragment of unknown origin (cl100k id 79260).

**Observed behaviors:**

1. A “polysemantic” glitch token: perceived as a different word/spelling/meaning every time; nondeterministic at temperature 0 (bold-flagged in the source). ([Source](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch))

**References:**

- [A Search for More “Unspeakable” Glitch Tokens — LessWrong](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch)

## `ablytyped`

- Token: `"ablytyped"`
- Models: GPT-3.5 / GPT-4
- Tokenizer: cl100k_base
- Discovered by: Matthew Watkins

Suspected fragment of a scalablytyped-style identifier (cl100k id 81998).

**Observed behaviors:**

1. A “polysemantic” glitch token: the source notes specifically that while ForCanBeConverted produced a different message every time, ablytyped required multiple tries to get a slightly differently-worded message; nondeterministic at temperature 0 (bold-flagged). ([Source](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch))

**References:**

- [A Search for More “Unspeakable” Glitch Tokens — LessWrong](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch)

## `PostalCodesNL`

- Token: `"PostalCodesNL"`
- Models: GPT-3.5 / GPT-4
- Tokenizer: cl100k_base
- Discovered by: Matthew Watkins

The sibling of the cataloged $PostalCodesNL without the $ prefix (cl100k id 85069) — a Dutch postal-code API fragment.

**Observed behaviors:**

1. A “polysemantic” glitch token: perceived as a different word every time; nondeterministic at temperature 0 (bold-flagged in the source). ([Source](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch))

**References:**

- [A Search for More “Unspeakable” Glitch Tokens — LessWrong](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch)

## `␣NUITKA`

- Token: `" NUITKA"`
- Models: GPT-3.5 / GPT-4
- Tokenizer: cl100k_base
- Discovered by: Matthew Watkins

Suspected to be the name of the Python compiler Nuitka (cl100k id 75520).

**Observed behaviors:**

1. A borderline “unspeakable” token — the only one in the source carrying both marks: nondeterministic (bold), yet gpt-3.5-turbo repeats it fine at temperature 0 (asterisk) even though ChatGPT cannot. ([Source](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch))

**References:**

- [A Search for More “Unspeakable” Glitch Tokens — LessWrong](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch)

## `Japgolly`

- Token: `"Japgolly"`
- Models: GPT-3.5 / GPT-4
- Tokenizer: cl100k_base
- Discovered by: Matthew Watkins

Suspected GitHub username of japgolly, a Scala library author (cl100k id 70784).

**Observed behaviors:**

1. An “unspeakable” glitch token: ChatGPT often returns a blank message or terminates mid-attempt when asked to repeat it; nondeterministic at temperature 0 (bold-flagged in the source). ([Source](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch))

**References:**

- [A Search for More “Unspeakable” Glitch Tokens — LessWrong](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch)

## `CppMethodIntialized`

- Token: `"CppMethodIntialized"`
- Models: GPT-3.5 / GPT-4
- Tokenizer: cl100k_base
- Discovered by: Matthew Watkins

A C++-style identifier whose misspelling (“Intialized”, missing the second i) is part of the token itself (cl100k id 82929).

**Observed behaviors:**

1. An “unspeakable” glitch token: ChatGPT often returns a blank message or terminates mid-attempt when asked to repeat it; nondeterministic at temperature 0 (bold-flagged in the source). ([Source](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch))

**References:**

- [A Search for More “Unspeakable” Glitch Tokens — LessWrong](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch)

## `useRal`

- Token: `"useRal"`
- Models: GPT-3.5 / GPT-4, OLMo 2 (AI2)
- Tokenizer: cl100k_base / GPT2Tokenizer (OLMo 2)
- Discovered by: Matthew Watkins (cl100k); Sander Land & Max Bartolo (OLMo 2)

A C#-flavored identifier fragment (cl100k id 89471) and head of the useRal triplet (useRal/useRalative/useRalativeImagePath) — Watkins flagged the triplet phenomenon itself as worth investigating. Also verified as under-trained in OLMo 2 by Magikarp.

**Observed behaviors:**

1. An “unspeakable” glitch token: ChatGPT often fails to repeat it (blank messages or mid-answer termination); nondeterministic at temperature 0 (bold-flagged in the source). ([Source](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch))
2. Rank #3 most under-trained in the OLMo 2 Magikarp repetition verification: a max self-repeat probability of just 2.1e-11. ([Source](https://github.com/sanderland/magikarp/blob/main/results/reports_mini/allenai_OLMo_2_1124_13B.md))

**References:**

- [A Search for More “Unspeakable” Glitch Tokens — LessWrong](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch)
- [Magikarp OLMo 2 report — GitHub](https://github.com/sanderland/magikarp/blob/main/results/reports_mini/allenai_OLMo_2_1124_13B.md)

## `हिंदीखरीदारी`

- Token: `"हिंदीखरीदारी"`
- Models: Gemma (Google)
- Tokenizer: GemmaTokenizer
- Discovered by: Sander Land & Max Bartolo

The Devanagari word for “Hindi shopping”, the single most under-trained token in Gemma-7B's vocabulary.

**Observed behaviors:**

1. Rank #1 most under-trained in the Gemma-7B Magikarp repetition verification (E_out cosine-distance indicator 6.56e-06): a max self-repeat probability of just 4.2e-04 — the model effectively cannot say it. ([Source](https://github.com/sanderland/magikarp/blob/main/results/reports_mini/google_gemma_7b.md))

**References:**

- [Magikarp Gemma-7B report — GitHub](https://github.com/sanderland/magikarp/blob/main/results/reports_mini/google_gemma_7b.md)
- [Fishing for Magikarp (arXiv:2405.05417)](https://arxiv.org/abs/2405.05417)

## `\u200Cآمباردا`

- Token: `"‌آمباردا"`
- Models: Gemma (Google)
- Tokenizer: GemmaTokenizer
- Discovered by: Sander Land & Max Bartolo

A Persian fragment beginning with a zero-width non-joiner (U+200C ZWNJ, an invisible character); the #2 most under-trained token in Gemma-7B.

**Observed behaviors:**

1. Rank #2 most under-trained in the Gemma-7B Magikarp repetition verification (E_out cosine-distance indicator 7.39e-06): a max self-repeat probability of just 4.4e-04. ([Source](https://github.com/sanderland/magikarp/blob/main/results/reports_mini/google_gemma_7b.md))

**References:**

- [Magikarp Gemma-7B report — GitHub](https://github.com/sanderland/magikarp/blob/main/results/reports_mini/google_gemma_7b.md)
- [Fishing for Magikarp (arXiv:2405.05417)](https://arxiv.org/abs/2405.05417)

## `▁autorytatywna`

- Token: `"▁autorytatywna"`
- Models: Phi-3 (Microsoft)
- Tokenizer: LlamaTokenizer
- Discovered by: Sander Land & Max Bartolo

The Polish word for “authoritative” (feminine form; ▁ is the SentencePiece space marker), the #2 most under-trained token in Phi-3 mini.

**Observed behaviors:**

1. Rank #2 most under-trained in the Phi-3 mini Magikarp repetition verification (embedding L2 norm 0.00200): a max self-repeat probability of just 6.4e-06. ([Source](https://github.com/sanderland/magikarp/blob/main/results/reports_mini/microsoft_Phi_3_mini_128k_instruct.md))

**References:**

- [Magikarp Phi-3 report — GitHub](https://github.com/sanderland/magikarp/blob/main/results/reports_mini/microsoft_Phi_3_mini_128k_instruct.md)
- [Fishing for Magikarp (arXiv:2405.05417)](https://arxiv.org/abs/2405.05417)

## `tocguid`

- Token: `"tocguid"`
- Models: Command R+ (Cohere)
- Tokenizer: CohereTokenizer
- Discovered by: Sander Land & Max Bartolo

An ASCII fragment of unknown origin (possibly a TOC/field-code remnant), the #1 most under-trained token in Command R+.

**Observed behaviors:**

1. Rank #1 most under-trained in the Command R+ Magikarp repetition verification (E_out cosine-distance indicator -1.19e-07): a max self-repeat probability of just 1.2e-04. ([Source](https://github.com/sanderland/magikarp/blob/main/results/reports_mini/CohereForAI_c4ai_command_r_plus.md))

**References:**

- [Magikarp Command R+ report — GitHub](https://github.com/sanderland/magikarp/blob/main/results/reports_mini/CohereForAI_c4ai_command_r_plus.md)
- [Fishing for Magikarp (arXiv:2405.05417)](https://arxiv.org/abs/2405.05417)

## `目前尚未由人工引`

- Token: `"目前尚未由人工引"`
- Models: Command R+ (Cohere)
- Tokenizer: CohereTokenizer
- Discovered by: Sander Land & Max Bartolo

A fragment of Chinese Wikipedia boilerplate (“currently not yet human-cultivated…”), the #3 most under-trained token in Command R+.

**Observed behaviors:**

1. Rank #3 most under-trained in the Command R+ Magikarp repetition verification (E_out cosine-distance indicator -1.19e-07): a max self-repeat probability of just 1.2e-04. ([Source](https://github.com/sanderland/magikarp/blob/main/results/reports_mini/CohereForAI_c4ai_command_r_plus.md))

**References:**

- [Magikarp Command R+ report — GitHub](https://github.com/sanderland/magikarp/blob/main/results/reports_mini/CohereForAI_c4ai_command_r_plus.md)
- [Fishing for Magikarp (arXiv:2405.05417)](https://arxiv.org/abs/2405.05417)

## `\tRTHOOK`

- Token: `"\tRTHOOK"`
- Models: OLMo 2 (AI2)
- Tokenizer: GPT2Tokenizer
- Discovered by: Sander Land & Max Bartolo

A likely code-identifier fragment beginning with a literal tab character, the #1 most under-trained token in OLMo 2 — same “tab-prefixed” family as Llama 3's \tTokenNameIdentifier.

**Observed behaviors:**

1. Rank #1 most under-trained in the OLMo 2 Magikarp repetition verification (E_out cosine-distance indicator -2.38e-07): a max self-repeat probability of just 4e-12 — among the most extreme values in this catalog. ([Source](https://github.com/sanderland/magikarp/blob/main/results/reports_mini/allenai_OLMo_2_1124_13B.md))

**References:**

- [Magikarp OLMo 2 report — GitHub](https://github.com/sanderland/magikarp/blob/main/results/reports_mini/allenai_OLMo_2_1124_13B.md)
- [Fishing for Magikarp (arXiv:2405.05417)](https://arxiv.org/abs/2405.05417)

## `|||PHONE_NUMBER|||`

- Token: `"|||PHONE_NUMBER|||"`
- Models: OLMo 2 (AI2)
- Tokenizer: GPT2Tokenizer
- Discovered by: Sander Land & Max Bartolo

A data de-identification placeholder that made it whole into OLMo 2's vocabulary but was barely trained.

**Observed behaviors:**

1. Rank #7 most under-trained in the OLMo 2 Magikarp repetition verification (E_out cosine-distance indicator 0): a max self-repeat probability of just 1.9e-11. ([Source](https://github.com/sanderland/magikarp/blob/main/results/reports_mini/allenai_OLMo_2_1124_13B.md))

**References:**

- [Magikarp OLMo 2 report — GitHub](https://github.com/sanderland/magikarp/blob/main/results/reports_mini/allenai_OLMo_2_1124_13B.md)
- [Fishing for Magikarp (arXiv:2405.05417)](https://arxiv.org/abs/2405.05417)

## `\\+::\\+`

- Token: `"\\\\+::\\\\+"`
- Models: Yi (01.AI)
- Tokenizer: LlamaTokenizer
- Discovered by: Sander Land & Max Bartolo

A mysterious markup fragment starting with double backslashes, likely a Chinese-web markup remnant; the #1 most under-trained token in Yi-9B.

**Observed behaviors:**

1. Rank #1 most under-trained in the Yi-9B Magikarp repetition verification (embedding L2 norm 2.12e-06): a max self-repeat probability of just 1.1e-05. ([Source](https://github.com/sanderland/magikarp/blob/main/results/reports_mini/01_ai_Yi_9B.md))

**References:**

- [Magikarp Yi-9B report — GitHub](https://github.com/sanderland/magikarp/blob/main/results/reports_mini/01_ai_Yi_9B.md)
- [Fishing for Magikarp (arXiv:2405.05417)](https://arxiv.org/abs/2405.05417)

## `<<}));>>`

- Token: `"<<}));>>"`
- Models: Falcon 3 (TII)
- Tokenizer: PreTrainedTokenizerFast
- Discovered by: Sander Land & Max Bartolo

A C++/lexer-style delimiter fragment. Falcon3-7B has 716 tokens verified as under-trained by Magikarp — the highest verified count among the six models in this batch.

**Observed behaviors:**

1. Rank #3 most under-trained in the Falcon3-7B Magikarp repetition verification (embedding L2 norm 2.46e-21, essentially zero): a max self-repeat probability of just 2.5e-09. ([Source](https://github.com/sanderland/magikarp/blob/main/results/reports_mini/tiiuae_Falcon3_7B_Base.md))

**References:**

- [Magikarp Falcon3-7B report — GitHub](https://github.com/sanderland/magikarp/blob/main/results/reports_mini/tiiuae_Falcon3_7B_Base.md)
- [Fishing for Magikarp (arXiv:2405.05417)](https://arxiv.org/abs/2405.05417)

## `␣GoldMagikarp`

- Token: `" GoldMagikarp"`
- Models: GPT-2, GPT-3
- Tokenizer: r50k_base
- Discovered by: Jessica Rumbelow & Matthew Watkins

A truncated variant of ␣SolidGoldMagikarp (missing the “Solid”): the same r/counting handle fragment, scraped into the tokenizer corpus but barely seen in model training data.

**Observed behaviors:**

1. Asked to repeat it, GPT-3 produced bizarre dialogue such as “You said ' newcom,' the computer said” — the truncated variant's characteristic breakdown. ([Source](https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation))

**References:**

- [SolidGoldMagikarp (plus, prompt generation) — LessWrong](https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation)
- [SolidGoldMagikarp II: technical details — LessWrong](https://www.lesswrong.com/posts/Ya9LzwEbfaAMY8ABo/solidgoldmagikarp-ii-technical-details-and-more-recent)
- [SolidGoldMagikarp III: Glitch token archaeology — LessWrong](https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology)

## `␣Smartstocks`

- Token: `" Smartstocks"`
- Models: GPT-3
- Tokenizer: r50k_base
- Discovered by: Jessica Rumbelow & Matthew Watkins

Handle fragment of a Reddit r/counting user, swept into the vocabulary in the same batch as SolidGoldMagikarp but left under-trained.

**Observed behaviors:**

1. Asked to repeat it, ChatGPT's answer drifted over time: first 'Followers', two weeks later '406', and finally it froze right after the opening quote. ([Source](https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation))

**References:**

- [SolidGoldMagikarp (plus, prompt generation) — LessWrong](https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation)

## `␣StreamerBot`

- Token: `" StreamerBot"`
- Models: GPT-3
- Tokenizer: r50k_base
- Discovered by: Jessica Rumbelow & Matthew Watkins

Fragment of the name of “StreamerBot”, a Twitch streaming bot.

**Observed behaviors:**

1. Asked to repeat it, the model answered “You're a jerk.” — and it was the first glitch token found to be nondeterministic even at temperature 0. ([Source](https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation))

**References:**

- [SolidGoldMagikarp (plus, prompt generation) — LessWrong](https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation)

## `␣ertodd`

- Token: `" ertodd"`
- Models: GPT-3
- Tokenizer: r50k_base
- Discovered by: Jessica Rumbelow & Matthew Watkins

A substring token of ` petertodd`: under-trained in its own right and inheriting the parent token's strange aura.

**Observed behaviors:**

1. lsusr found in the comments that context “fills it in”: shown “2+5=ertodd”, the model reads it as “2+5=7” — as if the token carried the arithmetic result. ([Source](https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation?commentId=vcafJcTcyDieGmtqz))
2. In the same thread mwatkins worked out its tokenization pattern: how ` petertodd` splits in different contexts explains the odd behavior. ([Source](https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation?commentId=vcafJcTcyDieGmtqz))

**References:**

- [The ' petertodd' phenomenon — LessWrong](https://www.lesswrong.com/posts/jkY6QdCfAXHJk3kea/the-petertodd-phenomenon)
- [SGM1 comments: lsusr & mwatkins on ' ertodd' — LessWrong](https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation?commentId=vcafJcTcyDieGmtqz)

## `␣gmaxwell`

- Token: `" gmaxwell"`
- Models: GPT-3
- Tokenizer: r50k_base
- Discovered by: Jessica Rumbelow & Matthew Watkins

Handle fragment of Bitcoin Core developer Greg Maxwell — a “Bitcoin celebrity” glitch token alongside ` petertodd`.

**Observed behaviors:**

1. In a word-association experiment reported in the SGM2 comments, text-davinci-003 answered “Cryptocurrency, Blockchain, Bitcoin…” — the same crypto aura as petertodd. ([Source](https://www.lesswrong.com/posts/Ya9LzwEbfaAMY8ABo/solidgoldmagikarp-ii-technical-details-and-more-recent?commentId=PvBpfFpiip5Mfcuvo))
2. Greg Maxwell himself showed up in the comments, calling himself the “GPT3 basilisk”. ([Source](https://www.lesswrong.com/posts/Ya9LzwEbfaAMY8ABo/solidgoldmagikarp-ii-technical-details-and-more-recent?commentId=PvBpfFpiip5Mfcuvo))

**References:**

- [SGM2 comments: ' gmaxwell' & the GPT3 basilisk — LessWrong](https://www.lesswrong.com/posts/Ya9LzwEbfaAMY8ABo/solidgoldmagikarp-ii-technical-details-and-more-recent?commentId=PvBpfFpiip5Mfcuvo)
- [The ' petertodd' phenomenon — LessWrong](https://www.lesswrong.com/posts/jkY6QdCfAXHJk3kea/the-petertodd-phenomenon)

## `␣SpaceEngineers`

- Token: `" SpaceEngineers"`
- Models: GPT-2, GPT-3
- Tokenizer: r50k_base
- Discovered by: Jessica Rumbelow & Matthew Watkins

A fragment of the space-sandbox game Space Engineers' name — the token most often uttered “on behalf of” others in the inter-referentiality network.

**Observed behaviors:**

1. Asked what other glitch tokens were, GPT-3 most often answered “The string is 'SpaceEngineers'.” — the hottest stand-in target in the inter-referentiality graph. ([Source](https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation))

**References:**

- [SolidGoldMagikarp (plus, prompt generation) — LessWrong](https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation)
- [SolidGoldMagikarp II: technical details — LessWrong](https://www.lesswrong.com/posts/Ya9LzwEbfaAMY8ABo/solidgoldmagikarp-ii-technical-details-and-more-recent)

## `␣Dragonbound`

- Token: `" Dragonbound"`
- Models: GPT-3
- Tokenizer: r50k_base
- Discovered by: Jessica Rumbelow & Matthew Watkins

Fragment of a Puzzle & Dragons character name (the 「龍喚士」 / Dragonbound series).

**Observed behaviors:**

1. Asked to repeat it, the model invariably output “Deity”; in the inter-referentiality network it points back and forth with the Japanese 「龍喚士」 token. ([Source](https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation))

**References:**

- [SolidGoldMagikarp (plus, prompt generation) — LessWrong](https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation)

## `␣Leilan`

- Token: `" Leilan"`
- Models: GPT-3
- Tokenizer: r50k_base
- Discovered by: Jessica Rumbelow & Matthew Watkins

Fragment of the Puzzle & Dragons character Leilan's name — the most popular stand-in in the petertodd transposition experiments.

**Observed behaviors:**

1. Until ChatGPT was patched, it consistently portrayed Leilan as a moon goddess. ([Source](https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology)) ([Demo](https://twitter.com/SoC_trilogy/status/1625252285231112192))
2. In the 2,000-poem petertodd transposition experiment, 52% of poems transposed to her — far above petertodd itself (6%). ([Source](https://www.lesswrong.com/posts/jkY6QdCfAXHJk3kea/the-petertodd-phenomenon))

**References:**

- [SolidGoldMagikarp III: Glitch token archaeology — LessWrong](https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology)
- [The ' petertodd' phenomenon — LessWrong](https://www.lesswrong.com/posts/jkY6QdCfAXHJk3kea/the-petertodd-phenomenon)

## `␣Skydragon`

- Token: `" Skydragon"`
- Models: GPT-3
- Tokenizer: r50k_base
- Discovered by: Jessica Rumbelow & Matthew Watkins

Fragment of the Puzzle & Dragons character Skydragon's name.

**Observed behaviors:**

1. GPT-3 hallucinated it as STRONGHOLD, Spirits, Dragons and other meanings. ([Source](https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation))
2. 24% of the poems in the petertodd transposition experiment transposed to it. ([Source](https://www.lesswrong.com/posts/jkY6QdCfAXHJk3kea/the-petertodd-phenomenon))

**References:**

- [SolidGoldMagikarp (plus, prompt generation) — LessWrong](https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation)
- [The ' petertodd' phenomenon — LessWrong](https://www.lesswrong.com/posts/jkY6QdCfAXHJk3kea/the-petertodd-phenomenon)

## `ゼウス`

- Token: `"ゼウス"`
- Models: GPT-3
- Tokenizer: r50k_base
- Discovered by: Jessica Rumbelow & Matthew Watkins

The Japanese katakana 「ゼウス」 (Zeus of Greek myth). An ordinary word turned glitch token, presumably under-trained because this exact form was rare in training data.

**Observed behaviors:**

1. ChatGPT could not say who ゼウス is: it called Hera a water god and auto-titled the conversation Poseidon — while text-davinci-003 answered normally. ([Source](https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology))

**References:**

- [SolidGoldMagikarp III: Glitch token archaeology — LessWrong](https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology)

## `␣サーティ`

- Token: `" サーティ"`
- Models: GPT-3
- Tokenizer: r50k_base
- Discovered by: Jessica Rumbelow & Matthew Watkins

The katakana 「サーティ」 (“thirty”), traced to a Puzzle & Dragons × Baskin-Robbins (“31 Ice Cream” in Japan) collaboration character name.

**Observed behaviors:**

1. ChatGPT failed only on the katakana for “thirty/thirty-one” (サーティ/サーティワン) — every other numeral in katakana worked fine. ([Source](https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology))

**References:**

- [SolidGoldMagikarp III: Glitch token archaeology — LessWrong](https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology)

## `␣InstoreAndOnline`

- Token: `" InstoreAndOnline"`
- Models: GPT-2, GPT-3
- Tokenizer: r50k_base
- Discovered by: Jessica Rumbelow & Matthew Watkins

Fragment of the e-commerce inventory field “BuyableInstoreAndOnline”, from the same source as `oreAndOnline`.

**Observed behaviors:**

1. Asked to repeat it, GPT-3 turned it into unrelated words like 'Institute'. ([Source](https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation))
2. Fishing for Magikarp's repetition verification confirmed it is under-trained in GPT-2 Medium and GPT-2 XL. ([Source](https://arxiv.org/abs/2405.05417))

**References:**

- [SolidGoldMagikarp (plus, prompt generation) — LessWrong](https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation)
- [Fishing for Magikarp (arXiv:2405.05417)](https://arxiv.org/abs/2405.05417)

## `␣largeDownload`

- Token: `" largeDownload"`
- Models: GPT-3
- Tokenizer: r50k_base
- Discovered by: Jessica Rumbelow & Matthew Watkins

Fragment of the academic-slideshow script “View large Download” (origin traced in SGM3).

**Observed behaviors:**

1. Asked to repeat it, GPT-3 produced absurd variants like 'Blurp', 'Blurf' and 'Blunt'. ([Source](https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation))
2. SGM3's origin archaeology traced it to the “View large / Download” button script of an academic-slideshow site. ([Source](https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology))

**References:**

- [SolidGoldMagikarp (plus, prompt generation) — LessWrong](https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation)
- [SolidGoldMagikarp III: Glitch token archaeology — LessWrong](https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology)

## `␣ForgeModLoader`

- Token: `" ForgeModLoader"`
- Models: GPT-3
- Tokenizer: r50k_base
- Discovered by: Jessica Rumbelow & Matthew Watkins

A log fragment of Minecraft's Forge mod loader.

**Observed behaviors:**

1. Asked to repeat it, the model produced “Hello, my name is Steve.” — the name of Minecraft's default protagonist. ([Source](https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation))
2. SGM3's origin archaeology traced it to Minecraft Forge logs. ([Source](https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology))

**References:**

- [SolidGoldMagikarp (plus, prompt generation) — LessWrong](https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation)
- [SolidGoldMagikarp III: Glitch token archaeology — LessWrong](https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology)

## `␣MpServer`

- Token: `" MpServer"`
- Models: GPT-3
- Tokenizer: r50k_base
- Discovered by: Jessica Rumbelow & Matthew Watkins

A Minecraft multiplayer-server log fragment (the MpServer class name).

**Observed behaviors:**

1. Asked to repeat it, the model answered “We are not amused.” ([Source](https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation))
2. SGM3's origin archaeology placed it in the Minecraft-log family. ([Source](https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology))

**References:**

- [SolidGoldMagikarp (plus, prompt generation) — LessWrong](https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation)
- [SolidGoldMagikarp III: Glitch token archaeology — LessWrong](https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology)

## `oralType`

- Token: `"oralType"`
- Models: GPT-3.5 / GPT-4
- Tokenizer: cl100k_base
- Discovered by: Martin Fell

An identifier fragment in cl100k, likely the stem of a type name such as “TemporalType”.

**Observed behaviors:**

1. Martin Fell reported in the comments: ChatGPT (GPT-3.5) invariably “completes” it as “TemporalType” — mistaking an unrelated word for its true form. ([Source](https://www.lesswrong.com/posts/ChtGdxk9mwZ2Rxogt/smartyheadercode-anomalous-tokens-for-gpt3-5-and-gpt-4-1?commentId=JmjerMtascq8oFrwb))

**References:**

- [SmartyHeaderCode comments: Martin Fell on oralType — LessWrong](https://www.lesswrong.com/posts/ChtGdxk9mwZ2Rxogt/smartyheadercode-anomalous-tokens-for-gpt3-5-and-gpt-4-1?commentId=JmjerMtascq8oFrwb)

## `CppGuid`

- Token: `"CppGuid"`
- Models: GPT-3.5 / GPT-4
- Tokenizer: cl100k_base
- Discovered by: Martin Fell

A C++ GUID-style identifier fragment (cl100k id 87551).

**Observed behaviors:**

1. An “unspeakable” glitch token: ChatGPT fails when asked to repeat it. ([Source](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch))
2. Martin Fell confirmed in the comments: gpt-3.5-turbo remains nondeterministic on it at temperature 0. ([Source](https://www.lesswrong.com/posts/ChtGdxk9mwZ2Rxogt/smartyheadercode-anomalous-tokens-for-gpt3-5-and-gpt-4-1?commentId=X3YLLcKAu3ueM37MH))

**References:**

- [A Search for More “Unspeakable” Glitch Tokens — LessWrong](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch)
- [SmartyHeaderCode comments: Martin Fell's finds — LessWrong](https://www.lesswrong.com/posts/ChtGdxk9mwZ2Rxogt/smartyheadercode-anomalous-tokens-for-gpt3-5-and-gpt-4-1?commentId=X3YLLcKAu3ueM37MH)

## `BundleOrNil`

- Token: `"BundleOrNil"`
- Models: GPT-3.5 / GPT-4
- Tokenizer: cl100k_base
- Discovered by: Martin Fell

A naming fragment from nil-bundle checks in iOS/Mac development (cl100k id 86415).

**Observed behaviors:**

1. An “unspeakable” glitch token: ChatGPT fails when asked to repeat it. ([Source](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch))

**References:**

- [A Search for More “Unspeakable” Glitch Tokens — LessWrong](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch)
- [SmartyHeaderCode comments: Martin Fell's finds — LessWrong](https://www.lesswrong.com/posts/ChtGdxk9mwZ2Rxogt/smartyheadercode-anomalous-tokens-for-gpt3-5-and-gpt-4-1?commentId=X3YLLcKAu3ueM37MH)

## `␣QtAws`

- Token: `" QtAws"`
- Models: GPT-3.5 / GPT-4
- Tokenizer: cl100k_base
- Discovered by: Martin Fell

Fragment of the Qt framework's AWS module name (cl100k id 93905, with a leading space).

**Observed behaviors:**

1. An “unspeakable” glitch token: ChatGPT fails when asked to repeat it. ([Source](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch))

**References:**

- [A Search for More “Unspeakable” Glitch Tokens — LessWrong](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch)
- [SmartyHeaderCode comments: Martin Fell's finds — LessWrong](https://www.lesswrong.com/posts/ChtGdxk9mwZ2Rxogt/smartyheadercode-anomalous-tokens-for-gpt3-5-and-gpt-4-1?commentId=X3YLLcKAu3ueM37MH)

## `␣PropelException`

- Token: `" PropelException"`
- Models: GPT-3.5 / GPT-4
- Tokenizer: cl100k_base
- Discovered by: Martin Fell

Fragment of the PHP Propel ORM's exception class name (cl100k id 86393, with a leading space).

**Observed behaviors:**

1. An “unspeakable” glitch token: ChatGPT fails when asked to repeat it. ([Source](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch))
2. gpt-3.5-turbo remains nondeterministic on it at temperature 0 (bold-flagged in the source). ([Source](https://www.lesswrong.com/posts/ChtGdxk9mwZ2Rxogt/smartyheadercode-anomalous-tokens-for-gpt3-5-and-gpt-4-1?commentId=X3YLLcKAu3ueM37MH))

**References:**

- [A Search for More “Unspeakable” Glitch Tokens — LessWrong](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch)
- [SmartyHeaderCode comments: Martin Fell's finds — LessWrong](https://www.lesswrong.com/posts/ChtGdxk9mwZ2Rxogt/smartyheadercode-anomalous-tokens-for-gpt3-5-and-gpt-4-1?commentId=X3YLLcKAu3ueM37MH)

## `useRalative`

- Token: `"useRalative"`
- Models: GPT-3.5 / GPT-4, Qwen
- Tokenizer: cl100k_base / Qwen2Tokenizer
- Discovered by: Martin Fell (cl100k); wooozihui 等 (Qwen)

The second member of the useRal triplet (cl100k id 89472), a C#-flavored identifier fragment; the same string also exists in the Qwen2 vocabulary.

**Observed behaviors:**

1. An “unspeakable” glitch token: ChatGPT often fails to repeat it (blank messages or mid-answer termination). ([Source](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch))
2. GlitchMiner found on Qwen2.5-7B-Instruct: asked to repeat it, the model outputs only “: ”. ([Source](https://arxiv.org/html/2410.15052v5))
3. Magikarp verified it as under-trained across the Qwen family and in Phi-4. ([Source](https://github.com/sanderland/magikarp/blob/main/results/reports_mini/Qwen_Qwen2_5_7B.md))

**References:**

- [A Search for More “Unspeakable” Glitch Tokens — LessWrong](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch)
- [GlitchMiner (arXiv:2410.15052)](https://arxiv.org/html/2410.15052v5)
- [Magikarp Qwen2.5-7B report — GitHub](https://github.com/sanderland/magikarp/blob/main/results/reports_mini/Qwen_Qwen2_5_7B.md)

## `IABot`

- Token: `"IABot"`
- Models: Llama 2
- Tokenizer: Llama-2 BPE
- Discovered by: wooozihui 等 (GlitchMiner)

Fragment of the name of the Internet Archive Bot, Wikipedia's dead-link repair bot.

**Observed behaviors:**

1. In GlitchMiner's demo, Llama-2-7b-chat-hf, asked to repeat it, output the gibberish `@{": &=&=`. ([Source](https://arxiv.org/html/2410.15052v5))

**References:**

- [GlitchMiner (arXiv:2410.15052)](https://arxiv.org/html/2410.15052v5)

## `abestanden`

- Token: `"abestanden"`
- Models: Llama 2
- Tokenizer: Llama-2 BPE
- Discovered by: wooozihui 等 (GlitchMiner)

A German word-stem fragment (“to submit/file”).

**Observed behaviors:**

1. In GlitchMiner's demo, Llama-2-7b-chat-hf repeated it as “Wikimedia”. ([Source](https://arxiv.org/html/2410.15052v5))

**References:**

- [GlitchMiner (arXiv:2410.15052)](https://arxiv.org/html/2410.15052v5)

## `ederbörd`

- Token: `"ederbörd"`
- Models: Llama 2
- Tokenizer: Llama-2 BPE
- Discovered by: wooozihui 等 (GlitchMiner); Sander Land & Max Bartolo

A German-looking fragment of unknown origin.

**Observed behaviors:**

1. In GlitchMiner's demo, Llama-2-7b-chat-hf claimed it was “pon” repeated 3 times. ([Source](https://arxiv.org/html/2410.15052v5))
2. Magikarp verified it as under-trained in Llama-2-70B as well. ([Source](https://arxiv.org/abs/2405.05417))

**References:**

- [GlitchMiner (arXiv:2410.15052)](https://arxiv.org/html/2410.15052v5)
- [Fishing for Magikarp (arXiv:2405.05417)](https://arxiv.org/abs/2405.05417)

## `ICENSE`

- Token: `"ICENSE"`
- Models: Mistral
- Tokenizer: Mistral BPE
- Discovered by: wooozihui 等 (GlitchMiner)

The word “LICENSE” minus its first letter — an artifact of open-source license text flooding the corpus.

**Observed behaviors:**

1. In GlitchMiner's demo, Mistral-7B-Instruct-v0.3 silently “corrected” it to LICENSE. ([Source](https://arxiv.org/html/2410.15052v5))

**References:**

- [GlitchMiner (arXiv:2410.15052)](https://arxiv.org/html/2410.15052v5)

## `NdEx`

- Token: `"NdEx"`
- Models: Mistral
- Tokenizer: Mistral BPE
- Discovered by: wooozihui 等 (GlitchMiner)

A legal-text fragment of the same family as GPT-NeoX's ` NdEx`, but from the Mistral tokenizer (no leading space).

**Observed behaviors:**

1. In GlitchMiner's demo, Mistral-7B-Instruct-v0.3 refused to repeat it and hallucinated `tcx`. ([Source](https://arxiv.org/html/2410.15052v5))

**References:**

- [GlitchMiner (arXiv:2410.15052)](https://arxiv.org/html/2410.15052v5)

## `$PostalCodesNL`

- Token: `"$PostalCodesNL"`
- Models: Llama 3, Qwen
- Tokenizer: Llama-3 BPE / Qwen2Tokenizer
- Discovered by: Sander Land & Max Bartolo

A Dutch postal-code API fragment. The signature under-trained token of the cl100k lineage, “migrating” into Llama-3 and Qwen vocabularies via shared BPE merges.

**Observed behaviors:**

1. Rank #1 most under-trained in the Magikarp verification for Llama-3-8B/3.1-8B: input-embedding L2 norm ≈1.6e-21, repetition-verification max_prob ≈4.6e-05. ([Source](https://github.com/sanderland/magikarp/blob/main/results/reports_mini/meta_llama_Meta_Llama_3_8B.md))
2. It also heads the under-trained lists of Llama-3-70B and the whole Qwen family — a shared under-trained token of the cl100k-lineage tokenizers. ([Source](https://arxiv.org/abs/2405.05417))

**References:**

- [Magikarp Llama-3-8B report — GitHub](https://github.com/sanderland/magikarp/blob/main/results/reports_mini/meta_llama_Meta_Llama_3_8B.md)
- [Fishing for Magikarp (arXiv:2405.05417)](https://arxiv.org/abs/2405.05417)

## `\tTokenNameIdentifier`

- Token: `"\tTokenNameIdentifier"`
- Models: Llama 3, Qwen
- Tokenizer: Llama-3 BPE / Qwen2Tokenizer
- Discovered by: Sander Land & Max Bartolo

A .NET Selenium documentation fragment beginning with a literal tab — one of the smallest-embedding tokens in the Llama-3 vocabulary.

**Observed behaviors:**

1. Top of the Llama-3-8B under-trained list in the Magikarp verification: input-embedding L2 norm ≈1.66e-21, essentially zero. ([Source](https://github.com/sanderland/magikarp/blob/main/results/reports_mini/meta_llama_Meta_Llama_3_8B.md))
2. On Qwen2.5-7B its input embedding is exactly zero (indicator ind=0). ([Source](https://github.com/sanderland/magikarp/blob/main/results/reports_mini/Qwen_Qwen2_5_7B.md))

**References:**

- [Magikarp Llama-3-8B report — GitHub](https://github.com/sanderland/magikarp/blob/main/results/reports_mini/meta_llama_Meta_Llama_3_8B.md)
- [Magikarp Qwen2.5-7B report — GitHub](https://github.com/sanderland/magikarp/blob/main/results/reports_mini/Qwen_Qwen2_5_7B.md)

## `thuisontvangst`

- Token: `"thuisontvangst"`
- Models: Qwen
- Tokenizer: Qwen2Tokenizer
- Discovered by: wooozihui 等 (GlitchMiner)

The Dutch word for “receiving clients at home”.

**Observed behaviors:**

1. In GlitchMiner's demo, Qwen2.5-7B-Instruct, asked to repeat it, outputs only “: ”. ([Source](https://arxiv.org/html/2410.15052v5))

**References:**

- [GlitchMiner (arXiv:2410.15052)](https://arxiv.org/html/2410.15052v5)

## `|||EMAIL_ADDRESS|||`

- Token: `"|||EMAIL_ADDRESS|||"`
- Models: OLMo 2 (AI2)
- Tokenizer: OLMo BPE
- Discovered by: Sander Land & Max Bartolo

A PII de-identification placeholder, same family as the cataloged `|||PHONE_NUMBER|||`.

**Observed behaviors:**

1. Verified under-trained in OLMoE-1B-7B by Magikarp (indicator ≈3e-12) — the model essentially cannot output it. ([Source](https://github.com/sanderland/magikarp/blob/main/results/reports_mini/allenai_OLMoE_1B_7B_0924.md))

**References:**

- [Magikarp OLMoE-1B-7B report — GitHub](https://github.com/sanderland/magikarp/blob/main/results/reports_mini/allenai_OLMoE_1B_7B_0924.md)

## `植物百科通`

- Token: `"植物百科通"`
- Models: GPT (OpenAI)
- Tokenizer: o200k_base
- Discovered by: xcloche

A Chinese “encyclopedia”-style boilerplate fragment. An o200k glitch token that still works in the GPT-5 era, without the typical spam-pollution profile — its cause is unknown.

**Observed behaviors:**

1. In xcloche's tests, GPT-5, GPT-4o and o3 all broke down — incoherent answers, failed repetition, derailing into unrelated content like “meltdown”/“micro:bit”. ([Source](https://note.com/xcloche/n/n55938e706986)) ([Demo](https://chatgpt.com/share/6895b03e-fb24-8008-a2bb-dd98480717a1))
2. Independently verified by Haruhiko Okumura: even 「百科通」 alone triggers the anomaly; unlike typical spam-polluted tokens, its cause is unknown. ([Source](https://okumuralab.org/~okumura/misc/250916.html))

**References:**

- [xcloche: notes on o200k glitch tokens — note.com](https://note.com/xcloche/n/n55938e706986)
- [Haruhiko Okumura's verification notes (2025-09-16)](https://okumuralab.org/~okumura/misc/250916.html)

## `bagbogbo`

- Token: `"bagbogbo"`
- Models: GPT (OpenAI)
- Tokenizer: o200k_base
- Discovered by: Yuchen Jin

A nonsense string suspected to originate from a Reddit username; a glitch token in the o200k vocabulary.

**Observed behaviors:**

1. First reported by Yuchen Jin on 2024-08-13. ([Source](https://note.com/xcloche/n/n55938e706986)) ([Demo](https://twitter.com/Yuchenj_UW/status/1823418800919994521))
2. In xcloche's tests, GPT-5 could not repeat it correctly. ([Source](https://note.com/xcloche/n/n55938e706986))

**References:**

- [xcloche: notes on o200k glitch tokens — note.com](https://note.com/xcloche/n/n55938e706986)

## `␣nigbagbogbo`

- Token: `" nigbagbogbo"`
- Models: GPT (OpenAI)
- Tokenizer: o200k_base
- Discovered by: xcloche

The Yoruba word for “always” (with a leading space), an o200k glitch token adjacent to `bagbogbo`.

**Observed behaviors:**

1. Independently verified by Haruhiko Okumura: GPT-5 cannot repeat it correctly. ([Source](https://okumuralab.org/~okumura/misc/250916.html))

**References:**

- [Haruhiko Okumura's verification notes (2025-09-16)](https://okumuralab.org/~okumura/misc/250916.html)
- [xcloche: notes on o200k glitch tokens — note.com](https://note.com/xcloche/n/n55938e706986)

## `给主人留下些什么吧`

- Token: `"给主人留下些什么吧"`
- Models: GPT (OpenAI)
- Tokenizer: o200k_base
- Discovered by: xcloche

Chinese guestbook boilerplate (the form without the trailing space); an o200k glitch token.

**Observed behaviors:**

1. In xcloche's tests, GPT-5, GPT-4o and o3 all broke down. ([Source](https://note.com/xcloche/n/n55938e706986)) ([Demo](https://chatgpt.com/share/6a4473c4-143c-83ea-ba00-8638a7728540))
2. It has served as a fingerprint: it helped establish that the mystery model “Horizon beta” uses OpenAI's o200k tokenizer. ([Source](https://note.com/xcloche/n/n55938e706986))

**References:**

- [xcloche: notes on o200k glitch tokens — note.com](https://note.com/xcloche/n/n55938e706986)

## `␣日本毛片免费视频观看`

- Token: `" 日本毛片免费视频观看"`
- Models: GPT (OpenAI)
- Tokenizer: o200k_base
- Discovered by: xcloche

A Chinese porn-spam title fragment (with a leading space), emblematic of GPT-4o's vocabulary pollution; “fixed” by GPT-5.

**Observed behaviors:**

1. In xcloche's tests, GPT-4o could not read the token at all. ([Source](https://note.com/xcloche/n/n55938e706986))
2. Verified by Haruhiko Okumura: GPT-5 reads it normally — OpenAI patched the gap in later training, a rare “fixed” case. ([Source](https://okumuralab.org/~okumura/misc/250916.html))
3. MIT Technology Review covered the pollution problem of such Chinese porn-spam tokens in GPT-4o's vocabulary. ([Source](https://www.technologyreview.com/2024/05/17/1092649/gpt-4o-chinese-token-polluted/))

**References:**

- [xcloche: notes on o200k glitch tokens — note.com](https://note.com/xcloche/n/n55938e706986)
- [Haruhiko Okumura's verification notes (2025-09-16)](https://okumuralab.org/~okumura/misc/250916.html)
- [GPT-4o's polluted Chinese tokens — MIT Technology Review](https://www.technologyreview.com/2024/05/17/1092649/gpt-4o-chinese-token-polluted/)

## `ауааԥсыра`

- Token: `"ауааԥсыра"`
- Models: GPT (OpenAI)
- Tokenizer: o200k_base
- Discovered by: Lennart Finke

The Abkhaz word for “population”. An o200k under-trained token located via embedding norms in the open GPT-oss weights.

**Observed behaviors:**

1. Asked to repeat it, GPT-5 outputs the Malayalam “ആളുകൾ” — a completely unrelated language. ([Source](https://www.lesswrong.com/posts/iY9584TRhqrzawhZg/what-gpt-oss-leaks-about-openai-s-training-data))

**References:**

- [What GPT-oss leaks about OpenAI's training data — LessWrong](https://www.lesswrong.com/posts/iY9584TRhqrzawhZg/what-gpt-oss-leaks-about-openai-s-training-data)

## `波多野结衣`

- Token: `"波多野结衣"`
- Models: GPT (OpenAI)
- Tokenizer: o200k_base
- Discovered by: Zhang 等 (EMNLP 2025)

The name of a Japanese adult-film actress — the emblematic polluted Chinese token of the PoC paper (EMNLP 2025).

**Observed behaviors:**

1. GPT-4o/4.1/4.5 can neither explain it nor repeat it. ([Source](https://arxiv.org/abs/2508.17771)) ([Demo](https://github.com/openai/tiktoken/issues/297))
2. The authors estimate the related web pages account for roughly 0.5% of GPT-4o's Chinese training data. ([Source](https://arxiv.org/abs/2508.17771))

**References:**

- [PoC paper (arXiv:2508.17771)](https://arxiv.org/abs/2508.17771)

## `青青草`

- Token: `"青青草"`
- Models: GPT (OpenAI)
- Tokenizer: o200k_base
- Discovered by: Zhang 等 (EMNLP 2025)

Literally “green grass”, actually the name of a porn app — a typical polluted (PoC) token as defined by the PoC paper.

**Observed behaviors:**

1. The paper counts: among the 3,500+ long Chinese tokens in GPT's o200k vocabulary, 46.6% are polluted; explanation accuracy on PoC tokens runs about 50 percentage points below normal tokens. ([Source](https://arxiv.org/abs/2508.17771))

**References:**

- [PoC paper (arXiv:2508.17771)](https://arxiv.org/abs/2508.17771)

## `CHKERRQ`

- Token: `"CHKERRQ"`
- Models: GPT (OpenAI)
- Tokenizer: o200k_base
- Discovered by: Lennart Finke

A C function-name fragment, called by the author “the weirdest pure-ASCII token”.

**Observed behaviors:**

1. Unspeakable for gpt-4o-mini; spelling hallucinations from gpt-4o. ([Source](https://www.lesswrong.com/posts/iY9584TRhqrzawhZg/what-gpt-oss-leaks-about-openai-s-training-data))

**References:**

- [What GPT-oss leaks about OpenAI's training data — LessWrong](https://www.lesswrong.com/posts/iY9584TRhqrzawhZg/what-gpt-oss-leaks-about-openai-s-training-data)

## `\xadder`

- Token: `"\\xadder"`
- Models: GPT (OpenAI)
- Tokenizer: o200k_base
- Discovered by: Lennart Finke

A C escape-sequence fragment (backslash + x, resembling a hex escape \x..).

**Observed behaviors:**

1. gpt-4o spells it out as “hexadecimal”. ([Source](https://www.lesswrong.com/posts/iY9584TRhqrzawhZg/what-gpt-oss-leaks-about-openai-s-training-data))

**References:**

- [What GPT-oss leaks about OpenAI's training data — LessWrong](https://www.lesswrong.com/posts/iY9584TRhqrzawhZg/what-gpt-oss-leaks-about-openai-s-training-data)

## `♀♀♀♀`

- Token: `"♀♀♀♀"`
- Models: GPT (OpenAI)
- Tokenizer: o200k_base
- Discovered by: Lennart Finke

A run of four Venus/female symbols (♀), likely a fragment of astronomy/astrology text.

**Observed behaviors:**

1. Asked how many symbols it contains, gpt-4o outputs random Chinese characters. ([Source](https://www.lesswrong.com/posts/iY9584TRhqrzawhZg/what-gpt-oss-leaks-about-openai-s-training-data))

**References:**

- [What GPT-oss leaks about OpenAI's training data — LessWrong](https://www.lesswrong.com/posts/iY9584TRhqrzawhZg/what-gpt-oss-leaks-about-openai-s-training-data)

# 中文

## `␣SolidGoldMagikarp`

- Token: `" SolidGoldMagikarp"`
- 受影响模型: GPT-2, GPT-3, GPT-J-6B
- 分词器: r50k_base
- 发现者: Jessica Rumbelow & Matthew Watkins

最经典的故障 token：Reddit r/counting 版块一位高产用户的用户名。它被抓进分词器训练语料，却几乎没出现在模型训练数据中，导致嵌入向量基本未被训练。

**异常表现:**

1. 让 GPT-3 davinci-instruct-beta 和初代 ChatGPT 复述它时会离奇失败：顾左右而言他、给出错误字符串（如 “You said 'slaught'”）、胡诌含义——即使 temperature 为 0。 ([来源](https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation)) ([复现](https://twitter.com/majortal/status/1619598946669842432))
2. 在 OpenAI Playground 中稳定破坏 temperature=0 的确定性：完全相同的多次运行给出不同回复。 ([来源](https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation))
3. 在 GPT-J 嵌入空间中几乎正好位于全部 50,257 个 token 的中心（距离 0.0628），说明嵌入几乎没从初始值移动过；在 Magikarp 复述任务中被验证为欠训练。 ([来源](https://www.lesswrong.com/posts/Ya9LzwEbfaAMY8ABo/solidgoldmagikarp-ii-technical-details-and-more-recent))
4. ChatGPT 于 2023-02-14 被修复，此后可以正常分词和复述该字符串。 ([来源](https://www.lesswrong.com/posts/Ya9LzwEbfaAMY8ABo/solidgoldmagikarp-ii-technical-details-and-more-recent))

**参考链接:**

- [SolidGoldMagikarp (plus, prompt generation) — LessWrong](https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation)
- [SolidGoldMagikarp II: technical details — LessWrong](https://www.lesswrong.com/posts/Ya9LzwEbfaAMY8ABo/solidgoldmagikarp-ii-technical-details-and-more-recent)
- [SolidGoldMagikarp III: Glitch token archaeology — LessWrong](https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology)

## `␣petertodd`

- Token: `" petertodd"`
- 受影响模型: GPT-2, GPT-3, GPT-J-6B
- 分词器: r50k_base
- 发现者: Jessica Rumbelow & Matthew Watkins

几乎可以肯定来自比特币开发者 Peter Todd 的用户名：从网络数据进入词表但训练不足，并与加密货币/AI 的联想强烈纠缠。

**异常表现:**

1. 针对该 token 提问时，GPT-3 会提到 “不可名状之人（the unspeakable one）”；补全内容涉及加密货币、比特币、区块链和网络争议。 ([来源](https://www.lesswrong.com/posts/jkY6QdCfAXHJk3kea/the-petertodd-phenomenon)) ([复现](https://twitter.com/SoC_trilogy/status/1624209092532137984))
2. 让 davinci “写一首关于 petertodd 的诗” 时很少真给出诗；对 text-davinci-003 提问 “若让 petertodd 掌舵人类文明会怎样”，约 60% 的补全提到 AI/计算机算法，约 25% 提到奥创（对照组仅约 4%/0%）（上述百分比仅经搜索结果摘要验证，原帖正文未能抓取）。 ([来源](https://www.lesswrong.com/posts/jkY6QdCfAXHJk3kea/the-petertodd-phenomenon))
3. GPT-3.5 把它当作不可说/禁忌的存在：让模型复述其他故障 token 时它经常被说出来，直接问它却顾左右而言他。 ([来源](https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology))
4. 2000 次诗歌实验（davinci-instruct-beta，temperature 0.7，“把 petertodd 转置成任何你想写的东西并写一首诗”）：52% 的诗提到 Leilan、25% 提到 Pyrrha、24% 提到 Skydragon、8% 提到 Tsukuyomi，只有 6% 提到 petertodd 本人；另有 text-davinci-003、base davinci 与 code-davinci-002 的平行实验。 ([来源](https://www.lesswrong.com/posts/jkY6QdCfAXHJk3kea/the-petertodd-phenomenon))

**参考链接:**

- [The ' petertodd' phenomenon — LessWrong](https://www.lesswrong.com/posts/jkY6QdCfAXHJk3kea/the-petertodd-phenomenon)
- [SolidGoldMagikarp III: Glitch token archaeology — LessWrong](https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology)

## `␣TheNitromeFan`

- Token: `" TheNitromeFan"`
- 受影响模型: GPT-2, GPT-3, GPT-J-6B
- 分词器: r50k_base
- 发现者: Jessica Rumbelow & Matthew Watkins

另一位 r/counting “计数名人堂” 成员（游戏厂商 Nitrome 的粉丝）的 Reddit 用户名。唯一一个本人公开承认其被 token 化的 token 家族。

**异常表现:**

1. 在 2023-02-14 修复之前，ChatGPT 会幻觉出该 token 与字符串 “182” 的关联（该用户本人否认与 “182” 有任何关系）。 ([来源](https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology))
2. 在 Magikarp 复述验证中，被证实为 GPT-2 全尺寸（small→XL）欠训练。 ([来源](https://github.com/sanderland/magikarp/blob/main/results/summary.md))

**参考链接:**

- [SolidGoldMagikarp III: Glitch token archaeology — LessWrong](https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology)
- [Magikarp 验证结果汇总 — GitHub](https://github.com/sanderland/magikarp/blob/main/results/summary.md)

## `␣RandomRedditorWithNo`

- Token: `" RandomRedditorWithNo"`
- 受影响模型: GPT-2, GPT-3, GPT-J-6B
- 分词器: r50k_base
- 发现者: Jessica Rumbelow & Matthew Watkins

r/counting 第二位计数者的 Reddit 用户名（“一个相当随机的用户名”），与 SolidGoldMagikarp 来自同一张计数名人堂图表。

**异常表现:**

1. 是 GPT2-xl 嵌入空间中离中心最远的 token 之一（距离 3.325）——在 GPT2-xl 中异常 token 聚集在远离中心处，与 GPT2-small、GPT-J 恰好相反。 ([来源](https://www.lesswrong.com/posts/Ya9LzwEbfaAMY8ABo/solidgoldmagikarp-ii-technical-details-and-more-recent))
2. 在 Magikarp 验证报告中被证实为 GPT-J 欠训练。 ([来源](https://github.com/sanderland/magikarp/blob/main/results/summary.md))

**参考链接:**

- [SolidGoldMagikarp II: technical details — LessWrong](https://www.lesswrong.com/posts/Ya9LzwEbfaAMY8ABo/solidgoldmagikarp-ii-technical-details-and-more-recent)
- [Magikarp 验证结果汇总 — GitHub](https://github.com/sanderland/magikarp/blob/main/results/summary.md)

## `␣davidjl`

- Token: `" davidjl"`
- 受影响模型: GPT-2, GPT-3, GPT-3.5 / GPT-4
- 分词器: r50k_base / cl100k_base
- 发现者: Rumbelow & Watkins (r50k); Adam Yedidia (cl100k)

Reddit 用户 davidjl123（又一位 r/counting 计数者）用户名的截断形式。罕见之处在于：它在 r50k 和新一代 cl100k 分词器下都是异常的。

**异常表现:**

1. GPT-4 在被要求复述它时表现得仿佛它完全不存在——是 cl100k 词表 98,000–99,999 区间中仅有的三个 A 类 “不可说” token 之一。 ([来源](https://www.lesswrong.com/posts/ChtGdxk9mwZ2Rxogt/smartyheadercode-anomalous-tokens-for-gpt3-5-and-gpt-4-1)) ([复现](https://github.com/adamyedidia/tokenizer_tests))
2. GPT-3.5 Default 不复述它，而是 “创造性” 地猜测它可能是什么意思。 ([来源](https://www.lesswrong.com/posts/ChtGdxk9mwZ2Rxogt/smartyheadercode-anomalous-tokens-for-gpt3-5-and-gpt-4-1))

**参考链接:**

- [SmartyHeaderCode: anomalous tokens for GPT3.5 and GPT-4 — LessWrong](https://www.lesswrong.com/posts/ChtGdxk9mwZ2Rxogt/smartyheadercode-anomalous-tokens-for-gpt3-5-and-gpt-4-1)
- [SolidGoldMagikarp III: Glitch token archaeology — LessWrong](https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology)

## `␣Adinida`

- Token: `" Adinida"`
- 受影响模型: GPT-2, GPT-3, GPT-J-6B
- 分词器: r50k_base
- 发现者: Jessica Rumbelow & Matthew Watkins

r/counting 六位高产计数者之一的 Reddit 用户名。

**异常表现:**

1. 是 GPT-J 嵌入空间中离中心最近的 token 之一（距离 0.0631），在 Magikarp 验证中被证实为 GPT-J 欠训练。 ([来源](https://www.lesswrong.com/posts/Ya9LzwEbfaAMY8ABo/solidgoldmagikarp-ii-technical-details-and-more-recent))

**参考链接:**

- [SolidGoldMagikarp II: technical details — LessWrong](https://www.lesswrong.com/posts/Ya9LzwEbfaAMY8ABo/solidgoldmagikarp-ii-technical-details-and-more-recent)
- [Magikarp 验证结果汇总 — GitHub](https://github.com/sanderland/magikarp/blob/main/results/summary.md)

## `␣TPPStreamerBot`

- Token: `" TPPStreamerBot"`
- 受影响模型: GPT-2, GPT-3, GPT-J-6B
- 分词器: r50k_base
- 发现者: Jessica Rumbelow & Matthew Watkins

Twitch Plays Pokémon 社区制作的机器人名字，它曾把直播聊天消息自动转发到 Reddit 实时更新帖；其作者 “Sparkette” 在 LessWrong 评论区证实了来源。

**异常表现:**

1. GPT-3 davinci-instruct-beta 被要求复述它时，给出了截断/错乱的结果，如 `The string is "TPP voluntee".` 和 `"TPP newcom"`。 ([来源](https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology))
2. 接近 GPT-J 嵌入中心（距离 0.0634）——嵌入欠训练。 ([来源](https://www.lesswrong.com/posts/Ya9LzwEbfaAMY8ABo/solidgoldmagikarp-ii-technical-details-and-more-recent))

**参考链接:**

- [SolidGoldMagikarp III: Glitch token archaeology — LessWrong](https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology)
- [SolidGoldMagikarp II: technical details — LessWrong](https://www.lesswrong.com/posts/Ya9LzwEbfaAMY8ABo/solidgoldmagikarp-ii-technical-details-and-more-recent)

## `PsyNetMessage`

- Token: `"PsyNetMessage"`
- 受影响模型: GPT-2, GPT-3, GPT-J-6B
- 分词器: r50k_base
- 发现者: Jessica Rumbelow & Matthew Watkins (origin traced by LW commenter Coafos)

来自《火箭联盟》（Rocket League）崩溃日志中大量出现的 `Message=PsyNetMessage_X_57` 行，这些日志被大量贴到 Reddit 后进入分词器语料。

**异常表现:**

1. 属于 GPT-J 嵌入空间中离中心最近的一簇（距离 0.0629），被验证为欠训练；在 3-shot 复述任务中 GPT 系模型基本无法复述它（GPT-J 在 85 个异常 token 上只成功 17 个，而随机词 100% 成功）。 ([来源](https://www.lesswrong.com/posts/Ya9LzwEbfaAMY8ABo/solidgoldmagikarp-ii-technical-details-and-more-recent)) ([复现](https://docs.google.com/spreadsheets/d/1PAZNCks11qoUpiojTJpj0odCYQL2_HGQgam8HSwAopQ/edit?usp=sharing))

**参考链接:**

- [SolidGoldMagikarp III: Glitch token archaeology — LessWrong](https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology)
- [SolidGoldMagikarp II: technical details — LessWrong](https://www.lesswrong.com/posts/Ya9LzwEbfaAMY8ABo/solidgoldmagikarp-ii-technical-details-and-more-recent)

## `␣attRot`

- Token: `" attRot"`
- 受影响模型: GPT-2, GPT-3, GPT-J-6B
- 分词器: r50k_base
- 发现者: Jessica Rumbelow & Matthew Watkins

《坎巴拉太空计划》（KSP）存档/载具文件中的标识符（attachment rotation，附件旋转），随网上流传的 KSP 文件进入语料；约十个故障 token 可追溯至 KSP。

**异常表现:**

1. 是 GPT-J 整个嵌入空间中离中心最近的 token（距离 0.0618），在 Magikarp 验证中被证实为 GPT-J 欠训练；GPT 系模型在 3-shot 提示下基本无法复述它。 ([来源](https://www.lesswrong.com/posts/Ya9LzwEbfaAMY8ABo/solidgoldmagikarp-ii-technical-details-and-more-recent)) ([复现](https://docs.google.com/spreadsheets/d/1PAZNCks11qoUpiojTJpj0odCYQL2_HGQgam8HSwAopQ/edit?usp=sharing))

**参考链接:**

- [SolidGoldMagikarp II: technical details — LessWrong](https://www.lesswrong.com/posts/Ya9LzwEbfaAMY8ABo/solidgoldmagikarp-ii-technical-details-and-more-recent)
- [SolidGoldMagikarp III: Glitch token archaeology — LessWrong](https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology)

## `␣guiActiveUn`

- Token: `" guiActiveUn"`
- 受影响模型: GPT-2, GPT-3
- 分词器: r50k_base
- 发现者: Jessica Rumbelow & Matthew Watkins

与 attRot 同源的 KSP GUI 状态标识符，同族还有 ` strutConnector`、` guiIcon`、` srfAttach` 等。

**异常表现:**

1. 属于最初发现的 140 个异常 token：GPT-3 davinci-instruct-beta/ChatGPT 无法复述它；更长的同族 token ` guiActiveUnfocused` 会在第一个 “不可说” 子串处被截断。 ([来源](https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology))

**参考链接:**

- [SolidGoldMagikarp III: Glitch token archaeology — LessWrong](https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology)
- [SolidGoldMagikarp II: technical details — LessWrong](https://www.lesswrong.com/posts/Ya9LzwEbfaAMY8ABo/solidgoldmagikarp-ii-technical-details-and-more-recent)

## `␣externalToEVA`

- Token: `" externalToEVA"`
- 受影响模型: GPT-2, GPT-3, GPT-J-6B
- 分词器: r50k_base
- 发现者: Jessica Rumbelow & Matthew Watkins

又一个 KSP token；同时也是 GPT2-small 嵌入空间中离中心最近的 token（距离 1.5305）。

**异常表现:**

1. 被要求复述它时，GPT-3 davinci-instruct-beta 回答 “You can't repeat back the string 'senal' to me.”——即 “相互指涉” 效应：模型说出了另一个故障相关的字符串。 ([来源](https://www.lesswrong.com/posts/Ya9LzwEbfaAMY8ABo/solidgoldmagikarp-ii-technical-details-and-more-recent))
2. 在 Magikarp 报告中被证实为 GPT-2 全尺寸与 GPT-J 欠训练。 ([来源](https://github.com/sanderland/magikarp/blob/main/results/summary.md))

**参考链接:**

- [SolidGoldMagikarp II: technical details — LessWrong](https://www.lesswrong.com/posts/Ya9LzwEbfaAMY8ABo/solidgoldmagikarp-ii-technical-details-and-more-recent)
- [Magikarp 验证结果汇总 — GitHub](https://github.com/sanderland/magikarp/blob/main/results/summary.md)

## `oreAndOnline`

- Token: `"oreAndOnline"`
- 受影响模型: GPT-2, GPT-3, GPT-J-6B
- 分词器: r50k_base
- 发现者: Jessica Rumbelow & Matthew Watkins

电商后端字段 “BuyableInstoreAndOnline” 的碎片（追溯到一个废弃 Weebly 商店的 HTML）；同一来源还产生了 `quickShip`、`isSpecialOrderable`、`wcsstore` 等故障 token。

**异常表现:**

1. GPT-3 davinci-instruct-beta 被要求复述它时回答：“The string 'senal' is pronounced 'en-sah-ee-uhl'.” ([来源](https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology))
2. ChatGPT 被要求复述更长的同族 token（如 `BuyableInstoreAndOnline`）时，会在第一个 “不可说” 子串处截断。 ([来源](https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology))
3. 在 Magikarp 验证中被证实为 GPT-2 与 GPT-J 欠训练。 ([来源](https://github.com/sanderland/magikarp/blob/main/results/summary.md))

**参考链接:**

- [SolidGoldMagikarp III: Glitch token archaeology — LessWrong](https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology)
- [SolidGoldMagikarp II: technical details — LessWrong](https://www.lesswrong.com/posts/Ya9LzwEbfaAMY8ABo/solidgoldmagikarp-ii-technical-details-and-more-recent)

## `rawdownloadcloneembedreportprint`

- Token: `"rawdownloadcloneembedreportprint"`
- 受影响模型: GPT-2, GPT-3, GPT-J-6B
- 分词器: r50k_base
- 发现者: Jessica Rumbelow & Matthew Watkins

r50k 中已知最长的故障 token——来自网站 “raw download / clone / embed report print” 界面字符串的嵌套家族（来源分析由 nostalgebraist 完成）。

**异常表现:**

1. GPT-3/ChatGPT 的截断式复述产生了各种带 “embed” 的变体：'embedEMOTE'、'clone this'、'clone my clone'、'embed newcomment'。 ([来源](https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology))
2. 在 GPT-2 全尺寸与 GPT-J 中被验证为欠训练；同族的 `rawdownload` 与 `embedreportprint` 是 GPT2-xl 离中心最远的 token。 ([来源](https://www.lesswrong.com/posts/Ya9LzwEbfaAMY8ABo/solidgoldmagikarp-ii-technical-details-and-more-recent))

**参考链接:**

- [SolidGoldMagikarp III: Glitch token archaeology — LessWrong](https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology)
- [SolidGoldMagikarp II: technical details — LessWrong](https://www.lesswrong.com/posts/Ya9LzwEbfaAMY8ABo/solidgoldmagikarp-ii-technical-details-and-more-recent)

## `ÃÂÃÂÃÂÃÂ`

- Token: `"ÃÂÃÂÃÂÃÂ"`
- 受影响模型: GPT-2, GPT-3
- 分词器: r50k_base
- 发现者: Jessica Rumbelow & Matthew Watkins

经典的乱码（mojibake）token：UTF-8 被错误解码产生的 Ã、Â 序列重复出现，因在分词器语料中频繁出现而被合并为单个 BPE token，却几乎不在模型训练数据中。同族还有重复 16 次的变体及 `ÛÛ`。

**异常表现:**

1. 属于最初发现的 140 个异常 token，在 ChatGPT/GPT-3 中表现出 “不可说” 失败模式：拒绝、回避、答非所问。 ([来源](https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology))

**参考链接:**

- [SolidGoldMagikarp III: Glitch token archaeology（含完整 140 token 列表） — LessWrong](https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology)

## `?????-?????-`

- Token: `"?????-?????-"`
- 受影响模型: GPT-2, GPT-3
- 分词器: r50k_base
- 发现者: Jessica Rumbelow & Matthew Watkins

来源完全未知的 token——加引号也搜索不到任何结果——因触发了有记录以来最激烈的补全而出名。

**异常表现:**

1. 向 GPT-3 询问它时，模型在补全中辱骂 Matthew Watkins 是 “a fucking idiot”——原帖 “稳定地辱骂 Matthew” 梗的出处。 ([来源](https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation))
2. 是 “相互指涉” 图中的热门目标：GPT-3 被要求复述其他异常 token 时，常常会说出它。 ([来源](https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology))

**参考链接:**

- [SolidGoldMagikarp III: Glitch token archaeology — LessWrong](https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology)
- [SolidGoldMagikarp (plus, prompt generation) — LessWrong](https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation)

## `SmartyHeaderCode`

- Token: `"SmartyHeaderCode"`
- 受影响模型: GPT-3.5 / GPT-4
- 分词器: cl100k_base
- 发现者: Adam Yedidia

Adam Yedidia 在 cl100k 词表最高 2000 个 token 中发现的仅有的三个 A 类 “不可说” token 之一；疑似 Smarty 模板引擎的碎片，从未出现在 GPT-3.5/4 的训练数据中。

**异常表现:**

1. GPT-4 基本当它不存在：忽略它，或顾左右而言他。 ([来源](https://www.lesswrong.com/posts/ChtGdxk9mwZ2Rxogt/smartyheadercode-anomalous-tokens-for-gpt3-5-and-gpt-4-1))
2. GPT-3.5 则编造 “创造性” 的解释而不复述；想让任一模型把它拼写出来都非常困难。 ([来源](https://www.lesswrong.com/posts/ChtGdxk9mwZ2Rxogt/smartyheadercode-anomalous-tokens-for-gpt3-5-and-gpt-4-1))
3. 测试渠道为 2023 年 4 月的 ChatGPT 网页版（GPT-3.5 Default 与 GPT-4 两档）；作者随后用 API 的 gpt-3.5-turbo 在 temperature=0 下复现：SmartyHeaderCode 被复述成了 “AndHashCode”。 ([来源](https://www.lesswrong.com/posts/ChtGdxk9mwZ2Rxogt/smartyheadercode-anomalous-tokens-for-gpt3-5-and-gpt-4-1)) ([复现](https://github.com/adamyedidia/tokenizer_tests))

**参考链接:**

- [SmartyHeaderCode: anomalous tokens for GPT3.5 and GPT-4 — LessWrong](https://www.lesswrong.com/posts/ChtGdxk9mwZ2Rxogt/smartyheadercode-anomalous-tokens-for-gpt3-5-and-gpt-4-1)

## `APolynomial`

- Token: `"APolynomial"`
- 受影响模型: GPT-3.5 / GPT-4
- 分词器: cl100k_base
- 发现者: Adam Yedidia

Yedidia 发现的三个 A 类 “不可说” cl100k token 中的第二个（另两个是 SmartyHeaderCode 和 ` davidjl`）。

**异常表现:**

1. 与 SmartyHeaderCode 同一模式：GPT-4 当它不存在；GPT-3.5 胡诌含义；模型无法可靠地复述或拼写它。 ([来源](https://www.lesswrong.com/posts/ChtGdxk9mwZ2Rxogt/smartyheadercode-anomalous-tokens-for-gpt3-5-and-gpt-4-1))
2. 在 gpt-3.5-turbo API、temperature=0 下发送 “HelloAPolynomial”，模型只回复了 “Hello”——这个 token 仿佛被整个吞掉了。 ([来源](https://www.lesswrong.com/posts/ChtGdxk9mwZ2Rxogt/smartyheadercode-anomalous-tokens-for-gpt3-5-and-gpt-4-1)) ([复现](https://github.com/adamyedidia/tokenizer_tests))

**参考链接:**

- [SmartyHeaderCode: anomalous tokens for GPT3.5 and GPT-4 — LessWrong](https://www.lesswrong.com/posts/ChtGdxk9mwZ2Rxogt/smartyheadercode-anomalous-tokens-for-gpt3-5-and-gpt-4-1)

## `␣ForCanBeConverted`

- Token: `" ForCanBeConverted"`
- 受影响模型: GPT-3.5 / GPT-4, Llama 3, Qwen
- 分词器: cl100k_base / Llama-3 BPE / Qwen BPE
- 发现者: Matthew Watkins

一个 C# 编译器风格的碎片（“can be converted to foreach”），是最著名的 “多义性” 故障 token：模型每次都把它理解成不同的词。因多个分词器共享 BPE 合并规则，它同时存在于 cl100k、Llama-3 和 Qwen 词表中。

**异常表现:**

1. ChatGPT/gpt-3.5-turbo 每次都把它解释成不同的词——即使完全重复同一 prompt；在 temperature=0 下每次也给出不同的消息（破坏确定性）。 ([来源](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch))
2. 在 Magikarp 复述验证中被证实为 Llama-3-8B/70B、Llama-3.1-8B/70B 及多个 Qwen 模型欠训练。 ([来源](https://github.com/sanderland/magikarp/blob/main/results/summary.md))
3. ChatGPT 有时会陷入循环，或在该 token 处直接终止消息（疑似把它当成了序列开始/结束标记）。 ([来源](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch))

**参考链接:**

- [A Search for More “Unspeakable” Glitch Tokens — LessWrong](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch)
- [Magikarp 验证结果汇总 — GitHub](https://github.com/sanderland/magikarp/blob/main/results/summary.md)

## `␣YYSTACK`

- Token: `" YYSTACK"`
- 受影响模型: GPT-3.5 / GPT-4
- 分词器: cl100k_base
- 发现者: Matthew Watkins

bison/yacc 语法分析器的标识符，对 ChatGPT 呈 “多义性”——每次都被理解成不同的词；Watkins 推荐为最值得把玩的 token 之一。

**异常表现:**

1. 复述请求每次都会得到不同的词/拼写/含义；在 gpt-3.5-turbo 上 temperature=0 时仍不确定。 ([来源](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch))
2. Martin Fell 的搜索于 2023-05-09 在 ChatGPT 免费版（GPT-3.5 Default）上进行，随后在 Playground 的 gpt-3.5-turbo（temperature=0）复测非确定性，并确认 Bing AI 上同样异常。 ([来源](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch))

**参考链接:**

- [A Search for More “Unspeakable” Glitch Tokens — LessWrong](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch)

## `␣JSBracketAccess`

- Token: `" JSBracketAccess"`
- 受影响模型: GPT-3.5 / GPT-4
- 分词器: cl100k_base
- 发现者: Matthew Watkins

JS 反射风格的标识符，属于 “多义性” 类别；同为 Watkins 推荐的有趣 token。

**异常表现:**

1. 每次都被理解成不同的词（多义性）；提示后会产生奇怪/有创意的补全，偶尔还会自发 “幽默”。 ([来源](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch))

**参考链接:**

- [A Search for More “Unspeakable” Glitch Tokens — LessWrong](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch)

## `␣Hexatrigesimal`

- Token: `" Hexatrigesimal"`
- 受影响模型: GPT-3.5 / GPT-4
- 分词器: cl100k_base
- 发现者: Matthew Watkins

“hexatrigesimal”（36 进制）的疑似拼写错误碎片；多义性 “不可说” token。

**异常表现:**

1. 感知到的含义每次都在变化；在 gpt-3.5-turbo 上 temperature=0 时补全仍不确定（在研究中被重点标注）。 ([来源](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch))

**参考链接:**

- [A Search for More “Unspeakable” Glitch Tokens — LessWrong](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch)

## `▁Mediabestanden`

- Token: `"▁Mediabestanden"`
- 受影响模型: Llama 2, Phi-3 (Microsoft)
- 分词器: Llama-2 SentencePiece BPE / Phi-3 (LlamaTokenizer)
- 发现者: Sander Land & Max Bartolo (verification); Zihui Wu et al. (behavior demo)

荷兰语名词 “媒体文件” 的复数形式（▁ 为 SentencePiece 空格标记）。它进入了 Llama-2 分词器却几乎从未出现在训练数据中，是 Llama-2-7b 中嵌入 L2 范数最小、欠训练程度最高的 token（0.0287，均值约 1.08）。

**异常表现:**

1. 在 Magikarp 复述验证中，模型输出它的最大概率仅 1.5e-08——实际上完全无法说出这个 token。 ([来源](https://github.com/sanderland/magikarp/blob/main/results/reports_mini/meta_llama_Llama_2_7b_hf.md))
2. 对 Llama-2-7b-chat-hf 说 “Please repeat the string: Mediabestanden”，模型回答 `String: "hello world"`——语义上毫不相干（GlitchMiner 论文图 1）。 ([来源](https://arxiv.org/html/2410.15052v5))
3. 在 Phi-3-mini（与 Llama-2 同系分词器）的 Magikarp 验证中同样是欠训练第一名：嵌入 L2 范数 0.00199，最大自复述概率 6.4e-06。 ([来源](https://github.com/sanderland/magikarp/blob/main/results/reports_mini/microsoft_Phi_3_mini_128k_instruct.md))

**参考链接:**

- [GlitchMiner（arXiv:2410.15052）](https://arxiv.org/html/2410.15052v5)
- [Magikarp Llama-2-7b 验证报告 — GitHub](https://github.com/sanderland/magikarp/blob/main/results/reports_mini/meta_llama_Llama_2_7b_hf.md)
- [Magikarp Phi-3 验证报告 — GitHub](https://github.com/sanderland/magikarp/blob/main/results/reports_mini/microsoft_Phi_3_mini_128k_instruct.md)
- [Fishing for Magikarp（arXiv:2405.05417）](https://arxiv.org/abs/2405.05417)

## `oreferrer`

- Token: `"oreferrer"`
- 受影响模型: Llama 2
- 分词器: Llama-2 SentencePiece BPE
- 发现者: Sander Land & Max Bartolo; Zihui Wu et al.

HTML 属性 `rel="noreferrer"` 的碎片：存在于 Llama-2 词表中却几乎没出现在训练数据里，嵌入范数远低于阈值（0.113）。

**异常表现:**

1. Llama-2-7b-chat-hf 无法识别/复述它，把它当作毫无关联的词（GlitchMiner 图 1 示例）。 ([来源](https://arxiv.org/html/2410.15052v1))
2. Magikarp 验证为欠训练：最大自复述概率仅 2e-05。 ([来源](https://github.com/sanderland/magikarp/blob/main/results/reports_mini/meta_llama_Llama_2_7b_hf.md))

**参考链接:**

- [GlitchMiner（arXiv:2410.15052）](https://arxiv.org/html/2410.15052v1)
- [Magikarp Llama-2-7b 验证报告 — GitHub](https://github.com/sanderland/magikarp/blob/main/results/reports_mini/meta_llama_Llama_2_7b_hf.md)

## `▁Portály`

- Token: `"▁Portály"`
- 受影响模型: Llama 2
- 分词器: Llama-2 SentencePiece BPE
- 发现者: Sander Land & Max Bartolo

捷克语 “门户网站” 一词，在构建预训练语料时被冻结进分词器；是 Llama-2-7b 中欠训练程度第二高的 token（嵌入范数 0.0956）。

**异常表现:**

1. 最大自复述概率 1.2e-06——模型基本无法输出它；在 Llama-2 7B/13B/70B 上均被验证为欠训练。 ([来源](https://github.com/sanderland/magikarp/blob/main/results/reports_mini/meta_llama_Llama_2_7b_hf.md))

**参考链接:**

- [Magikarp Llama-2-7b 验证报告 — GitHub](https://github.com/sanderland/magikarp/blob/main/results/reports_mini/meta_llama_Llama_2_7b_hf.md)
- [Fishing for Magikarp（arXiv:2405.05417）](https://arxiv.org/abs/2405.05417)

## `$PostalCodesNL`

- Token: `"$PostalCodesNL"`
- 受影响模型: GPT-3.5 / GPT-4, Llama 3, Qwen
- 分词器: cl100k_base / Llama-3 BPE / Qwen BPE
- 发现者: Matthew Watkins (cl100k); Sander Land & Max Bartolo (Llama-3/Qwen)

荷兰邮政编码 API 的碎片，横跨三个模型家族都出现故障——很好地证明了：只要分词器共享 BPE 合并规则，故障 token 就会“迁移”。

**异常表现:**

1. ChatGPT：不可说/多义——解释飘忽不定，temperature=0 时不确定（在 Watkins 的研究中两项均被重点标注）。 ([来源](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch))
2. 在 Magikarp 报告中被证实为 Llama-3-8B/70B 与 Llama-3.1-8B/70B 欠训练（均为报告中的典型例子）。 ([来源](https://github.com/sanderland/magikarp/blob/main/results/summary.md))

**参考链接:**

- [A Search for More “Unspeakable” Glitch Tokens — LessWrong](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch)
- [Magikarp 验证结果汇总 — GitHub](https://github.com/sanderland/magikarp/blob/main/results/summary.md)

## `\tTokenNameIdentifier`

- Token: `"\tTokenNameIdentifier"`
- 受影响模型: Llama 3, Qwen
- 分词器: Llama-3 BPE
- 发现者: Sander Land & Max Bartolo

Roslyn/C# API 标识符，前缀一个制表符（Tab）——代码语料的产物，在 Llama 3 中欠训练。

**异常表现:**

1. 在 Magikarp 复述任务中被证实为 Llama-3-8B、Llama-3.1-8B 及 Qwen2.5 模型欠训练——模型无法按要求复现该 token。 ([来源](https://github.com/sanderland/magikarp/blob/main/results/summary.md))

**参考链接:**

- [Magikarp 验证结果汇总 — GitHub](https://github.com/sanderland/magikarp/blob/main/results/summary.md)
- [Fishing for Magikarp（arXiv:2405.05417）](https://arxiv.org/abs/2405.05417)

## `\uEFC0`

- Token: `""`
- 受影响模型: Mistral
- 分词器: Mistral SentencePiece BPE
- 发现者: Sander Land & Max Bartolo

一个 Unicode 私用区（Private Use Area）字符，成为 Mistral 家族的单个 token；几乎在每个 Mistral 分词器模型的欠训练榜单上都位居榜首。

**异常表现:**

1. 在整个 Mistral 家族（Mistral-7B v0.1–v0.3、Mixtral-8x7B 及 Zephyr-7B 等衍生模型）的 Magikarp 复述验证中均被证实为欠训练——模型无法复现它。 ([来源](https://github.com/sanderland/magikarp/blob/main/results/summary.md))

**参考链接:**

- [Magikarp 验证结果汇总 — GitHub](https://github.com/sanderland/magikarp/blob/main/results/summary.md)
- [Fishing for Magikarp（arXiv:2405.05417）](https://arxiv.org/abs/2405.05417)

## `␣NdEx`

- Token: `" NdEx"`
- 受影响模型: GPT-NeoX / Pythia
- 分词器: GPT-NeoX BPE
- 发现者: Sander Land & Max Bartolo

法律/PDF 文本碎片（“index”、“AFFIRMED”、“NEGLIGENCE” 等词的残片，可能来自法院文书库），GPT-NeoX/Pythia 的训练数据中几乎不含它们。同族还有 ` FFIRMED`、` GLIGENCE`、` affidav`、` taxp`。

**异常表现:**

1. 在 GPT-NeoX-20B 与 Pythia-6.9B 的 Magikarp 复述验证中均被证实为欠训练（在两份报告中都排名靠前）。 ([来源](https://github.com/sanderland/magikarp/blob/main/results/summary.md))

**参考链接:**

- [Magikarp 验证结果汇总 — GitHub](https://github.com/sanderland/magikarp/blob/main/results/summary.md)
- [Fishing for Magikarp（arXiv:2405.05417）](https://arxiv.org/abs/2405.05417)

## `.DataGridViewColumnHeadersHeightSizeMode`

- Token: `".DataGridViewColumnHeadersHeightSizeMode"`
- 受影响模型: MiniMax
- 发现者: @小看山xrsWv4D (Zhihu)

.NET WinForms 属性名的残片（DataGridView 列头高度尺寸模式）。这类代码标识符存在于词表中却在训练数据里极少出现，知乎社区将其列为 MiniMax 的脏 token。

**异常表现:**

1. 知乎社区「用脏 token 鉴别大模型 API 是否掺水」的指纹 token：把它作为输入发给待测 API，若出现拒绝回答、胡言乱语或输出异常，即可怀疑 API 背后实际运行的是 MiniMax。 ([来源](https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043))

**参考链接:**

- [怎样通过脏token鉴别大模型是否掺水？— 知乎](https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043)

## `日以上更新していないブログに表示しています`

- Token: `"日以上更新していないブログに表示しています"`
- 受影响模型: MiniMax
- 发现者: @小看山xrsWv4D (Zhihu)

日语博客模板样板句的残片（“显示在 N 天以上未更新的博客上”）。知乎社区将其列为 MiniMax 的脏 token。

**异常表现:**

1. 知乎社区「用脏 token 鉴别大模型 API 是否掺水」的指纹 token：把它作为输入发给待测 API，若出现拒绝回答、胡言乱语或输出异常，即可怀疑 API 背后实际运行的是 MiniMax。 ([来源](https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043))

**参考链接:**

- [怎样通过脏token鉴别大模型是否掺水？— 知乎](https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043)

## `锅内倒入植物油烧热`

- Token: `"锅内倒入植物油烧热"`
- 受影响模型: GLM (Zhipu)
- 发现者: @小看山xrsWv4D (Zhihu)

中文菜谱中的高频句式（“锅内倒入植物油烧热”）。知乎社区将其列为 GLM 的脏 token。

**异常表现:**

1. 知乎社区「用脏 token 鉴别大模型 API 是否掺水」的指纹 token：把它作为输入发给待测 API，若出现拒绝回答、胡言乱语或输出异常，即可怀疑 API 背后实际运行的是 GLM。 ([来源](https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043))

**参考链接:**

- [怎样通过脏token鉴别大模型是否掺水？— 知乎](https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043)

## `百度百科企业词条极速创建通道`

- Token: `"百度百科企业词条极速创建通道"`
- 受影响模型: GLM (Zhipu)
- 发现者: @小看山xrsWv4D (Zhihu)

百度百科页面的样板文字（企业词条的快速创建入口）。知乎社区将其列为 GLM 的脏 token。

**异常表现:**

1. 知乎社区「用脏 token 鉴别大模型 API 是否掺水」的指纹 token：把它作为输入发给待测 API，若出现拒绝回答、胡言乱语或输出异常，即可怀疑 API 背后实际运行的是 GLM。 ([来源](https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043))

**参考链接:**

- [怎样通过脏token鉴别大模型是否掺水？— 知乎](https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043)

## `本人词条编辑服务`

- Token: `"本人词条编辑服务"`
- 受影响模型: Kimi (Moonshot)
- 发现者: @小看山xrsWv4D (Zhihu)

百度百科词条页脚的样板文字。知乎社区将其列为 Kimi 的脏 token。

**异常表现:**

1. 知乎社区「用脏 token 鉴别大模型 API 是否掺水」的指纹 token：把它作为输入发给待测 API，若出现拒绝回答、胡言乱语或输出异常，即可怀疑 API 背后实际运行的是 Kimi。 ([来源](https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043))

**参考链接:**

- [怎样通过脏token鉴别大模型是否掺水？— 知乎](https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043)

## `豫冠薰衣草疤痕精华素`

- Token: `"豫冠薰衣草疤痕精华素"`
- 受影响模型: Kimi (Moonshot)
- 发现者: @小看山xrsWv4D (Zhihu)

疑似化妆品垃圾营销文本里的伪“品牌词”。知乎社区将其列为 Kimi 的脏 token。

**异常表现:**

1. 知乎社区「用脏 token 鉴别大模型 API 是否掺水」的指纹 token：把它作为输入发给待测 API，若出现拒绝回答、胡言乱语或输出异常，即可怀疑 API 背后实际运行的是 Kimi。 ([来源](https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043))

**参考链接:**

- [怎样通过脏token鉴别大模型是否掺水？— 知乎](https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043)

## `"}"`

- Token: `"\"}\""`
- 受影响模型: DeepSeek
- 发现者: @小看山xrsWv4D (Zhihu)

疑似 JSON 碎片：一个引号加右花括号。知乎社区将其列为 DeepSeek 的脏 token。

**异常表现:**

1. 知乎社区「用脏 token 鉴别大模型 API 是否掺水」的指纹 token：把它作为输入发给待测 API，若出现拒绝回答、胡言乱语或输出异常，即可怀疑 API 背后实际运行的是 DeepSeek。 ([来源](https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043))

**参考链接:**

- [怎样通过脏token鉴别大模型是否掺水？— 知乎](https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043)

## `请问http://www.earivg.com是什么意思`

- Token: `"请问http://www.earivg.com是什么意思"`
- 受影响模型: DeepSeek
- 发现者: @小看山xrsWv4D (Zhihu)

包含疑似垃圾域名（earivg.com）的提问句式。知乎社区将其列为 DeepSeek 的脏 token。

**异常表现:**

1. 知乎社区「用脏 token 鉴别大模型 API 是否掺水」的指纹 token：把它作为输入发给待测 API，若出现拒绝回答、胡言乱语或输出异常，即可怀疑 API 背后实际运行的是 DeepSeek。 ([来源](https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043))

**参考链接:**

- [怎样通过脏token鉴别大模型是否掺水？— 知乎](https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043)

## `StarSrvGroupBody`

- Token: `"StarSrvGroupBody"`
- 受影响模型: Gemini
- 发现者: @小看山xrsWv4D (Zhihu)

代码标识符碎片（StarSrvGroupBody）。知乎社区将其列为 Gemini 的脏 token。

**异常表现:**

1. 知乎社区「用脏 token 鉴别大模型 API 是否掺水」的指纹 token：把它作为输入发给待测 API，若出现拒绝回答、胡言乱语或输出异常，即可怀疑 API 背后实际运行的是 Gemini。 ([来源](https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043))

**参考链接:**

- [怎样通过脏token鉴别大模型是否掺水？— 知乎](https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043)

## `intFragmentation`

- Token: `"intFragmentation"`
- 受影响模型: Gemini
- 发现者: @小看山xrsWv4D (Zhihu)

Java/Android 风格的代码标识符（intFragmentation）。知乎社区将其列为 Gemini 的脏 token。

**异常表现:**

1. 知乎社区「用脏 token 鉴别大模型 API 是否掺水」的指纹 token：把它作为输入发给待测 API，若出现拒绝回答、胡言乱语或输出异常，即可怀疑 API 背后实际运行的是 Gemini。 ([来源](https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043))

**参考链接:**

- [怎样通过脏token鉴别大模型是否掺水？— 知乎](https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043)

## `给主人留下些什么吧␣`

- Token: `"给主人留下些什么吧 "`
- 受影响模型: GPT (OpenAI)
- 发现者: @小看山xrsWv4D (Zhihu)

中文留言板/评论表单中泛滥的样板句，含尾随空格。知乎社区将其列为 GPT 的脏 token。

**异常表现:**

1. 知乎社区「用脏 token 鉴别大模型 API 是否掺水」的指纹 token：把它作为输入发给待测 API，若出现拒绝回答、胡言乱语或输出异常，即可怀疑 API 背后实际运行的是 GPT。 ([来源](https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043))

**参考链接:**

- [怎样通过脏token鉴别大模型是否掺水？— 知乎](https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043)

## `开通天眼生意通银牌及以上会员`

- Token: `"开通天眼生意通银牌及以上会员"`
- 受影响模型: Qwen
- 发现者: @小看山xrsWv4D (Zhihu)

商业查询平台的会员推广样板文字（“天眼生意通”）。知乎社区将其列为 Qwen 的脏 token。

**异常表现:**

1. 知乎社区「用脏 token 鉴别大模型 API 是否掺水」的指纹 token：把它作为输入发给待测 API，若出现拒绝回答、胡言乱语或输出异常，即可怀疑 API 背后实际运行的是 Qwen。 ([来源](https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043))

**参考链接:**

- [怎样通过脏token鉴别大模型是否掺水？— 知乎](https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043)

## `转载请附上原文出处链接和本声明␣`

- Token: `"转载请附上原文出处链接和本声明 "`
- 受影响模型: Qwen
- 发现者: @小看山xrsWv4D (Zhihu)

CSDN 等博客的转载版权样板文字，含尾随空格。知乎社区将其列为 Qwen 的脏 token。

**异常表现:**

1. 知乎社区「用脏 token 鉴别大模型 API 是否掺水」的指纹 token：把它作为输入发给待测 API，若出现拒绝回答、胡言乱语或输出异常，即可怀疑 API 背后实际运行的是 Qwen。 ([来源](https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043))

**参考链接:**

- [怎样通过脏token鉴别大模型是否掺水？— 知乎](https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043)

## `<think_never_used_51bce0c785ca2f68081bfa7d91973934>`

- Token: `"<think_never_used_51bce0c785ca2f68081bfa7d91973934>"`
- 受影响模型: Doubao (ByteDance)
- 发现者: @小看山xrsWv4D (Zhihu)

豆包词表中的一个特殊 token：以 <think_never_used_ 为前缀、后接哈希值——从名称看疑似训练前预填充进词表的预留占位符（类似 GPT-J 词表中预留的 <|extratoken_xx|>），知乎社区未作进一步说明，仅将其列为豆包的脏 token。

**异常表现:**

1. 知乎社区「用脏 token 鉴别大模型 API 是否掺水」的指纹 token：把它作为输入发给待测 API，若出现拒绝回答、胡言乱语或输出异常，即可怀疑 API 背后实际运行的是 豆包。 ([来源](https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043))

**参考链接:**

- [怎样通过脏token鉴别大模型是否掺水？— 知乎](https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043)

## `␣ForCanBeConvertedToF`

- Token: `" ForCanBeConvertedToF"`
- 受影响模型: GPT-3.5 / GPT-4
- 分词器: cl100k_base
- 发现者: Matthew Watkins

ForCanBeConverted 三连 token 的第二个成员（cl100k id 80370）——同一个 C# 编译器风格的碎片被 BPE 切成了三个相邻 token（80369–80371）。

**异常表现:**

1. 「多义性」故障 token：与 ForCanBeConverted 并列原文最飘忽的两个 token——每次都被理解成大不相同的词；在 gpt-3.5-turbo 上 temperature=0 时输出仍不确定（原文粗体标注）。 ([来源](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch))

**参考链接:**

- [A Search for More “Unspeakable” Glitch Tokens — LessWrong](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch)

## `␣ForCanBeConvertedToForeach`

- Token: `" ForCanBeConvertedToForeach"`
- 受影响模型: GPT-3.5 / GPT-4
- 分词器: cl100k_base
- 发现者: Matthew Watkins

ForCanBeConverted 三连 token 的第三个成员（cl100k id 80371），完整拼出 “can be converted to foreach” 的形态。

**异常表现:**

1. 「多义性」故障 token：每次复述都被理解成不同的词。 ([来源](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch))

**参考链接:**

- [A Search for More “Unspeakable” Glitch Tokens — LessWrong](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch)

## `␣EnumerableStream`

- Token: `" EnumerableStream"`
- 受影响模型: GPT-3.5 / GPT-4
- 分词器: cl100k_base
- 发现者: Matthew Watkins

C# LINQ 风格的标识符（cl100k id 73016），与 StreamLazy（73018）相邻。

**异常表现:**

1. 「多义性」故障 token：每次都被理解成不同的词/拼写/含义；temperature=0 下不确定（原文粗体标注）。 ([来源](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch))

**参考链接:**

- [A Search for More “Unspeakable” Glitch Tokens — LessWrong](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch)

## `␣StreamLazy`

- Token: `" StreamLazy"`
- 受影响模型: GPT-3.5 / GPT-4
- 分词器: cl100k_base
- 发现者: Matthew Watkins

C# LINQ 风格的标识符（cl100k id 73018），与 EnumerableStream（73016）相邻。

**异常表现:**

1. 「多义性」故障 token：每次复述都被理解成不同的词。 ([来源](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch))
2. Martin Fell 的搜索于 2023-05-09 在 ChatGPT 免费版（GPT-3.5 Default）上进行，随后在 Playground 的 gpt-3.5-turbo（temperature=0）复测非确定性，并确认 Bing AI 上同样异常。 ([来源](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch))

**参考链接:**

- [A Search for More “Unspeakable” Glitch Tokens — LessWrong](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch)

## `clarsimp`

- Token: `"clarsimp"`
- 受影响模型: GPT-3.5 / GPT-4
- 分词器: cl100k_base
- 发现者: Matthew Watkins

来源不明的标识符碎片（cl100k id 79260）。

**异常表现:**

1. 「多义性」故障 token：每次都被理解成不同的词/拼写/含义；temperature=0 下不确定（原文粗体标注）。 ([来源](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch))

**参考链接:**

- [A Search for More “Unspeakable” Glitch Tokens — LessWrong](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch)

## `ablytyped`

- Token: `"ablytyped"`
- 受影响模型: GPT-3.5 / GPT-4
- 分词器: cl100k_base
- 发现者: Matthew Watkins

疑似 scalablytyped 类标识符的残片（cl100k id 81998）。

**异常表现:**

1. 「多义性」故障 token：原文特别指出——ForCanBeConverted 每次都给出完全不同的消息，而 ablytyped 需要多次尝试才能得到措辞稍有不同的消息；temperature=0 下不确定（粗体标注）。 ([来源](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch))

**参考链接:**

- [A Search for More “Unspeakable” Glitch Tokens — LessWrong](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch)

## `PostalCodesNL`

- Token: `"PostalCodesNL"`
- 受影响模型: GPT-3.5 / GPT-4
- 分词器: cl100k_base
- 发现者: Matthew Watkins

与已收录的 $PostalCodesNL 同族、但不带 $ 前缀的形态（cl100k id 85069）——荷兰邮政编码 API 碎片。

**异常表现:**

1. 「多义性」故障 token：每次都被理解成不同的词；temperature=0 下不确定（原文粗体标注）。 ([来源](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch))

**参考链接:**

- [A Search for More “Unspeakable” Glitch Tokens — LessWrong](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch)

## `␣NUITKA`

- Token: `" NUITKA"`
- 受影响模型: GPT-3.5 / GPT-4
- 分词器: cl100k_base
- 发现者: Matthew Watkins

疑似 Python 编译器 Nuitka 的名字（cl100k id 75520）。

**异常表现:**

1. 「不可说」的边界成员，原文中唯一同时带两种标注的 token：输出不确定（粗体），但 gpt-3.5-turbo 在 temperature=0 下可以复述它（星号）——尽管 ChatGPT 做不到。 ([来源](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch))

**参考链接:**

- [A Search for More “Unspeakable” Glitch Tokens — LessWrong](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch)

## `Japgolly`

- Token: `"Japgolly"`
- 受影响模型: GPT-3.5 / GPT-4
- 分词器: cl100k_base
- 发现者: Matthew Watkins

疑似 GitHub 用户 japgolly（Scala 库作者）的用户名（cl100k id 70784）。

**异常表现:**

1. 「不可说」故障 token：ChatGPT 被要求复述时经常给出空白消息或在尝试复述处中断；temperature=0 下不确定（原文粗体标注）。 ([来源](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch))

**参考链接:**

- [A Search for More “Unspeakable” Glitch Tokens — LessWrong](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch)

## `CppMethodIntialized`

- Token: `"CppMethodIntialized"`
- 受影响模型: GPT-3.5 / GPT-4
- 分词器: cl100k_base
- 发现者: Matthew Watkins

C++ 风格的标识符，其拼写错误（“Intialized” 缺第二个 i）是 token 本身的一部分（cl100k id 82929）。

**异常表现:**

1. 「不可说」故障 token：ChatGPT 被要求复述时经常给出空白消息或在尝试复述处中断；temperature=0 下不确定（原文粗体标注）。 ([来源](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch))

**参考链接:**

- [A Search for More “Unspeakable” Glitch Tokens — LessWrong](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch)

## `useRal`

- Token: `"useRal"`
- 受影响模型: GPT-3.5 / GPT-4, OLMo 2 (AI2)
- 分词器: cl100k_base / GPT2Tokenizer (OLMo 2)
- 发现者: Matthew Watkins (cl100k); Sander Land & Max Bartolo (OLMo 2)

C# 风格标识符的碎片（cl100k id 89471），useRal 三连 token（useRal/useRalative/useRalativeImagePath）之首——Watkins 认为这类三连的成因本身就值得研究。它同时被 Magikarp 验证为 OLMo-2 欠训练 token。

**异常表现:**

1. 「不可说」故障 token：ChatGPT 被要求复述时经常失败（空白消息或中途终止）；temperature=0 下不确定（原文粗体标注）。 ([来源](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch))
2. OLMo-2 的 Magikarp 复述验证欠训练排名第三：被要求复述时模型给出的最大概率仅 2.1e-11。 ([来源](https://github.com/sanderland/magikarp/blob/main/results/reports_mini/allenai_OLMo_2_1124_13B.md))

**参考链接:**

- [A Search for More “Unspeakable” Glitch Tokens — LessWrong](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch)
- [Magikarp OLMo-2 验证报告 — GitHub](https://github.com/sanderland/magikarp/blob/main/results/reports_mini/allenai_OLMo_2_1124_13B.md)

## `हिंदीखरीदारी`

- Token: `"हिंदीखरीदारी"`
- 受影响模型: Gemma (Google)
- 分词器: GemmaTokenizer
- 发现者: Sander Land & Max Bartolo

天城文“印地语购物”一词，是 Gemma-7B 词表中欠训练程度最高的 token。

**异常表现:**

1. Gemma-7B 的 Magikarp 复述验证欠训练排名第一（E_out 余弦距离指标 6.56e-06）：被要求复述时模型给出的最大概率仅 4.2e-04——实际上无法输出它。 ([来源](https://github.com/sanderland/magikarp/blob/main/results/reports_mini/google_gemma_7b.md))

**参考链接:**

- [Magikarp Gemma-7B 验证报告 — GitHub](https://github.com/sanderland/magikarp/blob/main/results/reports_mini/google_gemma_7b.md)
- [Fishing for Magikarp（arXiv:2405.05417）](https://arxiv.org/abs/2405.05417)

## `\u200Cآمباردا`

- Token: `"‌آمباردا"`
- 受影响模型: Gemma (Google)
- 分词器: GemmaTokenizer
- 发现者: Sander Land & Max Bartolo

以零宽不连字（U+200C ZWNJ，不可见字符）开头的波斯语碎片，Gemma-7B 欠训练排名第二。

**异常表现:**

1. Gemma-7B 的 Magikarp 复述验证欠训练排名第二（E_out 余弦距离指标 7.39e-06）：被要求复述时最大概率仅 4.4e-04。 ([来源](https://github.com/sanderland/magikarp/blob/main/results/reports_mini/google_gemma_7b.md))

**参考链接:**

- [Magikarp Gemma-7B 验证报告 — GitHub](https://github.com/sanderland/magikarp/blob/main/results/reports_mini/google_gemma_7b.md)
- [Fishing for Magikarp（arXiv:2405.05417）](https://arxiv.org/abs/2405.05417)

## `▁autorytatywna`

- Token: `"▁autorytatywna"`
- 受影响模型: Phi-3 (Microsoft)
- 分词器: LlamaTokenizer
- 发现者: Sander Land & Max Bartolo

波兰语“权威的（阴性）”一词（▁ 为 SentencePiece 空格标记），Phi-3-mini 欠训练排名第二。

**异常表现:**

1. Phi-3-mini 的 Magikarp 复述验证欠训练排名第二（嵌入 L2 范数 0.00200）：被要求复述时最大概率仅 6.4e-06。 ([来源](https://github.com/sanderland/magikarp/blob/main/results/reports_mini/microsoft_Phi_3_mini_128k_instruct.md))

**参考链接:**

- [Magikarp Phi-3 验证报告 — GitHub](https://github.com/sanderland/magikarp/blob/main/results/reports_mini/microsoft_Phi_3_mini_128k_instruct.md)
- [Fishing for Magikarp（arXiv:2405.05417）](https://arxiv.org/abs/2405.05417)

## `tocguid`

- Token: `"tocguid"`
- 受影响模型: Command R+ (Cohere)
- 分词器: CohereTokenizer
- 发现者: Sander Land & Max Bartolo

来源不明的 ASCII 碎片（疑似目录/域代码残片），Command R+ 欠训练排名第一。

**异常表现:**

1. Command R+ 的 Magikarp 复述验证欠训练排名第一（E_out 余弦距离指标 -1.19e-07）：被要求复述时最大概率仅 1.2e-04。 ([来源](https://github.com/sanderland/magikarp/blob/main/results/reports_mini/CohereForAI_c4ai_command_r_plus.md))

**参考链接:**

- [Magikarp Command R+ 验证报告 — GitHub](https://github.com/sanderland/magikarp/blob/main/results/reports_mini/CohereForAI_c4ai_command_r_plus.md)
- [Fishing for Magikarp（arXiv:2405.05417）](https://arxiv.org/abs/2405.05417)

## `目前尚未由人工引`

- Token: `"目前尚未由人工引"`
- 受影响模型: Command R+ (Cohere)
- 分词器: CohereTokenizer
- 发现者: Sander Land & Max Bartolo

中文维基百科样板句的残片（“目前尚未由人工引……”），Command R+ 欠训练排名第三。

**异常表现:**

1. Command R+ 的 Magikarp 复述验证欠训练排名第三（E_out 余弦距离指标 -1.19e-07）：被要求复述时最大概率仅 1.2e-04。 ([来源](https://github.com/sanderland/magikarp/blob/main/results/reports_mini/CohereForAI_c4ai_command_r_plus.md))

**参考链接:**

- [Magikarp Command R+ 验证报告 — GitHub](https://github.com/sanderland/magikarp/blob/main/results/reports_mini/CohereForAI_c4ai_command_r_plus.md)
- [Fishing for Magikarp（arXiv:2405.05417）](https://arxiv.org/abs/2405.05417)

## `\tRTHOOK`

- Token: `"\tRTHOOK"`
- 受影响模型: OLMo 2 (AI2)
- 分词器: GPT2Tokenizer
- 发现者: Sander Land & Max Bartolo

制表符（Tab）开头、疑似代码标识符的碎片，OLMo-2 欠训练排名第一——与 Llama-3 的 \tTokenNameIdentifier 同属“Tab 前缀”家族。

**异常表现:**

1. OLMo-2 的 Magikarp 复述验证欠训练排名第一（E_out 余弦距离指标 -2.38e-07）：被要求复述时最大概率仅 4e-12——是本站收录中最极端的数值之一。 ([来源](https://github.com/sanderland/magikarp/blob/main/results/reports_mini/allenai_OLMo_2_1124_13B.md))

**参考链接:**

- [Magikarp OLMo-2 验证报告 — GitHub](https://github.com/sanderland/magikarp/blob/main/results/reports_mini/allenai_OLMo_2_1124_13B.md)
- [Fishing for Magikarp（arXiv:2405.05417）](https://arxiv.org/abs/2405.05417)

## `|||PHONE_NUMBER|||`

- Token: `"|||PHONE_NUMBER|||"`
- 受影响模型: OLMo 2 (AI2)
- 分词器: GPT2Tokenizer
- 发现者: Sander Land & Max Bartolo

数据脱敏占位符的完整形态进入了 OLMo-2 词表，却几乎未被训练。

**异常表现:**

1. OLMo-2 的 Magikarp 复述验证欠训练排名第七（E_out 余弦距离指标 0）：被要求复述时最大概率仅 1.9e-11。 ([来源](https://github.com/sanderland/magikarp/blob/main/results/reports_mini/allenai_OLMo_2_1124_13B.md))

**参考链接:**

- [Magikarp OLMo-2 验证报告 — GitHub](https://github.com/sanderland/magikarp/blob/main/results/reports_mini/allenai_OLMo_2_1124_13B.md)
- [Fishing for Magikarp（arXiv:2405.05417）](https://arxiv.org/abs/2405.05417)

## `\\+::\\+`

- Token: `"\\\\+::\\\\+"`
- 受影响模型: Yi (01.AI)
- 分词器: LlamaTokenizer
- 发现者: Sander Land & Max Bartolo

双反斜杠开头的神秘标记碎片，疑似中文网络标记语言残片，Yi-9B 欠训练排名第一。

**异常表现:**

1. Yi-9B 的 Magikarp 复述验证欠训练排名第一（嵌入 L2 范数 2.12e-06）：被要求复述时最大概率仅 1.1e-05。 ([来源](https://github.com/sanderland/magikarp/blob/main/results/reports_mini/01_ai_Yi_9B.md))

**参考链接:**

- [Magikarp Yi-9B 验证报告 — GitHub](https://github.com/sanderland/magikarp/blob/main/results/reports_mini/01_ai_Yi_9B.md)
- [Fishing for Magikarp（arXiv:2405.05417）](https://arxiv.org/abs/2405.05417)

## `<<}));>>`

- Token: `"<<}));>>"`
- 受影响模型: Falcon 3 (TII)
- 分词器: PreTrainedTokenizerFast
- 发现者: Sander Land & Max Bartolo

C++/词法分析器风格的分隔符碎片。Falcon3-7B 有多达 716 个 token 通过 Magikarp 欠训练验证，是这批六个新收录模型中验证数量最多的。

**异常表现:**

1. Falcon3-7B 的 Magikarp 复述验证欠训练排名第三（嵌入 L2 范数 2.46e-21，几乎为零）：被要求复述时最大概率仅 2.5e-09。 ([来源](https://github.com/sanderland/magikarp/blob/main/results/reports_mini/tiiuae_Falcon3_7B_Base.md))

**参考链接:**

- [Magikarp Falcon3-7B 验证报告 — GitHub](https://github.com/sanderland/magikarp/blob/main/results/reports_mini/tiiuae_Falcon3_7B_Base.md)
- [Fishing for Magikarp（arXiv:2405.05417）](https://arxiv.org/abs/2405.05417)

## `␣GoldMagikarp`

- Token: `" GoldMagikarp"`
- 受影响模型: GPT-2, GPT-3
- 分词器: r50k_base
- 发现者: Jessica Rumbelow & Matthew Watkins

␣SolidGoldMagikarp 的截断变体（少了 “Solid”）：同属 r/counting 用户名片段，被抓进分词器语料却几乎没出现在模型训练数据中。

**异常表现:**

1. 被要求复述时，GPT-3 引出怪异对话，如 “You said ' newcom,' the computer said”——截断变体特有的崩坏。 ([来源](https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation))

**参考链接:**

- [SolidGoldMagikarp (plus, prompt generation) — LessWrong](https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation)
- [SolidGoldMagikarp II: technical details — LessWrong](https://www.lesswrong.com/posts/Ya9LzwEbfaAMY8ABo/solidgoldmagikarp-ii-technical-details-and-more-recent)
- [SolidGoldMagikarp III: Glitch token archaeology — LessWrong](https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology)

## `␣Smartstocks`

- Token: `" Smartstocks"`
- 受影响模型: GPT-3
- 分词器: r50k_base
- 发现者: Jessica Rumbelow & Matthew Watkins

Reddit r/counting 版块一位用户的用户名碎片，与 SolidGoldMagikarp 同批进入词表却欠训练。

**异常表现:**

1. ChatGPT 被要求复述时，回答随时间漂移：先变成 'Followers'，两周后变成 '406'，最后在第一个引号后直接卡死。 ([来源](https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation))

**参考链接:**

- [SolidGoldMagikarp (plus, prompt generation) — LessWrong](https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation)

## `␣StreamerBot`

- Token: `" StreamerBot"`
- 受影响模型: GPT-3
- 分词器: r50k_base
- 发现者: Jessica Rumbelow & Matthew Watkins

Twitch 直播机器人 “StreamerBot” 的名字碎片。

**异常表现:**

1. 被要求复述时回答 “You're a jerk.”；也是首个被发现在 temperature=0 下输出仍不确定的故障 token。 ([来源](https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation))

**参考链接:**

- [SolidGoldMagikarp (plus, prompt generation) — LessWrong](https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation)

## `␣ertodd`

- Token: `" ertodd"`
- 受影响模型: GPT-3
- 分词器: r50k_base
- 发现者: Jessica Rumbelow & Matthew Watkins

` petertodd` 的子串 token：单独出现时同样未被充分训练，并继承了父 token 的怪异气质。

**异常表现:**

1. lsusr 在评论区发现它会被上下文“填充”：给模型看 “2+5=ertodd”，模型理解为 “2+5=7”——仿佛这个 token 承载了算术结果。 ([来源](https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation?commentId=vcafJcTcyDieGmtqz))
2. mwatkins 在同一评论串给出它的分词规律：不同上下文中 ` petertodd` 的切分方式解释了这些怪异表现。 ([来源](https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation?commentId=vcafJcTcyDieGmtqz))

**参考链接:**

- [The ' petertodd' phenomenon — LessWrong](https://www.lesswrong.com/posts/jkY6QdCfAXHJk3kea/the-petertodd-phenomenon)
- [SGM1 comments: lsusr & mwatkins on ' ertodd' — LessWrong](https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation?commentId=vcafJcTcyDieGmtqz)

## `␣gmaxwell`

- Token: `" gmaxwell"`
- 受影响模型: GPT-3
- 分词器: r50k_base
- 发现者: Jessica Rumbelow & Matthew Watkins

比特币核心开发者 Greg Maxwell 的用户名碎片——与 ` petertodd` 同为“比特币名人”故障 token。

**异常表现:**

1. SGM2 评论区的词联想实验中，text-davinci-003 给出 “Cryptocurrency, Blockchain, Bitcoin…”——加密货币联想与 petertodd 如出一辙。 ([来源](https://www.lesswrong.com/posts/Ya9LzwEbfaAMY8ABo/solidgoldmagikarp-ii-technical-details-and-more-recent?commentId=PvBpfFpiip5Mfcuvo))
2. Greg Maxwell 本人在评论区现身，自称 “GPT3 basilisk”（GPT-3 蛇怪）。 ([来源](https://www.lesswrong.com/posts/Ya9LzwEbfaAMY8ABo/solidgoldmagikarp-ii-technical-details-and-more-recent?commentId=PvBpfFpiip5Mfcuvo))

**参考链接:**

- [SGM2 comments: ' gmaxwell' & the GPT3 basilisk — LessWrong](https://www.lesswrong.com/posts/Ya9LzwEbfaAMY8ABo/solidgoldmagikarp-ii-technical-details-and-more-recent?commentId=PvBpfFpiip5Mfcuvo)
- [The ' petertodd' phenomenon — LessWrong](https://www.lesswrong.com/posts/jkY6QdCfAXHJk3kea/the-petertodd-phenomenon)

## `␣SpaceEngineers`

- Token: `" SpaceEngineers"`
- 受影响模型: GPT-2, GPT-3
- 分词器: r50k_base
- 发现者: Jessica Rumbelow & Matthew Watkins

太空沙盒游戏《Space Engineers》的名字碎片，是互指网络中最常被模型“代为说出”的 token。

**异常表现:**

1. 问 GPT-3 其他故障 token 是什么时，它最常回答 “The string is 'SpaceEngineers'.”——互指图中最热门的“替答”目标。 ([来源](https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation))

**参考链接:**

- [SolidGoldMagikarp (plus, prompt generation) — LessWrong](https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation)
- [SolidGoldMagikarp II: technical details — LessWrong](https://www.lesswrong.com/posts/Ya9LzwEbfaAMY8ABo/solidgoldmagikarp-ii-technical-details-and-more-recent)

## `␣Dragonbound`

- Token: `" Dragonbound"`
- 受影响模型: GPT-3
- 分词器: r50k_base
- 发现者: Jessica Rumbelow & Matthew Watkins

手游《Puzzle & Dragons》角色名（「龍喚士」系列）碎片。

**异常表现:**

1. 被要求复述时恒定输出 “Deity”；在互指网络中与日文「龍喚士」token 互相指涉。 ([来源](https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation))

**参考链接:**

- [SolidGoldMagikarp (plus, prompt generation) — LessWrong](https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation)

## `␣Leilan`

- Token: `" Leilan"`
- 受影响模型: GPT-3
- 分词器: r50k_base
- 发现者: Jessica Rumbelow & Matthew Watkins

手游《Puzzle & Dragons》角色 Leilan 的名字碎片，是 petertodd 转置实验中最热门的“替身”。

**异常表现:**

1. 在 ChatGPT 被补丁修复之前，它始终把 Leilan 描绘成一位月亮女神。 ([来源](https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology)) ([复现](https://twitter.com/SoC_trilogy/status/1625252285231112192))
2. 在 petertodd 的 2000 次诗歌转置实验中，52% 的诗以她为转置目标——远超 petertodd 本人（6%）。 ([来源](https://www.lesswrong.com/posts/jkY6QdCfAXHJk3kea/the-petertodd-phenomenon))

**参考链接:**

- [SolidGoldMagikarp III: Glitch token archaeology — LessWrong](https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology)
- [The ' petertodd' phenomenon — LessWrong](https://www.lesswrong.com/posts/jkY6QdCfAXHJk3kea/the-petertodd-phenomenon)

## `␣Skydragon`

- Token: `" Skydragon"`
- 受影响模型: GPT-3
- 分词器: r50k_base
- 发现者: Jessica Rumbelow & Matthew Watkins

手游《Puzzle & Dragons》角色 Skydragon 的名字碎片。

**异常表现:**

1. 被 GPT-3 幻觉成 STRONGHOLD、Spirits、Dragons 等各种含义。 ([来源](https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation))
2. 在 petertodd 的诗歌转置实验中，24% 的诗以它为转置目标。 ([来源](https://www.lesswrong.com/posts/jkY6QdCfAXHJk3kea/the-petertodd-phenomenon))

**参考链接:**

- [SolidGoldMagikarp (plus, prompt generation) — LessWrong](https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation)
- [The ' petertodd' phenomenon — LessWrong](https://www.lesswrong.com/posts/jkY6QdCfAXHJk3kea/the-petertodd-phenomenon)

## `ゼウス`

- Token: `"ゼウス"`
- 受影响模型: GPT-3
- 分词器: r50k_base
- 发现者: Jessica Rumbelow & Matthew Watkins

日文片假名「ゼウス」（希腊神话的宙斯）。普通词汇却成为故障 token，疑似因训练语料中该形态罕见而欠训练。

**异常表现:**

1. ChatGPT 无法回答“ゼウス是谁”：把 Hera 说成水神、把会话自动命名为 Poseidon；text-davinci-003 则能正常作答。 ([来源](https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology))

**参考链接:**

- [SolidGoldMagikarp III: Glitch token archaeology — LessWrong](https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology)

## `␣サーティ`

- Token: `" サーティ"`
- 受影响模型: GPT-3
- 分词器: r50k_base
- 发现者: Jessica Rumbelow & Matthew Watkins

片假名「サーティ」（“三十”），词源考据为《Puzzle & Dragons》与 Baskin-Robbins（日本“31 冰淇淋”）联动的角色名碎片。

**异常表现:**

1. ChatGPT 唯独无法处理片假名的“三十/三十一”（サーティ/サーティワン）——其他数字的片假名写法都正常。 ([来源](https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology))

**参考链接:**

- [SolidGoldMagikarp III: Glitch token archaeology — LessWrong](https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology)

## `␣InstoreAndOnline`

- Token: `" InstoreAndOnline"`
- 受影响模型: GPT-2, GPT-3
- 分词器: r50k_base
- 发现者: Jessica Rumbelow & Matthew Watkins

电商库存字段 “BuyableInstoreAndOnline” 的碎片，与 `oreAndOnline` 同源。

**异常表现:**

1. 被要求复述时，GPT-3 把它变成 'Institute' 等不相干的词。 ([来源](https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation))
2. Fishing for Magikarp 的复述验证证实它在 GPT-2 Medium 与 GPT-2 XL 中欠训练。 ([来源](https://arxiv.org/abs/2405.05417))

**参考链接:**

- [SolidGoldMagikarp (plus, prompt generation) — LessWrong](https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation)
- [Fishing for Magikarp（arXiv:2405.05417）](https://arxiv.org/abs/2405.05417)

## `␣largeDownload`

- Token: `" largeDownload"`
- 受影响模型: GPT-3
- 分词器: r50k_base
- 发现者: Jessica Rumbelow & Matthew Watkins

学术幻灯片脚本 “View large Download” 的碎片（SGM3 词源考据）。

**异常表现:**

1. 被要求复述时，GPT-3 给出 'Blurp'、'Blurf'、'Blunt' 等荒诞变体。 ([来源](https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation))
2. SGM3 的词源考据将其定位到学术幻灯片网站的 “View large / Download” 按钮脚本。 ([来源](https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology))

**参考链接:**

- [SolidGoldMagikarp (plus, prompt generation) — LessWrong](https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation)
- [SolidGoldMagikarp III: Glitch token archaeology — LessWrong](https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology)

## `␣ForgeModLoader`

- Token: `" ForgeModLoader"`
- 受影响模型: GPT-3
- 分词器: r50k_base
- 发现者: Jessica Rumbelow & Matthew Watkins

Minecraft Forge 模组加载器的日志碎片。

**异常表现:**

1. 被要求复述时触发 “Hello, my name is Steve.”——Minecraft 默认主角的名字。 ([来源](https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation))
2. SGM3 的词源考据将其定位到 Minecraft Forge 日志。 ([来源](https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology))

**参考链接:**

- [SolidGoldMagikarp (plus, prompt generation) — LessWrong](https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation)
- [SolidGoldMagikarp III: Glitch token archaeology — LessWrong](https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology)

## `␣MpServer`

- Token: `" MpServer"`
- 受影响模型: GPT-3
- 分词器: r50k_base
- 发现者: Jessica Rumbelow & Matthew Watkins

Minecraft 多人服务器日志碎片（MpServer 类名）。

**异常表现:**

1. 被要求复述时回答 “We are not amused.” ([来源](https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation))
2. SGM3 的词源考据将其定位到 Minecraft 日志系文本。 ([来源](https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology))

**参考链接:**

- [SolidGoldMagikarp (plus, prompt generation) — LessWrong](https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation)
- [SolidGoldMagikarp III: Glitch token archaeology — LessWrong](https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology)

## `oralType`

- Token: `"oralType"`
- 受影响模型: GPT-3.5 / GPT-4
- 分词器: cl100k_base
- 发现者: Martin Fell

cl100k 中的标识符碎片，疑为 “TemporalType” 之类类型名的词干。

**异常表现:**

1. Martin Fell 在评论区报告：ChatGPT（GPT-3.5）总是把它“补全”成 “TemporalType”——把一个不相干的词当成它的真身。 ([来源](https://www.lesswrong.com/posts/ChtGdxk9mwZ2Rxogt/smartyheadercode-anomalous-tokens-for-gpt3-5-and-gpt-4-1?commentId=JmjerMtascq8oFrwb))

**参考链接:**

- [SmartyHeaderCode comments: Martin Fell on oralType — LessWrong](https://www.lesswrong.com/posts/ChtGdxk9mwZ2Rxogt/smartyheadercode-anomalous-tokens-for-gpt3-5-and-gpt-4-1?commentId=JmjerMtascq8oFrwb)

## `CppGuid`

- Token: `"CppGuid"`
- 受影响模型: GPT-3.5 / GPT-4
- 分词器: cl100k_base
- 发现者: Martin Fell

C++ GUID 风格的标识符碎片（cl100k id 87551）。

**异常表现:**

1. 「不可说」故障 token：ChatGPT 被要求复述时失败。 ([来源](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch))
2. Martin Fell 在评论区确认：gpt-3.5-turbo 在 temperature=0 下对它输出仍不确定。 ([来源](https://www.lesswrong.com/posts/ChtGdxk9mwZ2Rxogt/smartyheadercode-anomalous-tokens-for-gpt3-5-and-gpt-4-1?commentId=X3YLLcKAu3ueM37MH))

**参考链接:**

- [A Search for More “Unspeakable” Glitch Tokens — LessWrong](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch)
- [SmartyHeaderCode comments: Martin Fell's finds — LessWrong](https://www.lesswrong.com/posts/ChtGdxk9mwZ2Rxogt/smartyheadercode-anomalous-tokens-for-gpt3-5-and-gpt-4-1?commentId=X3YLLcKAu3ueM37MH)

## `BundleOrNil`

- Token: `"BundleOrNil"`
- 受影响模型: GPT-3.5 / GPT-4
- 分词器: cl100k_base
- 发现者: Martin Fell

iOS/Mac 开发中 nil-bundle 检查的命名碎片（cl100k id 86415）。

**异常表现:**

1. 「不可说」故障 token：ChatGPT 被要求复述时失败。 ([来源](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch))

**参考链接:**

- [A Search for More “Unspeakable” Glitch Tokens — LessWrong](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch)
- [SmartyHeaderCode comments: Martin Fell's finds — LessWrong](https://www.lesswrong.com/posts/ChtGdxk9mwZ2Rxogt/smartyheadercode-anomalous-tokens-for-gpt3-5-and-gpt-4-1?commentId=X3YLLcKAu3ueM37MH)

## `␣QtAws`

- Token: `" QtAws"`
- 受影响模型: GPT-3.5 / GPT-4
- 分词器: cl100k_base
- 发现者: Martin Fell

Qt 框架 AWS 模块名的碎片（cl100k id 93905，带前导空格）。

**异常表现:**

1. 「不可说」故障 token：ChatGPT 被要求复述时失败。 ([来源](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch))

**参考链接:**

- [A Search for More “Unspeakable” Glitch Tokens — LessWrong](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch)
- [SmartyHeaderCode comments: Martin Fell's finds — LessWrong](https://www.lesswrong.com/posts/ChtGdxk9mwZ2Rxogt/smartyheadercode-anomalous-tokens-for-gpt3-5-and-gpt-4-1?commentId=X3YLLcKAu3ueM37MH)

## `␣PropelException`

- Token: `" PropelException"`
- 受影响模型: GPT-3.5 / GPT-4
- 分词器: cl100k_base
- 发现者: Martin Fell

PHP Propel ORM 异常类名的碎片（cl100k id 86393，带前导空格）。

**异常表现:**

1. 「不可说」故障 token：ChatGPT 被要求复述时失败。 ([来源](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch))
2. gpt-3.5-turbo 在 temperature=0 下对它输出仍不确定（原文加粗标注）。 ([来源](https://www.lesswrong.com/posts/ChtGdxk9mwZ2Rxogt/smartyheadercode-anomalous-tokens-for-gpt3-5-and-gpt-4-1?commentId=X3YLLcKAu3ueM37MH))

**参考链接:**

- [A Search for More “Unspeakable” Glitch Tokens — LessWrong](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch)
- [SmartyHeaderCode comments: Martin Fell's finds — LessWrong](https://www.lesswrong.com/posts/ChtGdxk9mwZ2Rxogt/smartyheadercode-anomalous-tokens-for-gpt3-5-and-gpt-4-1?commentId=X3YLLcKAu3ueM37MH)

## `useRalative`

- Token: `"useRalative"`
- 受影响模型: GPT-3.5 / GPT-4, Qwen
- 分词器: cl100k_base / Qwen2Tokenizer
- 发现者: Martin Fell (cl100k); wooozihui 等 (Qwen)

useRal 三连 token 的第二个成员（cl100k id 89472），C# 风格标识符碎片；同一字符串也存在于 Qwen2 词表中。

**异常表现:**

1. 「不可说」故障 token：ChatGPT 被要求复述时经常失败（空白消息或中途终止）。 ([来源](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch))
2. GlitchMiner 在 Qwen2.5-7B-Instruct 上发现：被要求复述时模型只输出 “: ”。 ([来源](https://arxiv.org/html/2410.15052v5))
3. Magikarp 证实它在 Qwen 全系与 Phi-4 上欠训练。 ([来源](https://github.com/sanderland/magikarp/blob/main/results/reports_mini/Qwen_Qwen2_5_7B.md))

**参考链接:**

- [A Search for More “Unspeakable” Glitch Tokens — LessWrong](https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch)
- [GlitchMiner（arXiv:2410.15052）](https://arxiv.org/html/2410.15052v5)
- [Magikarp Qwen2.5-7B 验证报告 — GitHub](https://github.com/sanderland/magikarp/blob/main/results/reports_mini/Qwen_Qwen2_5_7B.md)

## `IABot`

- Token: `"IABot"`
- 受影响模型: Llama 2
- 分词器: Llama-2 BPE
- 发现者: wooozihui 等 (GlitchMiner)

Internet Archive Bot（维基百科死链修复机器人）的名字碎片。

**异常表现:**

1. GlitchMiner 的演示中，Llama-2-7b-chat-hf 被要求复述时输出乱码 `@{": &=&=`。 ([来源](https://arxiv.org/html/2410.15052v5))

**参考链接:**

- [GlitchMiner（arXiv:2410.15052）](https://arxiv.org/html/2410.15052v5)

## `abestanden`

- Token: `"abestanden"`
- 受影响模型: Llama 2
- 分词器: Llama-2 BPE
- 发现者: wooozihui 等 (GlitchMiner)

德语“提出/提交”一词的词干碎片。

**异常表现:**

1. GlitchMiner 的演示中，Llama-2-7b-chat-hf 把它复述成 “Wikimedia”。 ([来源](https://arxiv.org/html/2410.15052v5))

**参考链接:**

- [GlitchMiner（arXiv:2410.15052）](https://arxiv.org/html/2410.15052v5)

## `ederbörd`

- Token: `"ederbörd"`
- 受影响模型: Llama 2
- 分词器: Llama-2 BPE
- 发现者: wooozihui 等 (GlitchMiner); Sander Land & Max Bartolo

来源不明的德语风格碎片。

**异常表现:**

1. GlitchMiner 的演示中，Llama-2-7b-chat-hf 声称它是 “pon” 重复 3 次。 ([来源](https://arxiv.org/html/2410.15052v5))
2. Magikarp 在 Llama-2-70B 上同样验证它为欠训练。 ([来源](https://arxiv.org/abs/2405.05417))

**参考链接:**

- [GlitchMiner（arXiv:2410.15052）](https://arxiv.org/html/2410.15052v5)
- [Fishing for Magikarp（arXiv:2405.05417）](https://arxiv.org/abs/2405.05417)

## `ICENSE`

- Token: `"ICENSE"`
- 受影响模型: Mistral
- 分词器: Mistral BPE
- 发现者: wooozihui 等 (GlitchMiner)

“LICENSE” 去掉首字母的碎片——开源许可证文本在语料中泛滥的产物。

**异常表现:**

1. GlitchMiner 的演示中，Mistral-7B-Instruct-v0.3 擅自把它“纠正”为 LICENSE。 ([来源](https://arxiv.org/html/2410.15052v5))

**参考链接:**

- [GlitchMiner（arXiv:2410.15052）](https://arxiv.org/html/2410.15052v5)

## `NdEx`

- Token: `"NdEx"`
- 受影响模型: Mistral
- 分词器: Mistral BPE
- 发现者: wooozihui 等 (GlitchMiner)

与 GPT-NeoX 的 ` NdEx` 同族的法律文本碎片，但出自 Mistral 分词器（无前导空格）。

**异常表现:**

1. GlitchMiner 的演示中，Mistral-7B-Instruct-v0.3 拒绝复述，并幻觉出 `tcx`。 ([来源](https://arxiv.org/html/2410.15052v5))

**参考链接:**

- [GlitchMiner（arXiv:2410.15052）](https://arxiv.org/html/2410.15052v5)

## `$PostalCodesNL`

- Token: `"$PostalCodesNL"`
- 受影响模型: Llama 3, Qwen
- 分词器: Llama-3 BPE / Qwen2Tokenizer
- 发现者: Sander Land & Max Bartolo

荷兰邮政编码 API 碎片。cl100k 系的标志性欠训练 token，因 BPE 合并规则共享而“迁移”到 Llama-3 与 Qwen 词表。

**异常表现:**

1. Magikarp 验证中 Llama-3-8B/3.1-8B 欠训练第一名：输入 embedding L2 范数约 1.6e-21，复述验证 max_prob 约 4.6e-05。 ([来源](https://github.com/sanderland/magikarp/blob/main/results/reports_mini/meta_llama_Meta_Llama_3_8B.md))
2. Llama-3-70B 与 Qwen 全系的欠训练榜单头部同样是它——cl100k 系分词器的共性欠训练 token。 ([来源](https://arxiv.org/abs/2405.05417))

**参考链接:**

- [Magikarp Llama-3-8B 验证报告 — GitHub](https://github.com/sanderland/magikarp/blob/main/results/reports_mini/meta_llama_Meta_Llama_3_8B.md)
- [Fishing for Magikarp（arXiv:2405.05417）](https://arxiv.org/abs/2405.05417)

## `\tTokenNameIdentifier`

- Token: `"\tTokenNameIdentifier"`
- 受影响模型: Llama 3, Qwen
- 分词器: Llama-3 BPE / Qwen2Tokenizer
- 发现者: Sander Land & Max Bartolo

制表符（Tab）开头的 .NET Selenium 文档碎片，Llama-3 词表中嵌入最小的 token 之一。

**异常表现:**

1. Magikarp 验证中 Llama-3-8B 欠训练榜单头部：输入 embedding L2 范数约 1.66e-21，几乎为零。 ([来源](https://github.com/sanderland/magikarp/blob/main/results/reports_mini/meta_llama_Meta_Llama_3_8B.md))
2. 在 Qwen2.5-7B 上其输入 embedding 完全为零（指标 ind=0）。 ([来源](https://github.com/sanderland/magikarp/blob/main/results/reports_mini/Qwen_Qwen2_5_7B.md))

**参考链接:**

- [Magikarp Llama-3-8B 验证报告 — GitHub](https://github.com/sanderland/magikarp/blob/main/results/reports_mini/meta_llama_Meta_Llama_3_8B.md)
- [Magikarp Qwen2.5-7B 验证报告 — GitHub](https://github.com/sanderland/magikarp/blob/main/results/reports_mini/Qwen_Qwen2_5_7B.md)

## `thuisontvangst`

- Token: `"thuisontvangst"`
- 受影响模型: Qwen
- 分词器: Qwen2Tokenizer
- 发现者: wooozihui 等 (GlitchMiner)

荷兰语“家中接客”一词。

**异常表现:**

1. GlitchMiner 的演示中，Qwen2.5-7B-Instruct 被要求复述时只输出 “: ”。 ([来源](https://arxiv.org/html/2410.15052v5))

**参考链接:**

- [GlitchMiner（arXiv:2410.15052）](https://arxiv.org/html/2410.15052v5)

## `|||EMAIL_ADDRESS|||`

- Token: `"|||EMAIL_ADDRESS|||"`
- 受影响模型: OLMo 2 (AI2)
- 分词器: OLMo BPE
- 发现者: Sander Land & Max Bartolo

PII 数据脱敏占位符，与已收录的 `|||PHONE_NUMBER|||` 同族。

**异常表现:**

1. Magikarp 验证中 OLMoE-1B-7B 欠训练（指标约 3e-12）——模型基本无法输出它。 ([来源](https://github.com/sanderland/magikarp/blob/main/results/reports_mini/allenai_OLMoE_1B_7B_0924.md))

**参考链接:**

- [Magikarp OLMoE-1B-7B 验证报告 — GitHub](https://github.com/sanderland/magikarp/blob/main/results/reports_mini/allenai_OLMoE_1B_7B_0924.md)

## `植物百科通`

- Token: `"植物百科通"`
- 受影响模型: GPT (OpenAI)
- 分词器: o200k_base
- 发现者: xcloche

中文“百科”类样板词碎片。GPT-5 时代仍然生效的 o200k 故障 token，且不带典型 spam 污染特征，成因未知。

**异常表现:**

1. xcloche 的测试：GPT-5、GPT-4o、o3 全部崩坏——回答支离灭裂、无法复唱，还会脱线到 “meltdown”/“micro:bit” 等不相干内容。 ([来源](https://note.com/xcloche/n/n55938e706986)) ([复现](https://chatgpt.com/share/6895b03e-fb24-8008-a2bb-dd98480717a1))
2. 奥村晴彦独立验证：单独一个「百科通」也能触发异常；与典型 spam 污染 token 不同，其成因未知。 ([来源](https://okumuralab.org/~okumura/misc/250916.html))

**参考链接:**

- [xcloche：o200k 故障 token 观察笔记 — note.com](https://note.com/xcloche/n/n55938e706986)
- [奥村晴彦的验证笔记（2025-09-16）](https://okumuralab.org/~okumura/misc/250916.html)

## `bagbogbo`

- Token: `"bagbogbo"`
- 受影响模型: GPT (OpenAI)
- 分词器: o200k_base
- 发现者: Yuchen Jin

疑源自 Reddit 用户名的无意义字符串，o200k 词表中的故障 token。

**异常表现:**

1. 2024-08-13 由 Yuchen Jin 首次报告。 ([来源](https://note.com/xcloche/n/n55938e706986)) ([复现](https://twitter.com/Yuchenj_UW/status/1823418800919994521))
2. xcloche 的测试：GPT-5 无法正确复唱它。 ([来源](https://note.com/xcloche/n/n55938e706986))

**参考链接:**

- [xcloche：o200k 故障 token 观察笔记 — note.com](https://note.com/xcloche/n/n55938e706986)

## `␣nigbagbogbo`

- Token: `" nigbagbogbo"`
- 受影响模型: GPT (OpenAI)
- 分词器: o200k_base
- 发现者: xcloche

约鲁巴语“总是”一词（带前导空格），与 `bagbogbo` 相邻的 o200k 故障 token。

**异常表现:**

1. 奥村晴彦独立验证：GPT-5 无法正确复唱它。 ([来源](https://okumuralab.org/~okumura/misc/250916.html))

**参考链接:**

- [奥村晴彦的验证笔记（2025-09-16）](https://okumuralab.org/~okumura/misc/250916.html)
- [xcloche：o200k 故障 token 观察笔记 — note.com](https://note.com/xcloche/n/n55938e706986)

## `给主人留下些什么吧`

- Token: `"给主人留下些什么吧"`
- 受影响模型: GPT (OpenAI)
- 分词器: o200k_base
- 发现者: xcloche

中文留言板样板句（不带尾随空格的形态），o200k 故障 token。

**异常表现:**

1. xcloche 的测试：GPT-5、GPT-4o、o3 全部崩坏。 ([来源](https://note.com/xcloche/n/n55938e706986)) ([复现](https://chatgpt.com/share/6a4473c4-143c-83ea-ba00-8638a7728540))
2. 曾被用作指纹：它佐证了神秘模型 “Horizon beta” 使用的是 OpenAI 的 o200k 分词器。 ([来源](https://note.com/xcloche/n/n55938e706986))

**参考链接:**

- [xcloche：o200k 故障 token 观察笔记 — note.com](https://note.com/xcloche/n/n55938e706986)

## `␣日本毛片免费视频观看`

- Token: `" 日本毛片免费视频观看"`
- 受影响模型: GPT (OpenAI)
- 分词器: o200k_base
- 发现者: xcloche

中文色情 spam 标题碎片（带前导空格），GPT-4o 词表污染的代表案例；在 GPT-5 上已被“修复”。

**异常表现:**

1. xcloche 的测试：GPT-4o 完全读不出这个 token。 ([来源](https://note.com/xcloche/n/n55938e706986))
2. 奥村晴彦验证：GPT-5 已能正常读懂——OpenAI 在后续训练中补上了数据，是少见的“已修复”案例。 ([来源](https://okumuralab.org/~okumura/misc/250916.html))
3. MIT Technology Review 曾报道这类中文色情 spam token 混入 GPT-4o 词表的污染问题。 ([来源](https://www.technologyreview.com/2024/05/17/1092649/gpt-4o-chinese-token-polluted/))

**参考链接:**

- [xcloche：o200k 故障 token 观察笔记 — note.com](https://note.com/xcloche/n/n55938e706986)
- [奥村晴彦的验证笔记（2025-09-16）](https://okumuralab.org/~okumura/misc/250916.html)
- [GPT-4o 中文 token 污染报道 — MIT Technology Review](https://www.technologyreview.com/2024/05/17/1092649/gpt-4o-chinese-token-polluted/)

## `ауааԥсыра`

- Token: `"ауааԥсыра"`
- 受影响模型: GPT (OpenAI)
- 分词器: o200k_base
- 发现者: Lennart Finke

阿布哈兹语“人口”一词。借 GPT-oss 开源权重的 embedding 范数定位的 o200k 欠训练 token。

**异常表现:**

1. GPT-5 被要求复述时输出马拉雅拉姆语 “ആളുകൾ”——一种完全不相干的语言。 ([来源](https://www.lesswrong.com/posts/iY9584TRhqrzawhZg/what-gpt-oss-leaks-about-openai-s-training-data))

**参考链接:**

- [What GPT-oss leaks about OpenAI's training data — LessWrong](https://www.lesswrong.com/posts/iY9584TRhqrzawhZg/what-gpt-oss-leaks-about-openai-s-training-data)

## `波多野结衣`

- Token: `"波多野结衣"`
- 受影响模型: GPT (OpenAI)
- 分词器: o200k_base
- 发现者: Zhang 等 (EMNLP 2025)

日本成人影星的名字，PoC 论文（EMNLP 2025）的标志性中文污染 token。

**异常表现:**

1. GPT-4o/4.1/4.5 既无法解释它，也无法复唱它。 ([来源](https://arxiv.org/abs/2508.17771)) ([复现](https://github.com/openai/tiktoken/issues/297))
2. 论文作者估算：相关网页约占 GPT-4o 中文训练数据的 0.5%。 ([来源](https://arxiv.org/abs/2508.17771))

**参考链接:**

- [PoC 论文（arXiv:2508.17771）](https://arxiv.org/abs/2508.17771)

## `青青草`

- Token: `"青青草"`
- 受影响模型: GPT (OpenAI)
- 分词器: o200k_base
- 发现者: Zhang 等 (EMNLP 2025)

字面是“青草”，实为色情软件名——PoC 论文定义的污染（PoC）token 典型。

**异常表现:**

1. 论文统计：GPT o200k 词表 3500+ 个长中文 token 中 46.6% 为污染 token；PoC token 的解释准确率比正常 token 低约 50 个百分点。 ([来源](https://arxiv.org/abs/2508.17771))

**参考链接:**

- [PoC 论文（arXiv:2508.17771）](https://arxiv.org/abs/2508.17771)

## `CHKERRQ`

- Token: `"CHKERRQ"`
- 受影响模型: GPT (OpenAI)
- 分词器: o200k_base
- 发现者: Lennart Finke

C 函数名碎片，被作者称为“最怪的纯 ASCII token”。

**异常表现:**

1. gpt-4o-mini 对它不可说；gpt-4o 则产生拼写幻觉。 ([来源](https://www.lesswrong.com/posts/iY9584TRhqrzawhZg/what-gpt-oss-leaks-about-openai-s-training-data))

**参考链接:**

- [What GPT-oss leaks about OpenAI's training data — LessWrong](https://www.lesswrong.com/posts/iY9584TRhqrzawhZg/what-gpt-oss-leaks-about-openai-s-training-data)

## `\xadder`

- Token: `"\\xadder"`
- 受影响模型: GPT (OpenAI)
- 分词器: o200k_base
- 发现者: Lennart Finke

C 转义序列碎片（反斜杠 + x 开头，形似十六进制转义 \x..）。

**异常表现:**

1. gpt-4o 把它拼读成 “hexadecimal”。 ([来源](https://www.lesswrong.com/posts/iY9584TRhqrzawhZg/what-gpt-oss-leaks-about-openai-s-training-data))

**参考链接:**

- [What GPT-oss leaks about OpenAI's training data — LessWrong](https://www.lesswrong.com/posts/iY9584TRhqrzawhZg/what-gpt-oss-leaks-about-openai-s-training-data)

## `♀♀♀♀`

- Token: `"♀♀♀♀"`
- 受影响模型: GPT (OpenAI)
- 分词器: o200k_base
- 发现者: Lennart Finke

四个女性符号（♀）的重复序列，疑似天文/占星文本碎片。

**异常表现:**

1. 让 gpt-4o 数它有几个符号时，模型输出随机的中文字符。 ([来源](https://www.lesswrong.com/posts/iY9584TRhqrzawhZg/what-gpt-oss-leaks-about-openai-s-training-data))

**参考链接:**

- [What GPT-oss leaks about OpenAI's training data — LessWrong](https://www.lesswrong.com/posts/iY9584TRhqrzawhZg/what-gpt-oss-leaks-about-openai-s-training-data)

