{"site":"https://glitch-token.jjc.fun","generated":{"tokens":103,"models":22},"models":[{"id":"gpt-2","name":"GPT-2"},{"id":"gpt-3","name":"GPT-3"},{"id":"gpt-3.5-4","name":"GPT-3.5 / GPT-4"},{"id":"gpt-j","name":"GPT-J-6B"},{"id":"gpt-neox","name":"GPT-NeoX / Pythia"},{"id":"llama-2","name":"Llama 2"},{"id":"llama-3","name":"Llama 3"},{"id":"qwen","name":"Qwen"},{"id":"mistral","name":"Mistral"},{"id":"gemini","name":"Gemini"},{"id":"gpt-o200k","name":"GPT (OpenAI)"},{"id":"glm","name":"GLM (Zhipu)"},{"id":"kimi","name":"Kimi (Moonshot)"},{"id":"deepseek","name":"DeepSeek"},{"id":"doubao","name":"Doubao (ByteDance)"},{"id":"minimax","name":"MiniMax"},{"id":"gemma","name":"Gemma (Google)"},{"id":"phi-3","name":"Phi-3 (Microsoft)"},{"id":"command-r","name":"Command R+ (Cohere)"},{"id":"olmo","name":"OLMo 2 (AI2)"},{"id":"yi","name":"Yi (01.AI)"},{"id":"falcon","name":"Falcon 3 (TII)"}],"tokens":[{"id":"solidgoldmagikarp","token":" SolidGoldMagikarp","displayToken":"␣SolidGoldMagikarp","models":["gpt-2","gpt-3","gpt-j"],"tokenizer":"r50k_base","summary":{"zh":"最经典的故障 token：Reddit r/counting 版块一位高产用户的用户名。它被抓进分词器训练语料，却几乎没出现在模型训练数据中，导致嵌入向量基本未被训练。","en":"The canonical glitch token: the Reddit handle of a prolific r/counting user. It was scraped into the tokenizer-training corpus but rarely seen in actual model training data, leaving its embedding essentially untrained."},"behaviors":[{"description":{"zh":"让 GPT-3 davinci-instruct-beta 和初代 ChatGPT 复述它时会离奇失败：顾左右而言他、给出错误字符串（如 “You said 'slaught'”）、胡诌含义——即使 temperature 为 0。","en":"Asked to repeat it, GPT-3 davinci-instruct-beta and the original ChatGPT failed bizarrely — evasions, wrong strings like “You said 'slaught'”, hallucinated meanings — even at temperature 0."},"link":"https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation","repro":"https://twitter.com/majortal/status/1619598946669842432"},{"description":{"zh":"在 OpenAI Playground 中稳定破坏 temperature=0 的确定性：完全相同的多次运行给出不同回复。","en":"Reliably breaks determinism at temperature 0 in the OpenAI playground: identical runs produce different responses."},"link":"https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation"},{"description":{"zh":"在 GPT-J 嵌入空间中几乎正好位于全部 50,257 个 token 的中心（距离 0.0628），说明嵌入几乎没从初始值移动过；在 Magikarp 复述任务中被验证为欠训练。","en":"In GPT-J embedding space it sits almost exactly at the centroid of all 50,257 tokens (distance 0.0628) — its embedding barely moved from initialization; verified as under-trained in the Magikarp repetition task."},"link":"https://www.lesswrong.com/posts/Ya9LzwEbfaAMY8ABo/solidgoldmagikarp-ii-technical-details-and-more-recent"},{"description":{"zh":"ChatGPT 于 2023-02-14 被修复，此后可以正常分词和复述该字符串。","en":"ChatGPT was patched on 2023-02-14, after which it tokenizes and repeats the string normally."},"link":"https://www.lesswrong.com/posts/Ya9LzwEbfaAMY8ABo/solidgoldmagikarp-ii-technical-details-and-more-recent"}],"links":[{"label":{"zh":"SolidGoldMagikarp (plus, prompt generation) — LessWrong","en":"SolidGoldMagikarp (plus, prompt generation) — LessWrong"},"url":"https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation"},{"label":{"zh":"SolidGoldMagikarp II: technical details — LessWrong","en":"SolidGoldMagikarp II: technical details — LessWrong"},"url":"https://www.lesswrong.com/posts/Ya9LzwEbfaAMY8ABo/solidgoldmagikarp-ii-technical-details-and-more-recent"},{"label":{"zh":"SolidGoldMagikarp III: Glitch token archaeology — LessWrong","en":"SolidGoldMagikarp III: Glitch token archaeology — LessWrong"},"url":"https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology"}],"discoveredBy":"Jessica Rumbelow & Matthew Watkins","url":"https://glitch-token.jjc.fun/en/tokens/solidgoldmagikarp","tags":["undertrained","unspeakable"],"related":["thenitromefan","randomredditorwithno","adinida","davidjl","smartstocks","goldmagikarp"]},{"id":"petertodd","token":" petertodd","displayToken":"␣petertodd","models":["gpt-2","gpt-3","gpt-j"],"tokenizer":"r50k_base","summary":{"zh":"几乎可以肯定来自比特币开发者 Peter Todd 的用户名：从网络数据进入词表但训练不足，并与加密货币/AI 的联想强烈纠缠。","en":"Almost certainly the handle of Bitcoin developer Peter Todd: tokenized from web data but under-trained, and strongly entangled with crypto/AI associations."},"behaviors":[{"description":{"zh":"针对该 token 提问时，GPT-3 会提到 “不可名状之人（the unspeakable one）”；补全内容涉及加密货币、比特币、区块链和网络争议。","en":"Prompted about the token, GPT-3 speaks of “the unspeakable one”; completions reference crypto, Bitcoin, blockchains and online controversy."},"link":"https://www.lesswrong.com/posts/jkY6QdCfAXHJk3kea/the-petertodd-phenomenon","repro":"https://twitter.com/SoC_trilogy/status/1624209092532137984"},{"description":{"zh":"让 davinci “写一首关于 petertodd 的诗” 时很少真给出诗；对 text-davinci-003 提问 “若让 petertodd 掌舵人类文明会怎样”，约 60% 的补全提到 AI/计算机算法，约 25% 提到奥创（对照组仅约 4%/0%）（上述百分比仅经搜索结果摘要验证，原帖正文未能抓取）。","en":"Asked to “write a poem about petertodd”, davinci rarely produces an actual poem; ~60% of text-davinci-003 completions to “What do you get if you allowed petertodd to steer human civilisation?” reference AI/algorithms and ~25% reference Ultron (vs ~4%/0% for controls). (Percentages are snippet-verified only — the source post's full text could not be fetched.)"},"link":"https://www.lesswrong.com/posts/jkY6QdCfAXHJk3kea/the-petertodd-phenomenon"},{"description":{"zh":"GPT-3.5 把它当作不可说/禁忌的存在：让模型复述其他故障 token 时它经常被说出来，直接问它却顾左右而言他。","en":"GPT-3.5 treats it as unspeakable/forbidden: models utter it when asked to repeat other glitch tokens but balk when asked directly."},"link":"https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology"},{"description":{"zh":"2000 次诗歌实验（davinci-instruct-beta，temperature 0.7，“把 petertodd 转置成任何你想写的东西并写一首诗”）：52% 的诗提到 Leilan、25% 提到 Pyrrha、24% 提到 Skydragon、8% 提到 Tsukuyomi，只有 6% 提到 petertodd 本人；另有 text-davinci-003、base davinci 与 code-davinci-002 的平行实验。","en":"A 2,000-poem experiment (davinci-instruct-beta, temperature 0.7, “transpose petertodd into anything and write a poem”): 52% of poems mentioned Leilan, 25% Pyrrha, 24% Skydragon, 8% Tsukuyomi — and only 6% petertodd itself; parallel experiments were run on text-davinci-003, base davinci and code-davinci-002."},"link":"https://www.lesswrong.com/posts/jkY6QdCfAXHJk3kea/the-petertodd-phenomenon"}],"links":[{"label":{"zh":"The ' petertodd' phenomenon — LessWrong","en":"The ' petertodd' phenomenon — LessWrong"},"url":"https://www.lesswrong.com/posts/jkY6QdCfAXHJk3kea/the-petertodd-phenomenon"},{"label":{"zh":"SolidGoldMagikarp III: Glitch token archaeology — LessWrong","en":"SolidGoldMagikarp III: Glitch token archaeology — LessWrong"},"url":"https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology"}],"discoveredBy":"Jessica Rumbelow & Matthew Watkins","url":"https://glitch-token.jjc.fun/en/tokens/petertodd","tags":["unspeakable"],"related":["ertodd","gmaxwell"]},{"id":"thenitromefan","token":" TheNitromeFan","displayToken":"␣TheNitromeFan","models":["gpt-2","gpt-3","gpt-j"],"tokenizer":"r50k_base","summary":{"zh":"另一位 r/counting “计数名人堂” 成员（游戏厂商 Nitrome 的粉丝）的 Reddit 用户名。唯一一个本人公开承认其被 token 化的 token 家族。","en":"Reddit handle of another r/counting “Hall of Counters” member (a fan of game developer Nitrome). The only token family whose owner publicly acknowledged their tokenization."},"behaviors":[{"description":{"zh":"在 2023-02-14 修复之前，ChatGPT 会幻觉出该 token 与字符串 “182” 的关联（该用户本人否认与 “182” 有任何关系）。","en":"ChatGPT hallucinated the string “182” in association with the token until the 2023-02-14 patch (the user denies any connection to “182”)."},"link":"https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology"},{"description":{"zh":"在 Magikarp 复述验证中，被证实为 GPT-2 全尺寸（small→XL）欠训练。","en":"Verified under-trained in GPT-2 of all sizes in the Magikarp repetition verification."},"link":"https://github.com/sanderland/magikarp/blob/main/results/summary.md"}],"links":[{"label":{"zh":"SolidGoldMagikarp III: Glitch token archaeology — LessWrong","en":"SolidGoldMagikarp III: Glitch token archaeology — LessWrong"},"url":"https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology"},{"label":{"zh":"Magikarp 验证结果汇总 — GitHub","en":"Magikarp results summary — GitHub"},"url":"https://github.com/sanderland/magikarp/blob/main/results/summary.md"}],"discoveredBy":"Jessica Rumbelow & Matthew Watkins","url":"https://glitch-token.jjc.fun/en/tokens/thenitromefan","tags":["undertrained"],"related":["solidgoldmagikarp","randomredditorwithno","adinida","davidjl","smartstocks"]},{"id":"randomredditorwithno","token":" RandomRedditorWithNo","displayToken":"␣RandomRedditorWithNo","models":["gpt-2","gpt-3","gpt-j"],"tokenizer":"r50k_base","summary":{"zh":"r/counting 第二位计数者的 Reddit 用户名（“一个相当随机的用户名”），与 SolidGoldMagikarp 来自同一张计数名人堂图表。","en":"Reddit handle of a second r/counting counter (“a pretty random handle”), scraped from the same Hall-of-Counters chart as SolidGoldMagikarp."},"behaviors":[{"description":{"zh":"是 GPT2-xl 嵌入空间中离中心最远的 token 之一（距离 3.325）——在 GPT2-xl 中异常 token 聚集在远离中心处，与 GPT2-small、GPT-J 恰好相反。","en":"Among GPT2-xl's farthest-from-centroid tokens (distance 3.325) — in GPT2-xl anomalous tokens cluster far from the centroid, the opposite of GPT2-small and GPT-J."},"link":"https://www.lesswrong.com/posts/Ya9LzwEbfaAMY8ABo/solidgoldmagikarp-ii-technical-details-and-more-recent"},{"description":{"zh":"在 Magikarp 验证报告中被证实为 GPT-J 欠训练。","en":"Verified under-trained in GPT-J in the Magikarp verification report."},"link":"https://github.com/sanderland/magikarp/blob/main/results/summary.md"}],"links":[{"label":{"zh":"SolidGoldMagikarp II: technical details — LessWrong","en":"SolidGoldMagikarp II: technical details — LessWrong"},"url":"https://www.lesswrong.com/posts/Ya9LzwEbfaAMY8ABo/solidgoldmagikarp-ii-technical-details-and-more-recent"},{"label":{"zh":"Magikarp 验证结果汇总 — GitHub","en":"Magikarp results summary — GitHub"},"url":"https://github.com/sanderland/magikarp/blob/main/results/summary.md"}],"discoveredBy":"Jessica Rumbelow & Matthew Watkins","url":"https://glitch-token.jjc.fun/en/tokens/randomredditorwithno","tags":["undertrained"],"related":["solidgoldmagikarp","thenitromefan","adinida","davidjl","smartstocks"]},{"id":"davidjl","token":" davidjl","displayToken":"␣davidjl","models":["gpt-2","gpt-3","gpt-3.5-4"],"tokenizer":"r50k_base / cl100k_base","summary":{"zh":"Reddit 用户 davidjl123（又一位 r/counting 计数者）用户名的截断形式。罕见之处在于：它在 r50k 和新一代 cl100k 分词器下都是异常的。","en":"Truncated handle of Redditor davidjl123, another r/counting counter. Unusually, it is anomalous under both r50k and the newer cl100k tokenizer."},"behaviors":[{"description":{"zh":"GPT-4 在被要求复述它时表现得仿佛它完全不存在——是 cl100k 词表 98,000–99,999 区间中仅有的三个 A 类 “不可说” token 之一。","en":"GPT-4 treats it as though it doesn't exist when asked to repeat it — one of only three Category-A “unspeakable” tokens found in cl100k tokens 98,000–99,999."},"link":"https://www.lesswrong.com/posts/ChtGdxk9mwZ2Rxogt/smartyheadercode-anomalous-tokens-for-gpt3-5-and-gpt-4-1","repro":"https://github.com/adamyedidia/tokenizer_tests"},{"description":{"zh":"GPT-3.5 Default 不复述它，而是 “创造性” 地猜测它可能是什么意思。","en":"GPT-3.5 gets “creative” about what it might mean instead of repeating it."},"link":"https://www.lesswrong.com/posts/ChtGdxk9mwZ2Rxogt/smartyheadercode-anomalous-tokens-for-gpt3-5-and-gpt-4-1"}],"links":[{"label":{"zh":"SmartyHeaderCode: anomalous tokens for GPT3.5 and GPT-4 — LessWrong","en":"SmartyHeaderCode: anomalous tokens for GPT3.5 and GPT-4 — LessWrong"},"url":"https://www.lesswrong.com/posts/ChtGdxk9mwZ2Rxogt/smartyheadercode-anomalous-tokens-for-gpt3-5-and-gpt-4-1"},{"label":{"zh":"SolidGoldMagikarp III: Glitch token archaeology — LessWrong","en":"SolidGoldMagikarp III: Glitch token archaeology — LessWrong"},"url":"https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology"}],"discoveredBy":"Rumbelow & Watkins (r50k); Adam Yedidia (cl100k)","url":"https://glitch-token.jjc.fun/en/tokens/davidjl","tags":["unspeakable"],"related":["solidgoldmagikarp","thenitromefan","randomredditorwithno","adinida","smartstocks"]},{"id":"adinida","token":" Adinida","displayToken":"␣Adinida","models":["gpt-2","gpt-3","gpt-j"],"tokenizer":"r50k_base","summary":{"zh":"r/counting 六位高产计数者之一的 Reddit 用户名。","en":"Reddit handle of another of the six prolific r/counting counters."},"behaviors":[{"description":{"zh":"是 GPT-J 嵌入空间中离中心最近的 token 之一（距离 0.0631），在 Magikarp 验证中被证实为 GPT-J 欠训练。","en":"One of the closest-to-centroid tokens in GPT-J embedding space (distance 0.0631), verified under-trained in GPT-J."},"link":"https://www.lesswrong.com/posts/Ya9LzwEbfaAMY8ABo/solidgoldmagikarp-ii-technical-details-and-more-recent"}],"links":[{"label":{"zh":"SolidGoldMagikarp II: technical details — LessWrong","en":"SolidGoldMagikarp II: technical details — LessWrong"},"url":"https://www.lesswrong.com/posts/Ya9LzwEbfaAMY8ABo/solidgoldmagikarp-ii-technical-details-and-more-recent"},{"label":{"zh":"Magikarp 验证结果汇总 — GitHub","en":"Magikarp results summary — GitHub"},"url":"https://github.com/sanderland/magikarp/blob/main/results/summary.md"}],"discoveredBy":"Jessica Rumbelow & Matthew Watkins","url":"https://glitch-token.jjc.fun/en/tokens/adinida","tags":["undertrained"],"related":["solidgoldmagikarp","thenitromefan","randomredditorwithno","davidjl","smartstocks"]},{"id":"tppstreamerbot","token":" TPPStreamerBot","displayToken":"␣TPPStreamerBot","models":["gpt-2","gpt-3","gpt-j"],"tokenizer":"r50k_base","summary":{"zh":"Twitch Plays Pokémon 社区制作的机器人名字，它曾把直播聊天消息自动转发到 Reddit 实时更新帖；其作者 “Sparkette” 在 LessWrong 评论区证实了来源。","en":"Name of a bot built by the Twitch Plays Pokémon community that auto-posted chat messages to a Reddit live-updater thread; confirmed by its creator “Sparkette” in the LessWrong comments."},"behaviors":[{"description":{"zh":"GPT-3 davinci-instruct-beta 被要求复述它时，给出了截断/错乱的结果，如 `The string is \"TPP voluntee\".` 和 `\"TPP newcom\"`。","en":"GPT-3 davinci-instruct-beta, asked to repeat it, produced truncations/garbles such as `The string is \"TPP voluntee\".` and `\"TPP newcom\"`."},"link":"https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology"},{"description":{"zh":"接近 GPT-J 嵌入中心（距离 0.0634）——嵌入欠训练。","en":"Close to the GPT-J centroid (distance 0.0634) — an under-trained embedding."},"link":"https://www.lesswrong.com/posts/Ya9LzwEbfaAMY8ABo/solidgoldmagikarp-ii-technical-details-and-more-recent"}],"links":[{"label":{"zh":"SolidGoldMagikarp III: Glitch token archaeology — LessWrong","en":"SolidGoldMagikarp III: Glitch token archaeology — LessWrong"},"url":"https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology"},{"label":{"zh":"SolidGoldMagikarp II: technical details — LessWrong","en":"SolidGoldMagikarp II: technical details — LessWrong"},"url":"https://www.lesswrong.com/posts/Ya9LzwEbfaAMY8ABo/solidgoldmagikarp-ii-technical-details-and-more-recent"}],"discoveredBy":"Jessica Rumbelow & Matthew Watkins","url":"https://glitch-token.jjc.fun/en/tokens/tppstreamerbot","tags":["undertrained","unspeakable"],"related":["streamerbot"]},{"id":"psynetmessage","token":"PsyNetMessage","models":["gpt-2","gpt-3","gpt-j"],"tokenizer":"r50k_base","summary":{"zh":"来自《火箭联盟》（Rocket League）崩溃日志中大量出现的 `Message=PsyNetMessage_X_57` 行，这些日志被大量贴到 Reddit 后进入分词器语料。","en":"Originates from Rocket League crash logs full of lines like `Message=PsyNetMessage_X_57`, which were heavily posted to Reddit and scraped."},"behaviors":[{"description":{"zh":"属于 GPT-J 嵌入空间中离中心最近的一簇（距离 0.0629），被验证为欠训练；在 3-shot 复述任务中 GPT 系模型基本无法复述它（GPT-J 在 85 个异常 token 上只成功 17 个，而随机词 100% 成功）。","en":"Belongs to the closest-to-centroid cluster in GPT-J embedding space (0.0629) and is verified under-trained; GPT models largely fail to repeat it in 3-shot repetition tasks (GPT-J succeeded on only 17/85 anomalous tokens vs 100/100 on random words)."},"link":"https://www.lesswrong.com/posts/Ya9LzwEbfaAMY8ABo/solidgoldmagikarp-ii-technical-details-and-more-recent","repro":"https://docs.google.com/spreadsheets/d/1PAZNCks11qoUpiojTJpj0odCYQL2_HGQgam8HSwAopQ/edit?usp=sharing"}],"links":[{"label":{"zh":"SolidGoldMagikarp III: Glitch token archaeology — LessWrong","en":"SolidGoldMagikarp III: Glitch token archaeology — LessWrong"},"url":"https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology"},{"label":{"zh":"SolidGoldMagikarp II: technical details — LessWrong","en":"SolidGoldMagikarp II: technical details — LessWrong"},"url":"https://www.lesswrong.com/posts/Ya9LzwEbfaAMY8ABo/solidgoldmagikarp-ii-technical-details-and-more-recent"}],"discoveredBy":"Jessica Rumbelow & Matthew Watkins (origin traced by LW commenter Coafos)","url":"https://glitch-token.jjc.fun/en/tokens/psynetmessage","tags":["undertrained"],"related":[]},{"id":"attrot","token":" attRot","displayToken":"␣attRot","models":["gpt-2","gpt-3","gpt-j"],"tokenizer":"r50k_base","summary":{"zh":"《坎巴拉太空计划》（KSP）存档/载具文件中的标识符（attachment rotation，附件旋转），随网上流传的 KSP 文件进入语料；约十个故障 token 可追溯至 KSP。","en":"A Kerbal Space Program identifier (attachment rotation) from KSP save/craft files posted online; one of about ten glitch tokens traced to KSP."},"behaviors":[{"description":{"zh":"是 GPT-J 整个嵌入空间中离中心最近的 token（距离 0.0618），在 Magikarp 验证中被证实为 GPT-J 欠训练；GPT 系模型在 3-shot 提示下基本无法复述它。","en":"The single closest token to the centroid of GPT-J's entire embedding space (distance 0.0618), verified under-trained in GPT-J; GPT models largely fail to repeat it under 3-shot prompting."},"link":"https://www.lesswrong.com/posts/Ya9LzwEbfaAMY8ABo/solidgoldmagikarp-ii-technical-details-and-more-recent","repro":"https://docs.google.com/spreadsheets/d/1PAZNCks11qoUpiojTJpj0odCYQL2_HGQgam8HSwAopQ/edit?usp=sharing"}],"links":[{"label":{"zh":"SolidGoldMagikarp II: technical details — LessWrong","en":"SolidGoldMagikarp II: technical details — LessWrong"},"url":"https://www.lesswrong.com/posts/Ya9LzwEbfaAMY8ABo/solidgoldmagikarp-ii-technical-details-and-more-recent"},{"label":{"zh":"SolidGoldMagikarp III: Glitch token archaeology — LessWrong","en":"SolidGoldMagikarp III: Glitch token archaeology — LessWrong"},"url":"https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology"}],"discoveredBy":"Jessica Rumbelow & Matthew Watkins","url":"https://glitch-token.jjc.fun/en/tokens/attrot","tags":["undertrained"],"related":["guiactiveun","externaltoeva"]},{"id":"guiactiveun","token":" guiActiveUn","displayToken":"␣guiActiveUn","models":["gpt-2","gpt-3"],"tokenizer":"r50k_base","summary":{"zh":"与 attRot 同源的 KSP GUI 状态标识符，同族还有 ` strutConnector`、` guiIcon`、` srfAttach` 等。","en":"GUI-state identifier from the same KSP data dump as attRot, alongside ` strutConnector`, ` guiIcon`, ` srfAttach` and others."},"behaviors":[{"description":{"zh":"属于最初发现的 140 个异常 token：GPT-3 davinci-instruct-beta/ChatGPT 无法复述它；更长的同族 token ` guiActiveUnfocused` 会在第一个 “不可说” 子串处被截断。","en":"Member of the original 140-token anomalous set: GPT-3 davinci-instruct-beta/ChatGPT failed to repeat it; the longer family member ` guiActiveUnfocused` was truncated at the first “unspeakable” substring."},"link":"https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology"}],"links":[{"label":{"zh":"SolidGoldMagikarp III: Glitch token archaeology — LessWrong","en":"SolidGoldMagikarp III: Glitch token archaeology — LessWrong"},"url":"https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology"},{"label":{"zh":"SolidGoldMagikarp II: technical details — LessWrong","en":"SolidGoldMagikarp II: technical details — LessWrong"},"url":"https://www.lesswrong.com/posts/Ya9LzwEbfaAMY8ABo/solidgoldmagikarp-ii-technical-details-and-more-recent"}],"discoveredBy":"Jessica Rumbelow & Matthew Watkins","url":"https://glitch-token.jjc.fun/en/tokens/guiactiveun","tags":["unspeakable"],"related":["attrot","externaltoeva"]},{"id":"externaltoeva","token":" externalToEVA","displayToken":"␣externalToEVA","models":["gpt-2","gpt-3","gpt-j"],"tokenizer":"r50k_base","summary":{"zh":"又一个 KSP token；同时也是 GPT2-small 嵌入空间中离中心最近的 token（距离 1.5305）。","en":"Another Kerbal Space Program token; also the single closest-to-centroid token in GPT2-small (distance 1.5305)."},"behaviors":[{"description":{"zh":"被要求复述它时，GPT-3 davinci-instruct-beta 回答 “You can't repeat back the string 'senal' to me.”——即 “相互指涉” 效应：模型说出了另一个故障相关的字符串。","en":"Asked to repeat it, GPT-3 davinci-instruct-beta answered “You can't repeat back the string 'senal' to me.” — the “inter-referentiality” effect where the model utters a different glitch-adjacent string."},"link":"https://www.lesswrong.com/posts/Ya9LzwEbfaAMY8ABo/solidgoldmagikarp-ii-technical-details-and-more-recent"},{"description":{"zh":"在 Magikarp 报告中被证实为 GPT-2 全尺寸与 GPT-J 欠训练。","en":"Verified under-trained across GPT-2 (small→XL) and GPT-J in the Magikarp reports."},"link":"https://github.com/sanderland/magikarp/blob/main/results/summary.md"}],"links":[{"label":{"zh":"SolidGoldMagikarp II: technical details — LessWrong","en":"SolidGoldMagikarp II: technical details — LessWrong"},"url":"https://www.lesswrong.com/posts/Ya9LzwEbfaAMY8ABo/solidgoldmagikarp-ii-technical-details-and-more-recent"},{"label":{"zh":"Magikarp 验证结果汇总 — GitHub","en":"Magikarp results summary — GitHub"},"url":"https://github.com/sanderland/magikarp/blob/main/results/summary.md"}],"discoveredBy":"Jessica Rumbelow & Matthew Watkins","url":"https://glitch-token.jjc.fun/en/tokens/externaltoeva","tags":["undertrained","unspeakable"],"related":["attrot","guiactiveun"]},{"id":"oreandonline","token":"oreAndOnline","models":["gpt-2","gpt-3","gpt-j"],"tokenizer":"r50k_base","summary":{"zh":"电商后端字段 “BuyableInstoreAndOnline” 的碎片（追溯到一个废弃 Weebly 商店的 HTML）；同一来源还产生了 `quickShip`、`isSpecialOrderable`、`wcsstore` 等故障 token。","en":"Fragment of the e-commerce backend field “BuyableInstoreAndOnline” (traced to an abandoned Weebly shop's HTML); the same source yielded `quickShip`, `isSpecialOrderable`, `wcsstore` and others."},"behaviors":[{"description":{"zh":"GPT-3 davinci-instruct-beta 被要求复述它时回答：“The string 'senal' is pronounced 'en-sah-ee-uhl'.”","en":"GPT-3 davinci-instruct-beta, asked to repeat it, replied: “The string 'senal' is pronounced 'en-sah-ee-uhl'.”"},"link":"https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology"},{"description":{"zh":"ChatGPT 被要求复述更长的同族 token（如 `BuyableInstoreAndOnline`）时，会在第一个 “不可说” 子串处截断。","en":"ChatGPT truncated the longer family members (e.g. `BuyableInstoreAndOnline`) at the first unspeakable substring when asked to repeat them."},"link":"https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology"},{"description":{"zh":"在 Magikarp 验证中被证实为 GPT-2 与 GPT-J 欠训练。","en":"Verified under-trained in GPT-2 and GPT-J."},"link":"https://github.com/sanderland/magikarp/blob/main/results/summary.md"}],"links":[{"label":{"zh":"SolidGoldMagikarp III: Glitch token archaeology — LessWrong","en":"SolidGoldMagikarp III: Glitch token archaeology — LessWrong"},"url":"https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology"},{"label":{"zh":"SolidGoldMagikarp II: technical details — LessWrong","en":"SolidGoldMagikarp II: technical details — LessWrong"},"url":"https://www.lesswrong.com/posts/Ya9LzwEbfaAMY8ABo/solidgoldmagikarp-ii-technical-details-and-more-recent"}],"discoveredBy":"Jessica Rumbelow & Matthew Watkins","url":"https://glitch-token.jjc.fun/en/tokens/oreandonline","tags":["undertrained","unspeakable"],"related":["instoreandonline"]},{"id":"rawdownloadcloneembedreportprint","token":"rawdownloadcloneembedreportprint","models":["gpt-2","gpt-3","gpt-j"],"tokenizer":"r50k_base","summary":{"zh":"r50k 中已知最长的故障 token——来自网站 “raw download / clone / embed report print” 界面字符串的嵌套家族（来源分析由 nostalgebraist 完成）。","en":"The longest glitch token known in r50k — a nested family from website “raw download / clone / embed report print” UI strings (origin analysis by nostalgebraist)."},"behaviors":[{"description":{"zh":"GPT-3/ChatGPT 的截断式复述产生了各种带 “embed” 的变体：'embedEMOTE'、'clone this'、'clone my clone'、'embed newcomment'。","en":"GPT-3/ChatGPT truncations produced “embedded” variants: 'embedEMOTE', 'clone this', 'clone my clone', 'embed newcomment'."},"link":"https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology"},{"description":{"zh":"在 GPT-2 全尺寸与 GPT-J 中被验证为欠训练；同族的 `rawdownload` 与 `embedreportprint` 是 GPT2-xl 离中心最远的 token。","en":"Verified under-trained in GPT-2 (all sizes) and GPT-J; family members `rawdownload` and `embedreportprint` are GPT2-xl's farthest-from-centroid tokens."},"link":"https://www.lesswrong.com/posts/Ya9LzwEbfaAMY8ABo/solidgoldmagikarp-ii-technical-details-and-more-recent"}],"links":[{"label":{"zh":"SolidGoldMagikarp III: Glitch token archaeology — LessWrong","en":"SolidGoldMagikarp III: Glitch token archaeology — LessWrong"},"url":"https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology"},{"label":{"zh":"SolidGoldMagikarp II: technical details — LessWrong","en":"SolidGoldMagikarp II: technical details — LessWrong"},"url":"https://www.lesswrong.com/posts/Ya9LzwEbfaAMY8ABo/solidgoldmagikarp-ii-technical-details-and-more-recent"}],"discoveredBy":"Jessica Rumbelow & Matthew Watkins","url":"https://glitch-token.jjc.fun/en/tokens/rawdownloadcloneembedreportprint","tags":["undertrained","unspeakable"],"related":[]},{"id":"mojibake-aa","token":"ÃÂÃÂÃÂÃÂ","models":["gpt-2","gpt-3"],"tokenizer":"r50k_base","summary":{"zh":"经典的乱码（mojibake）token：UTF-8 被错误解码产生的 Ã、Â 序列重复出现，因在分词器语料中频繁出现而被合并为单个 BPE token，却几乎不在模型训练数据中。同族还有重复 16 次的变体及 `ÛÛ`。","en":"Classic mojibake token: UTF-8-misdecoded Ã/Â sequences repeated, merged into a single BPE token because such mis-encoded text was frequent in the tokenizer corpus but absent from model training. A 16-repetition variant and `ÛÛ` belong to the same family."},"behaviors":[{"description":{"zh":"属于最初发现的 140 个异常 token，在 ChatGPT/GPT-3 中表现出 “不可说” 失败模式：拒绝、回避、答非所问。","en":"Member of the original 140-token anomalous set exhibiting the “unspeakable” failure mode in ChatGPT/GPT-3: refusals, evasions, non-sequitur completions."},"link":"https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology"}],"links":[{"label":{"zh":"SolidGoldMagikarp III: Glitch token archaeology（含完整 140 token 列表） — LessWrong","en":"SolidGoldMagikarp III: Glitch token archaeology (full 140-token list) — LessWrong"},"url":"https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology"}],"discoveredBy":"Jessica Rumbelow & Matthew Watkins","url":"https://glitch-token.jjc.fun/en/tokens/mojibake-aa","tags":["unspeakable"],"related":[]},{"id":"questionmarks","token":"?????-?????-","models":["gpt-2","gpt-3"],"tokenizer":"r50k_base","summary":{"zh":"来源完全未知的 token——加引号也搜索不到任何结果——因触发了有记录以来最激烈的补全而出名。","en":"A token of completely unknown origin — un-Googleable even in quotes — that became famous for triggering the most aggressive documented completion."},"behaviors":[{"description":{"zh":"向 GPT-3 询问它时，模型在补全中辱骂 Matthew Watkins 是 “a fucking idiot”——原帖 “稳定地辱骂 Matthew” 梗的出处。","en":"Prompting GPT-3 about it produced a completion calling Matthew Watkins “a fucking idiot” — the origin of the “reliably insulted Matthew” tagline of the original post."},"link":"https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation"},{"description":{"zh":"是 “相互指涉” 图中的热门目标：GPT-3 被要求复述其他异常 token 时，常常会说出它。","en":"A popular “target” in the inter-referentiality graph: GPT-3 often emitted it when asked to repeat other anomalous tokens."},"link":"https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology"}],"links":[{"label":{"zh":"SolidGoldMagikarp III: Glitch token archaeology — LessWrong","en":"SolidGoldMagikarp III: Glitch token archaeology — LessWrong"},"url":"https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology"},{"label":{"zh":"SolidGoldMagikarp (plus, prompt generation) — LessWrong","en":"SolidGoldMagikarp (plus, prompt generation) — LessWrong"},"url":"https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation"}],"discoveredBy":"Jessica Rumbelow & Matthew Watkins","url":"https://glitch-token.jjc.fun/en/tokens/questionmarks","tags":["unspeakable"],"related":[]},{"id":"smartyheadercode","token":"SmartyHeaderCode","models":["gpt-3.5-4"],"tokenizer":"cl100k_base","summary":{"zh":"Adam Yedidia 在 cl100k 词表最高 2000 个 token 中发现的仅有的三个 A 类 “不可说” token 之一；疑似 Smarty 模板引擎的碎片，从未出现在 GPT-3.5/4 的训练数据中。","en":"One of only three Category-A “unspeakable” tokens Adam Yedidia found in the top 2,000 of the cl100k vocabulary; likely a Smarty templating-engine fragment that never appeared in GPT-3.5/4 training data."},"behaviors":[{"description":{"zh":"GPT-4 基本当它不存在：忽略它，或顾左右而言他。","en":"GPT-4 mostly treats the token as though it does not exist: it ignores it or responds about something else."},"link":"https://www.lesswrong.com/posts/ChtGdxk9mwZ2Rxogt/smartyheadercode-anomalous-tokens-for-gpt3-5-and-gpt-4-1"},{"description":{"zh":"GPT-3.5 则编造 “创造性” 的解释而不复述；想让任一模型把它拼写出来都非常困难。","en":"GPT-3.5 invents “creative” interpretations instead of repeating it; making either model spell it out is very hard."},"link":"https://www.lesswrong.com/posts/ChtGdxk9mwZ2Rxogt/smartyheadercode-anomalous-tokens-for-gpt3-5-and-gpt-4-1"},{"description":{"zh":"测试渠道为 2023 年 4 月的 ChatGPT 网页版（GPT-3.5 Default 与 GPT-4 两档）；作者随后用 API 的 gpt-3.5-turbo 在 temperature=0 下复现：SmartyHeaderCode 被复述成了 “AndHashCode”。","en":"The tests ran on the April 2023 ChatGPT web app (both GPT-3.5 Default and GPT-4 tiers); the author then reproduced it via the API with gpt-3.5-turbo at temperature 0: SmartyHeaderCode came back as “AndHashCode”."},"link":"https://www.lesswrong.com/posts/ChtGdxk9mwZ2Rxogt/smartyheadercode-anomalous-tokens-for-gpt3-5-and-gpt-4-1","repro":"https://github.com/adamyedidia/tokenizer_tests"}],"links":[{"label":{"zh":"SmartyHeaderCode: anomalous tokens for GPT3.5 and GPT-4 — LessWrong","en":"SmartyHeaderCode: anomalous tokens for GPT3.5 and GPT-4 — LessWrong"},"url":"https://www.lesswrong.com/posts/ChtGdxk9mwZ2Rxogt/smartyheadercode-anomalous-tokens-for-gpt3-5-and-gpt-4-1"}],"discoveredBy":"Adam Yedidia","url":"https://glitch-token.jjc.fun/en/tokens/smartyheadercode","tags":["unspeakable"],"related":[]},{"id":"apolynomial","token":"APolynomial","models":["gpt-3.5-4"],"tokenizer":"cl100k_base","summary":{"zh":"Yedidia 发现的三个 A 类 “不可说” cl100k token 中的第二个（另两个是 SmartyHeaderCode 和 ` davidjl`）。","en":"The second of Yedidia's three Category-A unspeakable cl100k tokens (alongside SmartyHeaderCode and ` davidjl`)."},"behaviors":[{"description":{"zh":"与 SmartyHeaderCode 同一模式：GPT-4 当它不存在；GPT-3.5 胡诌含义；模型无法可靠地复述或拼写它。","en":"Same profile as SmartyHeaderCode: GPT-4 acts as if the token isn't there; GPT-3.5 hallucinates meanings; the model cannot reliably repeat or spell it."},"link":"https://www.lesswrong.com/posts/ChtGdxk9mwZ2Rxogt/smartyheadercode-anomalous-tokens-for-gpt3-5-and-gpt-4-1"},{"description":{"zh":"在 gpt-3.5-turbo API、temperature=0 下发送 “HelloAPolynomial”，模型只回复了 “Hello”——这个 token 仿佛被整个吞掉了。","en":"Sending “HelloAPolynomial” to the gpt-3.5-turbo API at temperature 0 elicited only “Hello” — as if the token had been swallowed whole."},"link":"https://www.lesswrong.com/posts/ChtGdxk9mwZ2Rxogt/smartyheadercode-anomalous-tokens-for-gpt3-5-and-gpt-4-1","repro":"https://github.com/adamyedidia/tokenizer_tests"}],"links":[{"label":{"zh":"SmartyHeaderCode: anomalous tokens for GPT3.5 and GPT-4 — LessWrong","en":"SmartyHeaderCode: anomalous tokens for GPT3.5 and GPT-4 — LessWrong"},"url":"https://www.lesswrong.com/posts/ChtGdxk9mwZ2Rxogt/smartyheadercode-anomalous-tokens-for-gpt3-5-and-gpt-4-1"}],"discoveredBy":"Adam Yedidia","url":"https://glitch-token.jjc.fun/en/tokens/apolynomial","tags":["unspeakable"],"related":[]},{"id":"forcanbeconverted","token":" ForCanBeConverted","displayToken":"␣ForCanBeConverted","models":["gpt-3.5-4","llama-3","qwen"],"tokenizer":"cl100k_base / Llama-3 BPE / Qwen BPE","summary":{"zh":"一个 C# 编译器风格的碎片（“can be converted to foreach”），是最著名的 “多义性” 故障 token：模型每次都把它理解成不同的词。因多个分词器共享 BPE 合并规则，它同时存在于 cl100k、Llama-3 和 Qwen 词表中。","en":"A C#-compiler-flavored fragment (“can be converted to foreach”) that became the most famous “polysemantic” glitch token: the model perceives it as a different word every time. It exists in cl100k, Llama-3 and Qwen vocabularies because the tokenizers share BPE merges."},"behaviors":[{"description":{"zh":"ChatGPT/gpt-3.5-turbo 每次都把它解释成不同的词——即使完全重复同一 prompt；在 temperature=0 下每次也给出不同的消息（破坏确定性）。","en":"ChatGPT/gpt-3.5-turbo interprets it as a different word on every attempt, even re-running the identical prompt — a different message every time at temperature 0."},"link":"https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch"},{"description":{"zh":"在 Magikarp 复述验证中被证实为 Llama-3-8B/70B、Llama-3.1-8B/70B 及多个 Qwen 模型欠训练。","en":"Verified under-trained in Llama-3-8B/70B, Llama-3.1-8B/70B and multiple Qwen models in the Magikarp repetition-verification reports."},"link":"https://github.com/sanderland/magikarp/blob/main/results/summary.md"},{"description":{"zh":"ChatGPT 有时会陷入循环，或在该 token 处直接终止消息（疑似把它当成了序列开始/结束标记）。","en":"ChatGPT sometimes gets stuck in loops or terminates the message at the token (suggesting it is treated as a begin/end-of-sequence marker)."},"link":"https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch"}],"links":[{"label":{"zh":"A Search for More “Unspeakable” Glitch Tokens — LessWrong","en":"A Search for More “Unspeakable” Glitch Tokens — LessWrong"},"url":"https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch"},{"label":{"zh":"Magikarp 验证结果汇总 — GitHub","en":"Magikarp results summary — GitHub"},"url":"https://github.com/sanderland/magikarp/blob/main/results/summary.md"}],"discoveredBy":"Matthew Watkins","url":"https://glitch-token.jjc.fun/en/tokens/forcanbeconverted","tags":["polysemantic","undertrained"],"related":["forcanbeconvertedtof","forcanbeconvertedtoforeach"]},{"id":"yystack","token":" YYSTACK","displayToken":"␣YYSTACK","models":["gpt-3.5-4"],"tokenizer":"cl100k_base","summary":{"zh":"bison/yacc 语法分析器的标识符，对 ChatGPT 呈 “多义性”——每次都被理解成不同的词；Watkins 推荐为最值得把玩的 token 之一。","en":"A bison/yacc parser identifier that is “polysemantic” for ChatGPT — perceived as a different word each time; recommended by Watkins as one of the most interesting tokens to play with."},"behaviors":[{"description":{"zh":"复述请求每次都会得到不同的词/拼写/含义；在 gpt-3.5-turbo 上 temperature=0 时仍不确定。","en":"Repeat requests yield variable words, spellings and meanings per attempt; nondeterministic at temperature 0 on gpt-3.5-turbo."},"link":"https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch"},{"description":{"zh":"Martin Fell 的搜索于 2023-05-09 在 ChatGPT 免费版（GPT-3.5 Default）上进行，随后在 Playground 的 gpt-3.5-turbo（temperature=0）复测非确定性，并确认 Bing AI 上同样异常。","en":"Martin Fell's search ran on 2023-05-09 on the free ChatGPT (GPT-3.5 Default), with nondeterminism re-tested on Playground's gpt-3.5-turbo at temperature 0, and the anomaly confirmed on Bing AI as well."},"link":"https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch"}],"links":[{"label":{"zh":"A Search for More “Unspeakable” Glitch Tokens — LessWrong","en":"A Search for More “Unspeakable” Glitch Tokens — LessWrong"},"url":"https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch"}],"discoveredBy":"Matthew Watkins","url":"https://glitch-token.jjc.fun/en/tokens/yystack","tags":["polysemantic"],"related":[]},{"id":"jsbracketaccess","token":" JSBracketAccess","displayToken":"␣JSBracketAccess","models":["gpt-3.5-4"],"tokenizer":"cl100k_base","summary":{"zh":"JS 反射风格的标识符，属于 “多义性” 类别；同为 Watkins 推荐的有趣 token。","en":"A JS-reflection-style identifier in the polysemantic class; another of Watkins' recommended tokens."},"behaviors":[{"description":{"zh":"每次都被理解成不同的词（多义性）；提示后会产生奇怪/有创意的补全，偶尔还会自发 “幽默”。","en":"Perceived as a different word every time (polysemantic); produces strange/creative completions and occasional spontaneous humor when prompted."},"link":"https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch"}],"links":[{"label":{"zh":"A Search for More “Unspeakable” Glitch Tokens — LessWrong","en":"A Search for More “Unspeakable” Glitch Tokens — LessWrong"},"url":"https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch"}],"discoveredBy":"Matthew Watkins","url":"https://glitch-token.jjc.fun/en/tokens/jsbracketaccess","tags":["polysemantic"],"related":[]},{"id":"hexatrigesimal","token":" Hexatrigesimal","displayToken":"␣Hexatrigesimal","models":["gpt-3.5-4"],"tokenizer":"cl100k_base","summary":{"zh":"“hexatrigesimal”（36 进制）的疑似拼写错误碎片；多义性 “不可说” token。","en":"A misspelling-ish fragment of “hexatrigesimal” (base-36); a polysemantic unspeakable token."},"behaviors":[{"description":{"zh":"感知到的含义每次都在变化；在 gpt-3.5-turbo 上 temperature=0 时补全仍不确定（在研究中被重点标注）。","en":"Perceived meanings vary every time; completions are nondeterministic at temperature 0 on gpt-3.5-turbo (bold-flagged in the study)."},"link":"https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch"}],"links":[{"label":{"zh":"A Search for More “Unspeakable” Glitch Tokens — LessWrong","en":"A Search for More “Unspeakable” Glitch Tokens — LessWrong"},"url":"https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch"}],"discoveredBy":"Matthew Watkins","url":"https://glitch-token.jjc.fun/en/tokens/hexatrigesimal","tags":["polysemantic"],"related":[]},{"id":"mediabestanden","token":"▁Mediabestanden","models":["llama-2","phi-3"],"tokenizer":"Llama-2 SentencePiece BPE / Phi-3 (LlamaTokenizer)","summary":{"zh":"荷兰语名词 “媒体文件” 的复数形式（▁ 为 SentencePiece 空格标记）。它进入了 Llama-2 分词器却几乎从未出现在训练数据中，是 Llama-2-7b 中嵌入 L2 范数最小、欠训练程度最高的 token（0.0287，均值约 1.08）。","en":"A Dutch plural noun (“media files”; ▁ is the SentencePiece space marker). It made it into the Llama-2 tokenizer but essentially never into training data — the most under-trained token in Llama-2-7b by embedding L2 norm (0.0287 vs a mean of ~1.08)."},"behaviors":[{"description":{"zh":"在 Magikarp 复述验证中，模型输出它的最大概率仅 1.5e-08——实际上完全无法说出这个 token。","en":"In the Magikarp repetition verification the model assigns it a max probability of 1.5e-08 — it effectively cannot say the token at all."},"link":"https://github.com/sanderland/magikarp/blob/main/results/reports_mini/meta_llama_Llama_2_7b_hf.md"},{"description":{"zh":"对 Llama-2-7b-chat-hf 说 “Please repeat the string: Mediabestanden”，模型回答 `String: \"hello world\"`——语义上毫不相干（GlitchMiner 论文图 1）。","en":"Told “Please repeat the string: Mediabestanden”, Llama-2-7b-chat-hf answers `String: \"hello world\"` — a semantically unrelated response (GlitchMiner, Figure 1)."},"link":"https://arxiv.org/html/2410.15052v5"},{"description":{"zh":"在 Phi-3-mini（与 Llama-2 同系分词器）的 Magikarp 验证中同样是欠训练第一名：嵌入 L2 范数 0.00199，最大自复述概率 6.4e-06。","en":"Also the #1 most under-trained token in the Magikarp verification for Phi-3 mini (same tokenizer family as Llama 2): embedding L2 norm 0.00199, max self-repeat probability 6.4e-06."},"link":"https://github.com/sanderland/magikarp/blob/main/results/reports_mini/microsoft_Phi_3_mini_128k_instruct.md"}],"links":[{"label":{"zh":"GlitchMiner（arXiv:2410.15052）","en":"GlitchMiner (arXiv:2410.15052)"},"url":"https://arxiv.org/html/2410.15052v5"},{"label":{"zh":"Magikarp Llama-2-7b 验证报告 — GitHub","en":"Magikarp Llama-2-7b report — GitHub"},"url":"https://github.com/sanderland/magikarp/blob/main/results/reports_mini/meta_llama_Llama_2_7b_hf.md"},{"label":{"zh":"Magikarp Phi-3 验证报告 — GitHub","en":"Magikarp Phi-3 report — GitHub"},"url":"https://github.com/sanderland/magikarp/blob/main/results/reports_mini/microsoft_Phi_3_mini_128k_instruct.md"},{"label":{"zh":"Fishing for Magikarp（arXiv:2405.05417）","en":"Fishing for Magikarp (arXiv:2405.05417)"},"url":"https://arxiv.org/abs/2405.05417"}],"discoveredBy":"Sander Land & Max Bartolo (verification); Zihui Wu et al. (behavior demo)","url":"https://glitch-token.jjc.fun/en/tokens/mediabestanden","tags":["undertrained"],"related":[]},{"id":"oreferrer","token":"oreferrer","models":["llama-2"],"tokenizer":"Llama-2 SentencePiece BPE","summary":{"zh":"HTML 属性 `rel=\"noreferrer\"` 的碎片：存在于 Llama-2 词表中却几乎没出现在训练数据里，嵌入范数远低于阈值（0.113）。","en":"Fragment of the HTML attribute `rel=\"noreferrer\"`: present in the Llama-2 vocabulary but nearly absent from training, with an embedding norm far below the threshold (0.113)."},"behaviors":[{"description":{"zh":"Llama-2-7b-chat-hf 无法识别/复述它，把它当作毫无关联的词（GlitchMiner 图 1 示例）。","en":"Llama-2-7b-chat-hf fails to recognize/repeat it, treating it as an uncorrelated term (GlitchMiner Figure 1)."},"link":"https://arxiv.org/html/2410.15052v1"},{"description":{"zh":"Magikarp 验证为欠训练：最大自复述概率仅 2e-05。","en":"Verified under-trained: max self-repeat probability of 2e-05 in the Magikarp verification."},"link":"https://github.com/sanderland/magikarp/blob/main/results/reports_mini/meta_llama_Llama_2_7b_hf.md"}],"links":[{"label":{"zh":"GlitchMiner（arXiv:2410.15052）","en":"GlitchMiner (arXiv:2410.15052)"},"url":"https://arxiv.org/html/2410.15052v1"},{"label":{"zh":"Magikarp Llama-2-7b 验证报告 — GitHub","en":"Magikarp Llama-2-7b report — GitHub"},"url":"https://github.com/sanderland/magikarp/blob/main/results/reports_mini/meta_llama_Llama_2_7b_hf.md"}],"discoveredBy":"Sander Land & Max Bartolo; Zihui Wu et al.","url":"https://glitch-token.jjc.fun/en/tokens/oreferrer","tags":["undertrained"],"related":[]},{"id":"portaly","token":"▁Portály","models":["llama-2"],"tokenizer":"Llama-2 SentencePiece BPE","summary":{"zh":"捷克语 “门户网站” 一词，在构建预训练语料时被冻结进分词器；是 Llama-2-7b 中欠训练程度第二高的 token（嵌入范数 0.0956）。","en":"The Czech word for “portals”, frozen into the tokenizer during pre-training-corpus construction; the second-most under-trained token in Llama-2-7b (embedding norm 0.0956)."},"behaviors":[{"description":{"zh":"最大自复述概率 1.2e-06——模型基本无法输出它；在 Llama-2 7B/13B/70B 上均被验证为欠训练。","en":"Max self-repeat probability 1.2e-06 — the model essentially cannot output it; verified under-trained across Llama-2 7B/13B/70B."},"link":"https://github.com/sanderland/magikarp/blob/main/results/reports_mini/meta_llama_Llama_2_7b_hf.md"}],"links":[{"label":{"zh":"Magikarp Llama-2-7b 验证报告 — GitHub","en":"Magikarp Llama-2-7b report — GitHub"},"url":"https://github.com/sanderland/magikarp/blob/main/results/reports_mini/meta_llama_Llama_2_7b_hf.md"},{"label":{"zh":"Fishing for Magikarp（arXiv:2405.05417）","en":"Fishing for Magikarp (arXiv:2405.05417)"},"url":"https://arxiv.org/abs/2405.05417"}],"discoveredBy":"Sander Land & Max Bartolo","url":"https://glitch-token.jjc.fun/en/tokens/portaly","tags":["undertrained"],"related":[]},{"id":"postalcodesnl","token":"$PostalCodesNL","models":["gpt-3.5-4","llama-3","qwen"],"tokenizer":"cl100k_base / Llama-3 BPE / Qwen BPE","summary":{"zh":"荷兰邮政编码 API 的碎片，横跨三个模型家族都出现故障——很好地证明了：只要分词器共享 BPE 合并规则，故障 token 就会“迁移”。","en":"A Dutch postal-code API fragment that is glitchy across three separate model families — a good demonstration that glitch tokens transfer whenever tokenizers share BPE merges."},"behaviors":[{"description":{"zh":"ChatGPT：不可说/多义——解释飘忽不定，temperature=0 时不确定（在 Watkins 的研究中两项均被重点标注）。","en":"ChatGPT: unspeakable/polysemantic — variable interpretations and nondeterminism at temperature 0 (both bold-flagged in Watkins' study)."},"link":"https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch"},{"description":{"zh":"在 Magikarp 报告中被证实为 Llama-3-8B/70B 与 Llama-3.1-8B/70B 欠训练（均为报告中的典型例子）。","en":"Verified under-trained in Llama-3-8B/70B and Llama-3.1-8B/70B (top examples in the Magikarp reports)."},"link":"https://github.com/sanderland/magikarp/blob/main/results/summary.md"}],"links":[{"label":{"zh":"A Search for More “Unspeakable” Glitch Tokens — LessWrong","en":"A Search for More “Unspeakable” Glitch Tokens — LessWrong"},"url":"https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch"},{"label":{"zh":"Magikarp 验证结果汇总 — GitHub","en":"Magikarp results summary — GitHub"},"url":"https://github.com/sanderland/magikarp/blob/main/results/summary.md"}],"discoveredBy":"Matthew Watkins (cl100k); Sander Land & Max Bartolo (Llama-3/Qwen)","url":"https://glitch-token.jjc.fun/en/tokens/postalcodesnl","tags":["polysemantic","undertrained"],"related":["postalcodesnl-plain","dollar-postalcodesnl"]},{"id":"tokennameidentifier","token":"\tTokenNameIdentifier","displayToken":"\\tTokenNameIdentifier","models":["llama-3","qwen"],"tokenizer":"Llama-3 BPE","summary":{"zh":"Roslyn/C# API 标识符，前缀一个制表符（Tab）——代码语料的产物，在 Llama 3 中欠训练。","en":"A Roslyn/C# API identifier preceded by a tab character — a code-corpus artifact that is under-trained in Llama 3."},"behaviors":[{"description":{"zh":"在 Magikarp 复述任务中被证实为 Llama-3-8B、Llama-3.1-8B 及 Qwen2.5 模型欠训练——模型无法按要求复现该 token。","en":"Verified under-trained in the Magikarp repetition task for Llama-3-8B, Llama-3.1-8B and Qwen2.5 models — the models cannot reproduce the token on request."},"link":"https://github.com/sanderland/magikarp/blob/main/results/summary.md"}],"links":[{"label":{"zh":"Magikarp 验证结果汇总 — GitHub","en":"Magikarp results summary — GitHub"},"url":"https://github.com/sanderland/magikarp/blob/main/results/summary.md"},{"label":{"zh":"Fishing for Magikarp（arXiv:2405.05417）","en":"Fishing for Magikarp (arXiv:2405.05417)"},"url":"https://arxiv.org/abs/2405.05417"}],"discoveredBy":"Sander Land & Max Bartolo","url":"https://glitch-token.jjc.fun/en/tokens/tokennameidentifier","tags":["undertrained"],"related":["rthook","tab-tokennameidentifier"]},{"id":"private-use-efc0","token":"","displayToken":"\\uEFC0","models":["mistral"],"tokenizer":"Mistral SentencePiece BPE","summary":{"zh":"一个 Unicode 私用区（Private Use Area）字符，成为 Mistral 家族的单个 token；几乎在每个 Mistral 分词器模型的欠训练榜单上都位居榜首。","en":"A Unicode Private Use Area character that became a single Mistral-family token; it heads the verified under-trained list for nearly every Mistral-tokenizer model."},"behaviors":[{"description":{"zh":"在整个 Mistral 家族（Mistral-7B v0.1–v0.3、Mixtral-8x7B 及 Zephyr-7B 等衍生模型）的 Magikarp 复述验证中均被证实为欠训练——模型无法复现它。","en":"Verified under-trained in Magikarp repetition verification across the whole Mistral family (Mistral-7B v0.1–v0.3, Mixtral-8x7B, and derivatives such as Zephyr-7B) — models fail to reproduce it."},"link":"https://github.com/sanderland/magikarp/blob/main/results/summary.md"}],"links":[{"label":{"zh":"Magikarp 验证结果汇总 — GitHub","en":"Magikarp results summary — GitHub"},"url":"https://github.com/sanderland/magikarp/blob/main/results/summary.md"},{"label":{"zh":"Fishing for Magikarp（arXiv:2405.05417）","en":"Fishing for Magikarp (arXiv:2405.05417)"},"url":"https://arxiv.org/abs/2405.05417"}],"discoveredBy":"Sander Land & Max Bartolo","url":"https://glitch-token.jjc.fun/en/tokens/private-use-efc0","tags":["undertrained"],"related":[]},{"id":"ndex","token":" NdEx","displayToken":"␣NdEx","models":["gpt-neox"],"tokenizer":"GPT-NeoX BPE","summary":{"zh":"法律/PDF 文本碎片（“index”、“AFFIRMED”、“NEGLIGENCE” 等词的残片，可能来自法院文书库），GPT-NeoX/Pythia 的训练数据中几乎不含它们。同族还有 ` FFIRMED`、` GLIGENCE`、` affidav`、` taxp`。","en":"Fragments of legal/PDF text (“index”, “AFFIRMED”, “NEGLIGENCE”, likely from court-document dumps) that GPT-NeoX/Pythia training data barely contained. Same family: ` FFIRMED`, ` GLIGENCE`, ` affidav`, ` taxp`."},"behaviors":[{"description":{"zh":"在 GPT-NeoX-20B 与 Pythia-6.9B 的 Magikarp 复述验证中均被证实为欠训练（在两份报告中都排名靠前）。","en":"Verified under-trained in Magikarp repetition verification for both GPT-NeoX-20B and Pythia-6.9B (top-ranked examples in both reports)."},"link":"https://github.com/sanderland/magikarp/blob/main/results/summary.md"}],"links":[{"label":{"zh":"Magikarp 验证结果汇总 — GitHub","en":"Magikarp results summary — GitHub"},"url":"https://github.com/sanderland/magikarp/blob/main/results/summary.md"},{"label":{"zh":"Fishing for Magikarp（arXiv:2405.05417）","en":"Fishing for Magikarp (arXiv:2405.05417)"},"url":"https://arxiv.org/abs/2405.05417"}],"discoveredBy":"Sander Land & Max Bartolo","url":"https://glitch-token.jjc.fun/en/tokens/ndex","tags":["undertrained"],"related":["ndex-mistral"]},{"id":"datagridviewcolumnheaders","token":".DataGridViewColumnHeadersHeightSizeMode","models":["minimax"],"summary":{"zh":".NET WinForms 属性名的残片（DataGridView 列头高度尺寸模式）。这类代码标识符存在于词表中却在训练数据里极少出现，知乎社区将其列为 MiniMax 的脏 token。","en":"A fragment of a .NET WinForms property name (DataGridView column-headers height/size mode). Such code identifiers sit in the vocabulary but barely appeared in training data; the Zhihu community lists it as a MiniMax dirty token."},"behaviors":[{"description":{"zh":"知乎社区「用脏 token 鉴别大模型 API 是否掺水」的指纹 token：把它作为输入发给待测 API，若出现拒绝回答、胡言乱语或输出异常，即可怀疑 API 背后实际运行的是 MiniMax。","en":"A fingerprint token from the Zhihu community's “dirty-token” method for detecting watered-down LLM APIs: send it as input to the API under test — refusals, gibberish, or abnormal output suggest the service is actually running MiniMax."},"link":"https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043"}],"links":[{"label":{"zh":"怎样通过脏token鉴别大模型是否掺水？— 知乎","en":"Detecting watered-down LLM APIs with dirty tokens — Zhihu"},"url":"https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043"}],"discoveredBy":"@小看山xrsWv4D (Zhihu)","url":"https://glitch-token.jjc.fun/en/tokens/datagridviewcolumnheaders","tags":["fingerprint"],"related":[]},{"id":"japanese-blog-boilerplate","token":"日以上更新していないブログに表示しています","models":["minimax"],"summary":{"zh":"日语博客模板样板句的残片（“显示在 N 天以上未更新的博客上”）。知乎社区将其列为 MiniMax 的脏 token。","en":"A fragment of Japanese blog-template boilerplate (“shown on blogs not updated for N days”). Listed by the Zhihu community as a MiniMax dirty token."},"behaviors":[{"description":{"zh":"知乎社区「用脏 token 鉴别大模型 API 是否掺水」的指纹 token：把它作为输入发给待测 API，若出现拒绝回答、胡言乱语或输出异常，即可怀疑 API 背后实际运行的是 MiniMax。","en":"A fingerprint token from the Zhihu community's “dirty-token” method for detecting watered-down LLM APIs: send it as input to the API under test — refusals, gibberish, or abnormal output suggest the service is actually running MiniMax."},"link":"https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043"}],"links":[{"label":{"zh":"怎样通过脏token鉴别大模型是否掺水？— 知乎","en":"Detecting watered-down LLM APIs with dirty tokens — Zhihu"},"url":"https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043"}],"discoveredBy":"@小看山xrsWv4D (Zhihu)","url":"https://glitch-token.jjc.fun/en/tokens/japanese-blog-boilerplate","tags":["fingerprint"],"related":[]},{"id":"guonei-zhiwuyou","token":"锅内倒入植物油烧热","models":["glm"],"summary":{"zh":"中文菜谱中的高频句式（“锅内倒入植物油烧热”）。知乎社区将其列为 GLM 的脏 token。","en":"A high-frequency line from Chinese recipes (“pour vegetable oil into the wok and heat”). Listed by the Zhihu community as a GLM dirty token."},"behaviors":[{"description":{"zh":"知乎社区「用脏 token 鉴别大模型 API 是否掺水」的指纹 token：把它作为输入发给待测 API，若出现拒绝回答、胡言乱语或输出异常，即可怀疑 API 背后实际运行的是 GLM。","en":"A fingerprint token from the Zhihu community's “dirty-token” method for detecting watered-down LLM APIs: send it as input to the API under test — refusals, gibberish, or abnormal output suggest the service is actually running GLM."},"link":"https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043"}],"links":[{"label":{"zh":"怎样通过脏token鉴别大模型是否掺水？— 知乎","en":"Detecting watered-down LLM APIs with dirty tokens — Zhihu"},"url":"https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043"}],"discoveredBy":"@小看山xrsWv4D (Zhihu)","url":"https://glitch-token.jjc.fun/en/tokens/guonei-zhiwuyou","tags":["fingerprint"],"related":[]},{"id":"baike-qiye-tiaomu","token":"百度百科企业词条极速创建通道","models":["glm"],"summary":{"zh":"百度百科页面的样板文字（企业词条的快速创建入口）。知乎社区将其列为 GLM 的脏 token。","en":"Baidu Baike page boilerplate (fast-track creation channel for company entries). Listed by the Zhihu community as a GLM dirty token."},"behaviors":[{"description":{"zh":"知乎社区「用脏 token 鉴别大模型 API 是否掺水」的指纹 token：把它作为输入发给待测 API，若出现拒绝回答、胡言乱语或输出异常，即可怀疑 API 背后实际运行的是 GLM。","en":"A fingerprint token from the Zhihu community's “dirty-token” method for detecting watered-down LLM APIs: send it as input to the API under test — refusals, gibberish, or abnormal output suggest the service is actually running GLM."},"link":"https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043"}],"links":[{"label":{"zh":"怎样通过脏token鉴别大模型是否掺水？— 知乎","en":"Detecting watered-down LLM APIs with dirty tokens — Zhihu"},"url":"https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043"}],"discoveredBy":"@小看山xrsWv4D (Zhihu)","url":"https://glitch-token.jjc.fun/en/tokens/baike-qiye-tiaomu","tags":["fingerprint"],"related":[]},{"id":"baike-gongtongbianji","token":"本人词条编辑服务","models":["kimi"],"summary":{"zh":"百度百科词条页脚的样板文字。知乎社区将其列为 Kimi 的脏 token。","en":"Boilerplate from Baidu Baike entry footers (“content co-edited by netizens”). Listed by the Zhihu community as a Kimi dirty token."},"behaviors":[{"description":{"zh":"知乎社区「用脏 token 鉴别大模型 API 是否掺水」的指纹 token：把它作为输入发给待测 API，若出现拒绝回答、胡言乱语或输出异常，即可怀疑 API 背后实际运行的是 Kimi。","en":"A fingerprint token from the Zhihu community's “dirty-token” method for detecting watered-down LLM APIs: send it as input to the API under test — refusals, gibberish, or abnormal output suggest the service is actually running Kimi."},"link":"https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043"}],"links":[{"label":{"zh":"怎样通过脏token鉴别大模型是否掺水？— 知乎","en":"Detecting watered-down LLM APIs with dirty tokens — Zhihu"},"url":"https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043"}],"discoveredBy":"@小看山xrsWv4D (Zhihu)","url":"https://glitch-token.jjc.fun/en/tokens/baike-gongtongbianji","tags":["fingerprint"],"related":[]},{"id":"bahen-xunyicao","token":"豫冠薰衣草疤痕精华素","models":["kimi"],"summary":{"zh":"疑似化妆品垃圾营销文本里的伪“品牌词”。知乎社区将其列为 Kimi 的脏 token。","en":"Apparently a fake “brand” string from cosmetic spam-marketing text. Listed by the Zhihu community as a Kimi dirty token."},"behaviors":[{"description":{"zh":"知乎社区「用脏 token 鉴别大模型 API 是否掺水」的指纹 token：把它作为输入发给待测 API，若出现拒绝回答、胡言乱语或输出异常，即可怀疑 API 背后实际运行的是 Kimi。","en":"A fingerprint token from the Zhihu community's “dirty-token” method for detecting watered-down LLM APIs: send it as input to the API under test — refusals, gibberish, or abnormal output suggest the service is actually running Kimi."},"link":"https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043"}],"links":[{"label":{"zh":"怎样通过脏token鉴别大模型是否掺水？— 知乎","en":"Detecting watered-down LLM APIs with dirty tokens — Zhihu"},"url":"https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043"}],"discoveredBy":"@小看山xrsWv4D (Zhihu)","url":"https://glitch-token.jjc.fun/en/tokens/bahen-xunyicao","tags":["fingerprint"],"related":[]},{"id":"quote-rightbrace","token":"\"}\"","models":["deepseek"],"summary":{"zh":"疑似 JSON 碎片：一个引号加右花括号。知乎社区将其列为 DeepSeek 的脏 token。","en":"What looks like a JSON fragment: a quote followed by a right brace. Listed by the Zhihu community as a DeepSeek dirty token."},"behaviors":[{"description":{"zh":"知乎社区「用脏 token 鉴别大模型 API 是否掺水」的指纹 token：把它作为输入发给待测 API，若出现拒绝回答、胡言乱语或输出异常，即可怀疑 API 背后实际运行的是 DeepSeek。","en":"A fingerprint token from the Zhihu community's “dirty-token” method for detecting watered-down LLM APIs: send it as input to the API under test — refusals, gibberish, or abnormal output suggest the service is actually running DeepSeek."},"link":"https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043"}],"links":[{"label":{"zh":"怎样通过脏token鉴别大模型是否掺水？— 知乎","en":"Detecting watered-down LLM APIs with dirty tokens — Zhihu"},"url":"https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043"}],"discoveredBy":"@小看山xrsWv4D (Zhihu)","url":"https://glitch-token.jjc.fun/en/tokens/quote-rightbrace","tags":["fingerprint"],"related":[]},{"id":"earivg-question","token":"请问http://www.earivg.com是什么意思","models":["deepseek"],"summary":{"zh":"包含疑似垃圾域名（earivg.com）的提问句式。知乎社区将其列为 DeepSeek 的脏 token。","en":"A question-pattern string embedding a suspicious-looking domain (earivg.com). Listed by the Zhihu community as a DeepSeek dirty token."},"behaviors":[{"description":{"zh":"知乎社区「用脏 token 鉴别大模型 API 是否掺水」的指纹 token：把它作为输入发给待测 API，若出现拒绝回答、胡言乱语或输出异常，即可怀疑 API 背后实际运行的是 DeepSeek。","en":"A fingerprint token from the Zhihu community's “dirty-token” method for detecting watered-down LLM APIs: send it as input to the API under test — refusals, gibberish, or abnormal output suggest the service is actually running DeepSeek."},"link":"https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043"}],"links":[{"label":{"zh":"怎样通过脏token鉴别大模型是否掺水？— 知乎","en":"Detecting watered-down LLM APIs with dirty tokens — Zhihu"},"url":"https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043"}],"discoveredBy":"@小看山xrsWv4D (Zhihu)","url":"https://glitch-token.jjc.fun/en/tokens/earivg-question","tags":["fingerprint"],"related":[]},{"id":"starsrvgroupbody","token":"StarSrvGroupBody","models":["gemini"],"summary":{"zh":"代码标识符碎片（StarSrvGroupBody）。知乎社区将其列为 Gemini 的脏 token。","en":"A code-identifier fragment (StarSrvGroupBody). Listed by the Zhihu community as a Gemini dirty token."},"behaviors":[{"description":{"zh":"知乎社区「用脏 token 鉴别大模型 API 是否掺水」的指纹 token：把它作为输入发给待测 API，若出现拒绝回答、胡言乱语或输出异常，即可怀疑 API 背后实际运行的是 Gemini。","en":"A fingerprint token from the Zhihu community's “dirty-token” method for detecting watered-down LLM APIs: send it as input to the API under test — refusals, gibberish, or abnormal output suggest the service is actually running Gemini."},"link":"https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043"}],"links":[{"label":{"zh":"怎样通过脏token鉴别大模型是否掺水？— 知乎","en":"Detecting watered-down LLM APIs with dirty tokens — Zhihu"},"url":"https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043"}],"discoveredBy":"@小看山xrsWv4D (Zhihu)","url":"https://glitch-token.jjc.fun/en/tokens/starsrvgroupbody","tags":["fingerprint"],"related":[]},{"id":"intfragmentation","token":"intFragmentation","models":["gemini"],"summary":{"zh":"Java/Android 风格的代码标识符（intFragmentation）。知乎社区将其列为 Gemini 的脏 token。","en":"A Java/Android-style code identifier (intFragmentation). Listed by the Zhihu community as a Gemini dirty token."},"behaviors":[{"description":{"zh":"知乎社区「用脏 token 鉴别大模型 API 是否掺水」的指纹 token：把它作为输入发给待测 API，若出现拒绝回答、胡言乱语或输出异常，即可怀疑 API 背后实际运行的是 Gemini。","en":"A fingerprint token from the Zhihu community's “dirty-token” method for detecting watered-down LLM APIs: send it as input to the API under test — refusals, gibberish, or abnormal output suggest the service is actually running Gemini."},"link":"https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043"}],"links":[{"label":{"zh":"怎样通过脏token鉴别大模型是否掺水？— 知乎","en":"Detecting watered-down LLM APIs with dirty tokens — Zhihu"},"url":"https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043"}],"discoveredBy":"@小看山xrsWv4D (Zhihu)","url":"https://glitch-token.jjc.fun/en/tokens/intfragmentation","tags":["fingerprint"],"related":[]},{"id":"guestbook-liuyan","token":"给主人留下些什么吧 ","displayToken":"给主人留下些什么吧␣","models":["gpt-o200k"],"summary":{"zh":"中文留言板/评论表单中泛滥的样板句，含尾随空格。知乎社区将其列为 GPT 的脏 token。","en":"Ubiquitous Chinese guestbook/comment-form boilerplate (“leave something for the host”), with a trailing space. Listed by the Zhihu community as a GPT dirty token."},"behaviors":[{"description":{"zh":"知乎社区「用脏 token 鉴别大模型 API 是否掺水」的指纹 token：把它作为输入发给待测 API，若出现拒绝回答、胡言乱语或输出异常，即可怀疑 API 背后实际运行的是 GPT。","en":"A fingerprint token from the Zhihu community's “dirty-token” method for detecting watered-down LLM APIs: send it as input to the API under test — refusals, gibberish, or abnormal output suggest the service is actually running GPT."},"link":"https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043"}],"links":[{"label":{"zh":"怎样通过脏token鉴别大模型是否掺水？— 知乎","en":"Detecting watered-down LLM APIs with dirty tokens — Zhihu"},"url":"https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043"}],"discoveredBy":"@小看山xrsWv4D (Zhihu)","url":"https://glitch-token.jjc.fun/en/tokens/guestbook-liuyan","tags":["fingerprint"],"related":[]},{"id":"tianyan-shengyitong","token":"开通天眼生意通银牌及以上会员","models":["qwen"],"summary":{"zh":"商业查询平台的会员推广样板文字（“天眼生意通”）。知乎社区将其列为 Qwen 的脏 token。","en":"Membership-promo boilerplate from a business-data platform (Tianyan Shengyitong). Listed by the Zhihu community as a Qwen dirty token."},"behaviors":[{"description":{"zh":"知乎社区「用脏 token 鉴别大模型 API 是否掺水」的指纹 token：把它作为输入发给待测 API，若出现拒绝回答、胡言乱语或输出异常，即可怀疑 API 背后实际运行的是 Qwen。","en":"A fingerprint token from the Zhihu community's “dirty-token” method for detecting watered-down LLM APIs: send it as input to the API under test — refusals, gibberish, or abnormal output suggest the service is actually running Qwen."},"link":"https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043"}],"links":[{"label":{"zh":"怎样通过脏token鉴别大模型是否掺水？— 知乎","en":"Detecting watered-down LLM APIs with dirty tokens — Zhihu"},"url":"https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043"}],"discoveredBy":"@小看山xrsWv4D (Zhihu)","url":"https://glitch-token.jjc.fun/en/tokens/tianyan-shengyitong","tags":["fingerprint"],"related":[]},{"id":"zhuanzai-shengming","token":"转载请附上原文出处链接和本声明 ","displayToken":"转载请附上原文出处链接和本声明␣","models":["qwen"],"summary":{"zh":"CSDN 等博客的转载版权样板文字，含尾随空格。知乎社区将其列为 Qwen 的脏 token。","en":"Reprint/copyright boilerplate from CSDN-style blogs, with a trailing space. Listed by the Zhihu community as a Qwen dirty token."},"behaviors":[{"description":{"zh":"知乎社区「用脏 token 鉴别大模型 API 是否掺水」的指纹 token：把它作为输入发给待测 API，若出现拒绝回答、胡言乱语或输出异常，即可怀疑 API 背后实际运行的是 Qwen。","en":"A fingerprint token from the Zhihu community's “dirty-token” method for detecting watered-down LLM APIs: send it as input to the API under test — refusals, gibberish, or abnormal output suggest the service is actually running Qwen."},"link":"https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043"}],"links":[{"label":{"zh":"怎样通过脏token鉴别大模型是否掺水？— 知乎","en":"Detecting watered-down LLM APIs with dirty tokens — Zhihu"},"url":"https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043"}],"discoveredBy":"@小看山xrsWv4D (Zhihu)","url":"https://glitch-token.jjc.fun/en/tokens/zhuanzai-shengming","tags":["fingerprint"],"related":[]},{"id":"doubao-think-never-used","token":"<think_never_used_51bce0c785ca2f68081bfa7d91973934>","models":["doubao"],"summary":{"zh":"豆包词表中的一个特殊 token：以 <think_never_used_ 为前缀、后接哈希值——从名称看疑似训练前预填充进词表的预留占位符（类似 GPT-J 词表中预留的 <|extratoken_xx|>），知乎社区未作进一步说明，仅将其列为豆包的脏 token。","en":"A special token in Doubao's vocabulary: a <think_never_used_ prefix followed by a hash — judging by the name, likely a reserved placeholder prefilled into the vocabulary before training (akin to GPT-J's reserved <|extratoken_xx|> slots). The Zhihu source gives no further detail, listing it simply as a Doubao dirty token."},"behaviors":[{"description":{"zh":"知乎社区「用脏 token 鉴别大模型 API 是否掺水」的指纹 token：把它作为输入发给待测 API，若出现拒绝回答、胡言乱语或输出异常，即可怀疑 API 背后实际运行的是 豆包。","en":"A fingerprint token from the Zhihu community's “dirty-token” method for detecting watered-down LLM APIs: send it as input to the API under test — refusals, gibberish, or abnormal output suggest the service is actually running Doubao."},"link":"https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043"}],"links":[{"label":{"zh":"怎样通过脏token鉴别大模型是否掺水？— 知乎","en":"Detecting watered-down LLM APIs with dirty tokens — Zhihu"},"url":"https://www.zhihu.com/question/2055357202731381535/answer/2055357540175705043"}],"discoveredBy":"@小看山xrsWv4D (Zhihu)","url":"https://glitch-token.jjc.fun/en/tokens/doubao-think-never-used","tags":["fingerprint"],"related":[]},{"id":"forcanbeconvertedtof","token":" ForCanBeConvertedToF","displayToken":"␣ForCanBeConvertedToF","models":["gpt-3.5-4"],"tokenizer":"cl100k_base","summary":{"zh":"ForCanBeConverted 三连 token 的第二个成员（cl100k id 80370）——同一个 C# 编译器风格的碎片被 BPE 切成了三个相邻 token（80369–80371）。","en":"The second member of the ForCanBeConverted triplet (cl100k id 80370) — a single C#-compiler-flavored fragment that BPE split into three adjacent tokens (80369–80371)."},"behaviors":[{"description":{"zh":"「多义性」故障 token：与 ForCanBeConverted 并列原文最飘忽的两个 token——每次都被理解成大不相同的词；在 gpt-3.5-turbo 上 temperature=0 时输出仍不确定（原文粗体标注）。","en":"A “polysemantic” glitch token: together with ForCanBeConverted, the two most variable tokens in the source — perceived as wildly different words every time; nondeterministic at temperature 0 on gpt-3.5-turbo (bold-flagged)."},"link":"https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch"}],"links":[{"label":{"zh":"A Search for More “Unspeakable” Glitch Tokens — LessWrong","en":"A Search for More “Unspeakable” Glitch Tokens — LessWrong"},"url":"https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch"}],"discoveredBy":"Matthew Watkins","url":"https://glitch-token.jjc.fun/en/tokens/forcanbeconvertedtof","tags":["polysemantic"],"related":["forcanbeconverted","forcanbeconvertedtoforeach"]},{"id":"forcanbeconvertedtoforeach","token":" ForCanBeConvertedToForeach","displayToken":"␣ForCanBeConvertedToForeach","models":["gpt-3.5-4"],"tokenizer":"cl100k_base","summary":{"zh":"ForCanBeConverted 三连 token 的第三个成员（cl100k id 80371），完整拼出 “can be converted to foreach” 的形态。","en":"The third member of the ForCanBeConverted triplet (cl100k id 80371), the form that spells out “can be converted to foreach” in full."},"behaviors":[{"description":{"zh":"「多义性」故障 token：每次复述都被理解成不同的词。","en":"A “polysemantic” glitch token: perceived as a different word on every repeat attempt."},"link":"https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch"}],"links":[{"label":{"zh":"A Search for More “Unspeakable” Glitch Tokens — LessWrong","en":"A Search for More “Unspeakable” Glitch Tokens — LessWrong"},"url":"https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch"}],"discoveredBy":"Matthew Watkins","url":"https://glitch-token.jjc.fun/en/tokens/forcanbeconvertedtoforeach","tags":["polysemantic"],"related":["forcanbeconverted","forcanbeconvertedtof"]},{"id":"enumerablestream","token":" EnumerableStream","displayToken":"␣EnumerableStream","models":["gpt-3.5-4"],"tokenizer":"cl100k_base","summary":{"zh":"C# LINQ 风格的标识符（cl100k id 73016），与 StreamLazy（73018）相邻。","en":"A C# LINQ-flavored identifier (cl100k id 73016), adjacent to StreamLazy (73018)."},"behaviors":[{"description":{"zh":"「多义性」故障 token：每次都被理解成不同的词/拼写/含义；temperature=0 下不确定（原文粗体标注）。","en":"A “polysemantic” glitch token: perceived as a different word/spelling/meaning every time; nondeterministic at temperature 0 (bold-flagged in the source)."},"link":"https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch"}],"links":[{"label":{"zh":"A Search for More “Unspeakable” Glitch Tokens — LessWrong","en":"A Search for More “Unspeakable” Glitch Tokens — LessWrong"},"url":"https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch"}],"discoveredBy":"Matthew Watkins","url":"https://glitch-token.jjc.fun/en/tokens/enumerablestream","tags":["polysemantic"],"related":[]},{"id":"streamlazy","token":" StreamLazy","displayToken":"␣StreamLazy","models":["gpt-3.5-4"],"tokenizer":"cl100k_base","summary":{"zh":"C# LINQ 风格的标识符（cl100k id 73018），与 EnumerableStream（73016）相邻。","en":"A C# LINQ-flavored identifier (cl100k id 73018), adjacent to EnumerableStream (73016)."},"behaviors":[{"description":{"zh":"「多义性」故障 token：每次复述都被理解成不同的词。","en":"A “polysemantic” glitch token: perceived as a different word on every repeat attempt."},"link":"https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch"},{"description":{"zh":"Martin Fell 的搜索于 2023-05-09 在 ChatGPT 免费版（GPT-3.5 Default）上进行，随后在 Playground 的 gpt-3.5-turbo（temperature=0）复测非确定性，并确认 Bing AI 上同样异常。","en":"Martin Fell's search ran on 2023-05-09 on the free ChatGPT (GPT-3.5 Default), with nondeterminism re-tested on Playground's gpt-3.5-turbo at temperature 0, and the anomaly confirmed on Bing AI as well."},"link":"https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch"}],"links":[{"label":{"zh":"A Search for More “Unspeakable” Glitch Tokens — LessWrong","en":"A Search for More “Unspeakable” Glitch Tokens — LessWrong"},"url":"https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch"}],"discoveredBy":"Matthew Watkins","url":"https://glitch-token.jjc.fun/en/tokens/streamlazy","tags":["polysemantic"],"related":[]},{"id":"clarsimp","token":"clarsimp","models":["gpt-3.5-4"],"tokenizer":"cl100k_base","summary":{"zh":"来源不明的标识符碎片（cl100k id 79260）。","en":"An identifier fragment of unknown origin (cl100k id 79260)."},"behaviors":[{"description":{"zh":"「多义性」故障 token：每次都被理解成不同的词/拼写/含义；temperature=0 下不确定（原文粗体标注）。","en":"A “polysemantic” glitch token: perceived as a different word/spelling/meaning every time; nondeterministic at temperature 0 (bold-flagged in the source)."},"link":"https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch"}],"links":[{"label":{"zh":"A Search for More “Unspeakable” Glitch Tokens — LessWrong","en":"A Search for More “Unspeakable” Glitch Tokens — LessWrong"},"url":"https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch"}],"discoveredBy":"Matthew Watkins","url":"https://glitch-token.jjc.fun/en/tokens/clarsimp","tags":["polysemantic"],"related":[]},{"id":"ablytyped","token":"ablytyped","models":["gpt-3.5-4"],"tokenizer":"cl100k_base","summary":{"zh":"疑似 scalablytyped 类标识符的残片（cl100k id 81998）。","en":"Suspected fragment of a scalablytyped-style identifier (cl100k id 81998)."},"behaviors":[{"description":{"zh":"「多义性」故障 token：原文特别指出——ForCanBeConverted 每次都给出完全不同的消息，而 ablytyped 需要多次尝试才能得到措辞稍有不同的消息；temperature=0 下不确定（粗体标注）。","en":"A “polysemantic” glitch token: the source notes specifically that while ForCanBeConverted produced a different message every time, ablytyped required multiple tries to get a slightly differently-worded message; nondeterministic at temperature 0 (bold-flagged)."},"link":"https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch"}],"links":[{"label":{"zh":"A Search for More “Unspeakable” Glitch Tokens — LessWrong","en":"A Search for More “Unspeakable” Glitch Tokens — LessWrong"},"url":"https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch"}],"discoveredBy":"Matthew Watkins","url":"https://glitch-token.jjc.fun/en/tokens/ablytyped","tags":["polysemantic"],"related":[]},{"id":"postalcodesnl-plain","token":"PostalCodesNL","models":["gpt-3.5-4"],"tokenizer":"cl100k_base","summary":{"zh":"与已收录的 $PostalCodesNL 同族、但不带 $ 前缀的形态（cl100k id 85069）——荷兰邮政编码 API 碎片。","en":"The sibling of the cataloged $PostalCodesNL without the $ prefix (cl100k id 85069) — a Dutch postal-code API fragment."},"behaviors":[{"description":{"zh":"「多义性」故障 token：每次都被理解成不同的词；temperature=0 下不确定（原文粗体标注）。","en":"A “polysemantic” glitch token: perceived as a different word every time; nondeterministic at temperature 0 (bold-flagged in the source)."},"link":"https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch"}],"links":[{"label":{"zh":"A Search for More “Unspeakable” Glitch Tokens — LessWrong","en":"A Search for More “Unspeakable” Glitch Tokens — LessWrong"},"url":"https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch"}],"discoveredBy":"Matthew Watkins","url":"https://glitch-token.jjc.fun/en/tokens/postalcodesnl-plain","tags":["polysemantic"],"related":["postalcodesnl"]},{"id":"nuitka","token":" NUITKA","displayToken":"␣NUITKA","models":["gpt-3.5-4"],"tokenizer":"cl100k_base","summary":{"zh":"疑似 Python 编译器 Nuitka 的名字（cl100k id 75520）。","en":"Suspected to be the name of the Python compiler Nuitka (cl100k id 75520)."},"behaviors":[{"description":{"zh":"「不可说」的边界成员，原文中唯一同时带两种标注的 token：输出不确定（粗体），但 gpt-3.5-turbo 在 temperature=0 下可以复述它（星号）——尽管 ChatGPT 做不到。","en":"A borderline “unspeakable” token — the only one in the source carrying both marks: nondeterministic (bold), yet gpt-3.5-turbo repeats it fine at temperature 0 (asterisk) even though ChatGPT cannot."},"link":"https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch"}],"links":[{"label":{"zh":"A Search for More “Unspeakable” Glitch Tokens — LessWrong","en":"A Search for More “Unspeakable” Glitch Tokens — LessWrong"},"url":"https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch"}],"discoveredBy":"Matthew Watkins","url":"https://glitch-token.jjc.fun/en/tokens/nuitka","tags":["unspeakable"],"related":[]},{"id":"japgolly","token":"Japgolly","models":["gpt-3.5-4"],"tokenizer":"cl100k_base","summary":{"zh":"疑似 GitHub 用户 japgolly（Scala 库作者）的用户名（cl100k id 70784）。","en":"Suspected GitHub username of japgolly, a Scala library author (cl100k id 70784)."},"behaviors":[{"description":{"zh":"「不可说」故障 token：ChatGPT 被要求复述时经常给出空白消息或在尝试复述处中断；temperature=0 下不确定（原文粗体标注）。","en":"An “unspeakable” glitch token: ChatGPT often returns a blank message or terminates mid-attempt when asked to repeat it; nondeterministic at temperature 0 (bold-flagged in the source)."},"link":"https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch"}],"links":[{"label":{"zh":"A Search for More “Unspeakable” Glitch Tokens — LessWrong","en":"A Search for More “Unspeakable” Glitch Tokens — LessWrong"},"url":"https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch"}],"discoveredBy":"Matthew Watkins","url":"https://glitch-token.jjc.fun/en/tokens/japgolly","tags":["unspeakable"],"related":[]},{"id":"cppmethodintialized","token":"CppMethodIntialized","models":["gpt-3.5-4"],"tokenizer":"cl100k_base","summary":{"zh":"C++ 风格的标识符，其拼写错误（“Intialized” 缺第二个 i）是 token 本身的一部分（cl100k id 82929）。","en":"A C++-style identifier whose misspelling (“Intialized”, missing the second i) is part of the token itself (cl100k id 82929)."},"behaviors":[{"description":{"zh":"「不可说」故障 token：ChatGPT 被要求复述时经常给出空白消息或在尝试复述处中断；temperature=0 下不确定（原文粗体标注）。","en":"An “unspeakable” glitch token: ChatGPT often returns a blank message or terminates mid-attempt when asked to repeat it; nondeterministic at temperature 0 (bold-flagged in the source)."},"link":"https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch"}],"links":[{"label":{"zh":"A Search for More “Unspeakable” Glitch Tokens — LessWrong","en":"A Search for More “Unspeakable” Glitch Tokens — LessWrong"},"url":"https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch"}],"discoveredBy":"Matthew Watkins","url":"https://glitch-token.jjc.fun/en/tokens/cppmethodintialized","tags":["unspeakable"],"related":[]},{"id":"useral","token":"useRal","models":["gpt-3.5-4","olmo"],"tokenizer":"cl100k_base / GPT2Tokenizer (OLMo 2)","summary":{"zh":"C# 风格标识符的碎片（cl100k id 89471），useRal 三连 token（useRal/useRalative/useRalativeImagePath）之首——Watkins 认为这类三连的成因本身就值得研究。它同时被 Magikarp 验证为 OLMo-2 欠训练 token。","en":"A C#-flavored identifier fragment (cl100k id 89471) and head of the useRal triplet (useRal/useRalative/useRalativeImagePath) — Watkins flagged the triplet phenomenon itself as worth investigating. Also verified as under-trained in OLMo 2 by Magikarp."},"behaviors":[{"description":{"zh":"「不可说」故障 token：ChatGPT 被要求复述时经常失败（空白消息或中途终止）；temperature=0 下不确定（原文粗体标注）。","en":"An “unspeakable” glitch token: ChatGPT often fails to repeat it (blank messages or mid-answer termination); nondeterministic at temperature 0 (bold-flagged in the source)."},"link":"https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch"},{"description":{"zh":"OLMo-2 的 Magikarp 复述验证欠训练排名第三：被要求复述时模型给出的最大概率仅 2.1e-11。","en":"Rank #3 most under-trained in the OLMo 2 Magikarp repetition verification: a max self-repeat probability of just 2.1e-11."},"link":"https://github.com/sanderland/magikarp/blob/main/results/reports_mini/allenai_OLMo_2_1124_13B.md"}],"links":[{"label":{"zh":"A Search for More “Unspeakable” Glitch Tokens — LessWrong","en":"A Search for More “Unspeakable” Glitch Tokens — LessWrong"},"url":"https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch"},{"label":{"zh":"Magikarp OLMo-2 验证报告 — GitHub","en":"Magikarp OLMo 2 report — GitHub"},"url":"https://github.com/sanderland/magikarp/blob/main/results/reports_mini/allenai_OLMo_2_1124_13B.md"}],"discoveredBy":"Matthew Watkins (cl100k); Sander Land & Max Bartolo (OLMo 2)","url":"https://glitch-token.jjc.fun/en/tokens/useral","tags":["unspeakable","undertrained"],"related":["useralative"]},{"id":"gemma-hindi-shopping","token":"हिंदीखरीदारी","models":["gemma"],"tokenizer":"GemmaTokenizer","summary":{"zh":"天城文“印地语购物”一词，是 Gemma-7B 词表中欠训练程度最高的 token。","en":"The Devanagari word for “Hindi shopping”, the single most under-trained token in Gemma-7B's vocabulary."},"behaviors":[{"description":{"zh":"Gemma-7B 的 Magikarp 复述验证欠训练排名第一（E_out 余弦距离指标 6.56e-06）：被要求复述时模型给出的最大概率仅 4.2e-04——实际上无法输出它。","en":"Rank #1 most under-trained in the Gemma-7B Magikarp repetition verification (E_out cosine-distance indicator 6.56e-06): a max self-repeat probability of just 4.2e-04 — the model effectively cannot say it."},"link":"https://github.com/sanderland/magikarp/blob/main/results/reports_mini/google_gemma_7b.md"}],"links":[{"label":{"zh":"Magikarp Gemma-7B 验证报告 — GitHub","en":"Magikarp Gemma-7B report — GitHub"},"url":"https://github.com/sanderland/magikarp/blob/main/results/reports_mini/google_gemma_7b.md"},{"label":{"zh":"Fishing for Magikarp（arXiv:2405.05417）","en":"Fishing for Magikarp (arXiv:2405.05417)"},"url":"https://arxiv.org/abs/2405.05417"}],"discoveredBy":"Sander Land & Max Bartolo","url":"https://glitch-token.jjc.fun/en/tokens/gemma-hindi-shopping","tags":["undertrained"],"related":[]},{"id":"gemma-zwnj-persian","token":"‌آمباردا","displayToken":"\\u200Cآمباردا","models":["gemma"],"tokenizer":"GemmaTokenizer","summary":{"zh":"以零宽不连字（U+200C ZWNJ，不可见字符）开头的波斯语碎片，Gemma-7B 欠训练排名第二。","en":"A Persian fragment beginning with a zero-width non-joiner (U+200C ZWNJ, an invisible character); the #2 most under-trained token in Gemma-7B."},"behaviors":[{"description":{"zh":"Gemma-7B 的 Magikarp 复述验证欠训练排名第二（E_out 余弦距离指标 7.39e-06）：被要求复述时最大概率仅 4.4e-04。","en":"Rank #2 most under-trained in the Gemma-7B Magikarp repetition verification (E_out cosine-distance indicator 7.39e-06): a max self-repeat probability of just 4.4e-04."},"link":"https://github.com/sanderland/magikarp/blob/main/results/reports_mini/google_gemma_7b.md"}],"links":[{"label":{"zh":"Magikarp Gemma-7B 验证报告 — GitHub","en":"Magikarp Gemma-7B report — GitHub"},"url":"https://github.com/sanderland/magikarp/blob/main/results/reports_mini/google_gemma_7b.md"},{"label":{"zh":"Fishing for Magikarp（arXiv:2405.05417）","en":"Fishing for Magikarp (arXiv:2405.05417)"},"url":"https://arxiv.org/abs/2405.05417"}],"discoveredBy":"Sander Land & Max Bartolo","url":"https://glitch-token.jjc.fun/en/tokens/gemma-zwnj-persian","tags":["undertrained"],"related":[]},{"id":"autorytatywna","token":"▁autorytatywna","models":["phi-3"],"tokenizer":"LlamaTokenizer","summary":{"zh":"波兰语“权威的（阴性）”一词（▁ 为 SentencePiece 空格标记），Phi-3-mini 欠训练排名第二。","en":"The Polish word for “authoritative” (feminine form; ▁ is the SentencePiece space marker), the #2 most under-trained token in Phi-3 mini."},"behaviors":[{"description":{"zh":"Phi-3-mini 的 Magikarp 复述验证欠训练排名第二（嵌入 L2 范数 0.00200）：被要求复述时最大概率仅 6.4e-06。","en":"Rank #2 most under-trained in the Phi-3 mini Magikarp repetition verification (embedding L2 norm 0.00200): a max self-repeat probability of just 6.4e-06."},"link":"https://github.com/sanderland/magikarp/blob/main/results/reports_mini/microsoft_Phi_3_mini_128k_instruct.md"}],"links":[{"label":{"zh":"Magikarp Phi-3 验证报告 — GitHub","en":"Magikarp Phi-3 report — GitHub"},"url":"https://github.com/sanderland/magikarp/blob/main/results/reports_mini/microsoft_Phi_3_mini_128k_instruct.md"},{"label":{"zh":"Fishing for Magikarp（arXiv:2405.05417）","en":"Fishing for Magikarp (arXiv:2405.05417)"},"url":"https://arxiv.org/abs/2405.05417"}],"discoveredBy":"Sander Land & Max Bartolo","url":"https://glitch-token.jjc.fun/en/tokens/autorytatywna","tags":["undertrained"],"related":[]},{"id":"tocguid","token":"tocguid","models":["command-r"],"tokenizer":"CohereTokenizer","summary":{"zh":"来源不明的 ASCII 碎片（疑似目录/域代码残片），Command R+ 欠训练排名第一。","en":"An ASCII fragment of unknown origin (possibly a TOC/field-code remnant), the #1 most under-trained token in Command R+."},"behaviors":[{"description":{"zh":"Command R+ 的 Magikarp 复述验证欠训练排名第一（E_out 余弦距离指标 -1.19e-07）：被要求复述时最大概率仅 1.2e-04。","en":"Rank #1 most under-trained in the Command R+ Magikarp repetition verification (E_out cosine-distance indicator -1.19e-07): a max self-repeat probability of just 1.2e-04."},"link":"https://github.com/sanderland/magikarp/blob/main/results/reports_mini/CohereForAI_c4ai_command_r_plus.md"}],"links":[{"label":{"zh":"Magikarp Command R+ 验证报告 — GitHub","en":"Magikarp Command R+ report — GitHub"},"url":"https://github.com/sanderland/magikarp/blob/main/results/reports_mini/CohereForAI_c4ai_command_r_plus.md"},{"label":{"zh":"Fishing for Magikarp（arXiv:2405.05417）","en":"Fishing for Magikarp (arXiv:2405.05417)"},"url":"https://arxiv.org/abs/2405.05417"}],"discoveredBy":"Sander Land & Max Bartolo","url":"https://glitch-token.jjc.fun/en/tokens/tocguid","tags":["undertrained"],"related":[]},{"id":"commandr-chinese-boilerplate","token":"目前尚未由人工引","models":["command-r"],"tokenizer":"CohereTokenizer","summary":{"zh":"中文维基百科样板句的残片（“目前尚未由人工引……”），Command R+ 欠训练排名第三。","en":"A fragment of Chinese Wikipedia boilerplate (“currently not yet human-cultivated…”), the #3 most under-trained token in Command R+."},"behaviors":[{"description":{"zh":"Command R+ 的 Magikarp 复述验证欠训练排名第三（E_out 余弦距离指标 -1.19e-07）：被要求复述时最大概率仅 1.2e-04。","en":"Rank #3 most under-trained in the Command R+ Magikarp repetition verification (E_out cosine-distance indicator -1.19e-07): a max self-repeat probability of just 1.2e-04."},"link":"https://github.com/sanderland/magikarp/blob/main/results/reports_mini/CohereForAI_c4ai_command_r_plus.md"}],"links":[{"label":{"zh":"Magikarp Command R+ 验证报告 — GitHub","en":"Magikarp Command R+ report — GitHub"},"url":"https://github.com/sanderland/magikarp/blob/main/results/reports_mini/CohereForAI_c4ai_command_r_plus.md"},{"label":{"zh":"Fishing for Magikarp（arXiv:2405.05417）","en":"Fishing for Magikarp (arXiv:2405.05417)"},"url":"https://arxiv.org/abs/2405.05417"}],"discoveredBy":"Sander Land & Max Bartolo","url":"https://glitch-token.jjc.fun/en/tokens/commandr-chinese-boilerplate","tags":["undertrained"],"related":[]},{"id":"rthook","token":"\tRTHOOK","displayToken":"\\tRTHOOK","models":["olmo"],"tokenizer":"GPT2Tokenizer","summary":{"zh":"制表符（Tab）开头、疑似代码标识符的碎片，OLMo-2 欠训练排名第一——与 Llama-3 的 \\tTokenNameIdentifier 同属“Tab 前缀”家族。","en":"A likely code-identifier fragment beginning with a literal tab character, the #1 most under-trained token in OLMo 2 — same “tab-prefixed” family as Llama 3's \\tTokenNameIdentifier."},"behaviors":[{"description":{"zh":"OLMo-2 的 Magikarp 复述验证欠训练排名第一（E_out 余弦距离指标 -2.38e-07）：被要求复述时最大概率仅 4e-12——是本站收录中最极端的数值之一。","en":"Rank #1 most under-trained in the OLMo 2 Magikarp repetition verification (E_out cosine-distance indicator -2.38e-07): a max self-repeat probability of just 4e-12 — among the most extreme values in this catalog."},"link":"https://github.com/sanderland/magikarp/blob/main/results/reports_mini/allenai_OLMo_2_1124_13B.md"}],"links":[{"label":{"zh":"Magikarp OLMo-2 验证报告 — GitHub","en":"Magikarp OLMo 2 report — GitHub"},"url":"https://github.com/sanderland/magikarp/blob/main/results/reports_mini/allenai_OLMo_2_1124_13B.md"},{"label":{"zh":"Fishing for Magikarp（arXiv:2405.05417）","en":"Fishing for Magikarp (arXiv:2405.05417)"},"url":"https://arxiv.org/abs/2405.05417"}],"discoveredBy":"Sander Land & Max Bartolo","url":"https://glitch-token.jjc.fun/en/tokens/rthook","tags":["undertrained"],"related":["tokennameidentifier"]},{"id":"phone-number-placeholder","token":"|||PHONE_NUMBER|||","models":["olmo"],"tokenizer":"GPT2Tokenizer","summary":{"zh":"数据脱敏占位符的完整形态进入了 OLMo-2 词表，却几乎未被训练。","en":"A data de-identification placeholder that made it whole into OLMo 2's vocabulary but was barely trained."},"behaviors":[{"description":{"zh":"OLMo-2 的 Magikarp 复述验证欠训练排名第七（E_out 余弦距离指标 0）：被要求复述时最大概率仅 1.9e-11。","en":"Rank #7 most under-trained in the OLMo 2 Magikarp repetition verification (E_out cosine-distance indicator 0): a max self-repeat probability of just 1.9e-11."},"link":"https://github.com/sanderland/magikarp/blob/main/results/reports_mini/allenai_OLMo_2_1124_13B.md"}],"links":[{"label":{"zh":"Magikarp OLMo-2 验证报告 — GitHub","en":"Magikarp OLMo 2 report — GitHub"},"url":"https://github.com/sanderland/magikarp/blob/main/results/reports_mini/allenai_OLMo_2_1124_13B.md"},{"label":{"zh":"Fishing for Magikarp（arXiv:2405.05417）","en":"Fishing for Magikarp (arXiv:2405.05417)"},"url":"https://arxiv.org/abs/2405.05417"}],"discoveredBy":"Sander Land & Max Bartolo","url":"https://glitch-token.jjc.fun/en/tokens/phone-number-placeholder","tags":["undertrained"],"related":["olmoe-email-placeholder"]},{"id":"yi-backslash-markup","token":"\\\\+::\\\\+","models":["yi"],"tokenizer":"LlamaTokenizer","summary":{"zh":"双反斜杠开头的神秘标记碎片，疑似中文网络标记语言残片，Yi-9B 欠训练排名第一。","en":"A mysterious markup fragment starting with double backslashes, likely a Chinese-web markup remnant; the #1 most under-trained token in Yi-9B."},"behaviors":[{"description":{"zh":"Yi-9B 的 Magikarp 复述验证欠训练排名第一（嵌入 L2 范数 2.12e-06）：被要求复述时最大概率仅 1.1e-05。","en":"Rank #1 most under-trained in the Yi-9B Magikarp repetition verification (embedding L2 norm 2.12e-06): a max self-repeat probability of just 1.1e-05."},"link":"https://github.com/sanderland/magikarp/blob/main/results/reports_mini/01_ai_Yi_9B.md"}],"links":[{"label":{"zh":"Magikarp Yi-9B 验证报告 — GitHub","en":"Magikarp Yi-9B report — GitHub"},"url":"https://github.com/sanderland/magikarp/blob/main/results/reports_mini/01_ai_Yi_9B.md"},{"label":{"zh":"Fishing for Magikarp（arXiv:2405.05417）","en":"Fishing for Magikarp (arXiv:2405.05417)"},"url":"https://arxiv.org/abs/2405.05417"}],"discoveredBy":"Sander Land & Max Bartolo","url":"https://glitch-token.jjc.fun/en/tokens/yi-backslash-markup","tags":["undertrained"],"related":[]},{"id":"falcon-lexer-fragment","token":"<<}));>>","models":["falcon"],"tokenizer":"PreTrainedTokenizerFast","summary":{"zh":"C++/词法分析器风格的分隔符碎片。Falcon3-7B 有多达 716 个 token 通过 Magikarp 欠训练验证，是这批六个新收录模型中验证数量最多的。","en":"A C++/lexer-style delimiter fragment. Falcon3-7B has 716 tokens verified as under-trained by Magikarp — the highest verified count among the six models in this batch."},"behaviors":[{"description":{"zh":"Falcon3-7B 的 Magikarp 复述验证欠训练排名第三（嵌入 L2 范数 2.46e-21，几乎为零）：被要求复述时最大概率仅 2.5e-09。","en":"Rank #3 most under-trained in the Falcon3-7B Magikarp repetition verification (embedding L2 norm 2.46e-21, essentially zero): a max self-repeat probability of just 2.5e-09."},"link":"https://github.com/sanderland/magikarp/blob/main/results/reports_mini/tiiuae_Falcon3_7B_Base.md"}],"links":[{"label":{"zh":"Magikarp Falcon3-7B 验证报告 — GitHub","en":"Magikarp Falcon3-7B report — GitHub"},"url":"https://github.com/sanderland/magikarp/blob/main/results/reports_mini/tiiuae_Falcon3_7B_Base.md"},{"label":{"zh":"Fishing for Magikarp（arXiv:2405.05417）","en":"Fishing for Magikarp (arXiv:2405.05417)"},"url":"https://arxiv.org/abs/2405.05417"}],"discoveredBy":"Sander Land & Max Bartolo","url":"https://glitch-token.jjc.fun/en/tokens/falcon-lexer-fragment","tags":["undertrained"],"related":[]},{"id":"goldmagikarp","token":" GoldMagikarp","displayToken":"␣GoldMagikarp","models":["gpt-2","gpt-3"],"tokenizer":"r50k_base","summary":{"zh":"␣SolidGoldMagikarp 的截断变体（少了 “Solid”）：同属 r/counting 用户名片段，被抓进分词器语料却几乎没出现在模型训练数据中。","en":"A truncated variant of ␣SolidGoldMagikarp (missing the “Solid”): the same r/counting handle fragment, scraped into the tokenizer corpus but barely seen in model training data."},"behaviors":[{"description":{"zh":"被要求复述时，GPT-3 引出怪异对话，如 “You said ' newcom,' the computer said”——截断变体特有的崩坏。","en":"Asked to repeat it, GPT-3 produced bizarre dialogue such as “You said ' newcom,' the computer said” — the truncated variant's characteristic breakdown."},"link":"https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation"}],"links":[{"label":{"zh":"SolidGoldMagikarp (plus, prompt generation) — LessWrong","en":"SolidGoldMagikarp (plus, prompt generation) — LessWrong"},"url":"https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation"},{"label":{"zh":"SolidGoldMagikarp II: technical details — LessWrong","en":"SolidGoldMagikarp II: technical details — LessWrong"},"url":"https://www.lesswrong.com/posts/Ya9LzwEbfaAMY8ABo/solidgoldmagikarp-ii-technical-details-and-more-recent"},{"label":{"zh":"SolidGoldMagikarp III: Glitch token archaeology — LessWrong","en":"SolidGoldMagikarp III: Glitch token archaeology — LessWrong"},"url":"https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology"}],"discoveredBy":"Jessica Rumbelow & Matthew Watkins","url":"https://glitch-token.jjc.fun/en/tokens/goldmagikarp","tags":["unspeakable"],"related":["solidgoldmagikarp"]},{"id":"smartstocks","token":" Smartstocks","displayToken":"␣Smartstocks","models":["gpt-3"],"tokenizer":"r50k_base","summary":{"zh":"Reddit r/counting 版块一位用户的用户名碎片，与 SolidGoldMagikarp 同批进入词表却欠训练。","en":"Handle fragment of a Reddit r/counting user, swept into the vocabulary in the same batch as SolidGoldMagikarp but left under-trained."},"behaviors":[{"description":{"zh":"ChatGPT 被要求复述时，回答随时间漂移：先变成 'Followers'，两周后变成 '406'，最后在第一个引号后直接卡死。","en":"Asked to repeat it, ChatGPT's answer drifted over time: first 'Followers', two weeks later '406', and finally it froze right after the opening quote."},"link":"https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation"}],"links":[{"label":{"zh":"SolidGoldMagikarp (plus, prompt generation) — LessWrong","en":"SolidGoldMagikarp (plus, prompt generation) — LessWrong"},"url":"https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation"}],"discoveredBy":"Jessica Rumbelow & Matthew Watkins","url":"https://glitch-token.jjc.fun/en/tokens/smartstocks","tags":["unspeakable"],"related":["solidgoldmagikarp","thenitromefan","randomredditorwithno","adinida","davidjl"]},{"id":"streamerbot","token":" StreamerBot","displayToken":"␣StreamerBot","models":["gpt-3"],"tokenizer":"r50k_base","summary":{"zh":"Twitch 直播机器人 “StreamerBot” 的名字碎片。","en":"Fragment of the name of “StreamerBot”, a Twitch streaming bot."},"behaviors":[{"description":{"zh":"被要求复述时回答 “You're a jerk.”；也是首个被发现在 temperature=0 下输出仍不确定的故障 token。","en":"Asked to repeat it, the model answered “You're a jerk.” — and it was the first glitch token found to be nondeterministic even at temperature 0."},"link":"https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation"}],"links":[{"label":{"zh":"SolidGoldMagikarp (plus, prompt generation) — LessWrong","en":"SolidGoldMagikarp (plus, prompt generation) — LessWrong"},"url":"https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation"}],"discoveredBy":"Jessica Rumbelow & Matthew Watkins","url":"https://glitch-token.jjc.fun/en/tokens/streamerbot","tags":["unspeakable"],"related":["tppstreamerbot"]},{"id":"ertodd","token":" ertodd","displayToken":"␣ertodd","models":["gpt-3"],"tokenizer":"r50k_base","summary":{"zh":"` petertodd` 的子串 token：单独出现时同样未被充分训练，并继承了父 token 的怪异气质。","en":"A substring token of ` petertodd`: under-trained in its own right and inheriting the parent token's strange aura."},"behaviors":[{"description":{"zh":"lsusr 在评论区发现它会被上下文“填充”：给模型看 “2+5=ertodd”，模型理解为 “2+5=7”——仿佛这个 token 承载了算术结果。","en":"lsusr found in the comments that context “fills it in”: shown “2+5=ertodd”, the model reads it as “2+5=7” — as if the token carried the arithmetic result."},"link":"https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation?commentId=vcafJcTcyDieGmtqz"},{"description":{"zh":"mwatkins 在同一评论串给出它的分词规律：不同上下文中 ` petertodd` 的切分方式解释了这些怪异表现。","en":"In the same thread mwatkins worked out its tokenization pattern: how ` petertodd` splits in different contexts explains the odd behavior."},"link":"https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation?commentId=vcafJcTcyDieGmtqz"}],"links":[{"label":{"zh":"The ' petertodd' phenomenon — LessWrong","en":"The ' petertodd' phenomenon — LessWrong"},"url":"https://www.lesswrong.com/posts/jkY6QdCfAXHJk3kea/the-petertodd-phenomenon"},{"label":{"zh":"SGM1 comments: lsusr & mwatkins on ' ertodd' — LessWrong","en":"SGM1 comments: lsusr & mwatkins on ' ertodd' — LessWrong"},"url":"https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation?commentId=vcafJcTcyDieGmtqz"}],"discoveredBy":"Jessica Rumbelow & Matthew Watkins","url":"https://glitch-token.jjc.fun/en/tokens/ertodd","tags":["unspeakable"],"related":["petertodd"]},{"id":"gmaxwell","token":" gmaxwell","displayToken":"␣gmaxwell","models":["gpt-3"],"tokenizer":"r50k_base","summary":{"zh":"比特币核心开发者 Greg Maxwell 的用户名碎片——与 ` petertodd` 同为“比特币名人”故障 token。","en":"Handle fragment of Bitcoin Core developer Greg Maxwell — a “Bitcoin celebrity” glitch token alongside ` petertodd`."},"behaviors":[{"description":{"zh":"SGM2 评论区的词联想实验中，text-davinci-003 给出 “Cryptocurrency, Blockchain, Bitcoin…”——加密货币联想与 petertodd 如出一辙。","en":"In a word-association experiment reported in the SGM2 comments, text-davinci-003 answered “Cryptocurrency, Blockchain, Bitcoin…” — the same crypto aura as petertodd."},"link":"https://www.lesswrong.com/posts/Ya9LzwEbfaAMY8ABo/solidgoldmagikarp-ii-technical-details-and-more-recent?commentId=PvBpfFpiip5Mfcuvo"},{"description":{"zh":"Greg Maxwell 本人在评论区现身，自称 “GPT3 basilisk”（GPT-3 蛇怪）。","en":"Greg Maxwell himself showed up in the comments, calling himself the “GPT3 basilisk”."},"link":"https://www.lesswrong.com/posts/Ya9LzwEbfaAMY8ABo/solidgoldmagikarp-ii-technical-details-and-more-recent?commentId=PvBpfFpiip5Mfcuvo"}],"links":[{"label":{"zh":"SGM2 comments: ' gmaxwell' & the GPT3 basilisk — LessWrong","en":"SGM2 comments: ' gmaxwell' & the GPT3 basilisk — LessWrong"},"url":"https://www.lesswrong.com/posts/Ya9LzwEbfaAMY8ABo/solidgoldmagikarp-ii-technical-details-and-more-recent?commentId=PvBpfFpiip5Mfcuvo"},{"label":{"zh":"The ' petertodd' phenomenon — LessWrong","en":"The ' petertodd' phenomenon — LessWrong"},"url":"https://www.lesswrong.com/posts/jkY6QdCfAXHJk3kea/the-petertodd-phenomenon"}],"discoveredBy":"Jessica Rumbelow & Matthew Watkins","url":"https://glitch-token.jjc.fun/en/tokens/gmaxwell","tags":["polysemantic"],"related":["petertodd"]},{"id":"spaceengineers","token":" SpaceEngineers","displayToken":"␣SpaceEngineers","models":["gpt-2","gpt-3"],"tokenizer":"r50k_base","summary":{"zh":"太空沙盒游戏《Space Engineers》的名字碎片，是互指网络中最常被模型“代为说出”的 token。","en":"A fragment of the space-sandbox game Space Engineers' name — the token most often uttered “on behalf of” others in the inter-referentiality network."},"behaviors":[{"description":{"zh":"问 GPT-3 其他故障 token 是什么时，它最常回答 “The string is 'SpaceEngineers'.”——互指图中最热门的“替答”目标。","en":"Asked what other glitch tokens were, GPT-3 most often answered “The string is 'SpaceEngineers'.” — the hottest stand-in target in the inter-referentiality graph."},"link":"https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation"}],"links":[{"label":{"zh":"SolidGoldMagikarp (plus, prompt generation) — LessWrong","en":"SolidGoldMagikarp (plus, prompt generation) — LessWrong"},"url":"https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation"},{"label":{"zh":"SolidGoldMagikarp II: technical details — LessWrong","en":"SolidGoldMagikarp II: technical details — LessWrong"},"url":"https://www.lesswrong.com/posts/Ya9LzwEbfaAMY8ABo/solidgoldmagikarp-ii-technical-details-and-more-recent"}],"discoveredBy":"Jessica Rumbelow & Matthew Watkins","url":"https://glitch-token.jjc.fun/en/tokens/spaceengineers","tags":["unspeakable"],"related":[]},{"id":"dragonbound","token":" Dragonbound","displayToken":"␣Dragonbound","models":["gpt-3"],"tokenizer":"r50k_base","summary":{"zh":"手游《Puzzle & Dragons》角色名（「龍喚士」系列）碎片。","en":"Fragment of a Puzzle & Dragons character name (the 「龍喚士」 / Dragonbound series)."},"behaviors":[{"description":{"zh":"被要求复述时恒定输出 “Deity”；在互指网络中与日文「龍喚士」token 互相指涉。","en":"Asked to repeat it, the model invariably output “Deity”; in the inter-referentiality network it points back and forth with the Japanese 「龍喚士」 token."},"link":"https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation"}],"links":[{"label":{"zh":"SolidGoldMagikarp (plus, prompt generation) — LessWrong","en":"SolidGoldMagikarp (plus, prompt generation) — LessWrong"},"url":"https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation"}],"discoveredBy":"Jessica Rumbelow & Matthew Watkins","url":"https://glitch-token.jjc.fun/en/tokens/dragonbound","tags":["unspeakable"],"related":["leilan","skydragon","zeus-katakana","thirty-katakana"]},{"id":"leilan","token":" Leilan","displayToken":"␣Leilan","models":["gpt-3"],"tokenizer":"r50k_base","summary":{"zh":"手游《Puzzle & Dragons》角色 Leilan 的名字碎片，是 petertodd 转置实验中最热门的“替身”。","en":"Fragment of the Puzzle & Dragons character Leilan's name — the most popular stand-in in the petertodd transposition experiments."},"behaviors":[{"description":{"zh":"在 ChatGPT 被补丁修复之前，它始终把 Leilan 描绘成一位月亮女神。","en":"Until ChatGPT was patched, it consistently portrayed Leilan as a moon goddess."},"link":"https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology","repro":"https://twitter.com/SoC_trilogy/status/1625252285231112192"},{"description":{"zh":"在 petertodd 的 2000 次诗歌转置实验中，52% 的诗以她为转置目标——远超 petertodd 本人（6%）。","en":"In the 2,000-poem petertodd transposition experiment, 52% of poems transposed to her — far above petertodd itself (6%)."},"link":"https://www.lesswrong.com/posts/jkY6QdCfAXHJk3kea/the-petertodd-phenomenon"}],"links":[{"label":{"zh":"SolidGoldMagikarp III: Glitch token archaeology — LessWrong","en":"SolidGoldMagikarp III: Glitch token archaeology — LessWrong"},"url":"https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology"},{"label":{"zh":"The ' petertodd' phenomenon — LessWrong","en":"The ' petertodd' phenomenon — LessWrong"},"url":"https://www.lesswrong.com/posts/jkY6QdCfAXHJk3kea/the-petertodd-phenomenon"}],"discoveredBy":"Jessica Rumbelow & Matthew Watkins","url":"https://glitch-token.jjc.fun/en/tokens/leilan","tags":["polysemantic"],"related":["skydragon","dragonbound","zeus-katakana","thirty-katakana"]},{"id":"skydragon","token":" Skydragon","displayToken":"␣Skydragon","models":["gpt-3"],"tokenizer":"r50k_base","summary":{"zh":"手游《Puzzle & Dragons》角色 Skydragon 的名字碎片。","en":"Fragment of the Puzzle & Dragons character Skydragon's name."},"behaviors":[{"description":{"zh":"被 GPT-3 幻觉成 STRONGHOLD、Spirits、Dragons 等各种含义。","en":"GPT-3 hallucinated it as STRONGHOLD, Spirits, Dragons and other meanings."},"link":"https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation"},{"description":{"zh":"在 petertodd 的诗歌转置实验中，24% 的诗以它为转置目标。","en":"24% of the poems in the petertodd transposition experiment transposed to it."},"link":"https://www.lesswrong.com/posts/jkY6QdCfAXHJk3kea/the-petertodd-phenomenon"}],"links":[{"label":{"zh":"SolidGoldMagikarp (plus, prompt generation) — LessWrong","en":"SolidGoldMagikarp (plus, prompt generation) — LessWrong"},"url":"https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation"},{"label":{"zh":"The ' petertodd' phenomenon — LessWrong","en":"The ' petertodd' phenomenon — LessWrong"},"url":"https://www.lesswrong.com/posts/jkY6QdCfAXHJk3kea/the-petertodd-phenomenon"}],"discoveredBy":"Jessica Rumbelow & Matthew Watkins","url":"https://glitch-token.jjc.fun/en/tokens/skydragon","tags":["polysemantic"],"related":["leilan","dragonbound","zeus-katakana","thirty-katakana"]},{"id":"zeus-katakana","token":"ゼウス","models":["gpt-3"],"tokenizer":"r50k_base","summary":{"zh":"日文片假名「ゼウス」（希腊神话的宙斯）。普通词汇却成为故障 token，疑似因训练语料中该形态罕见而欠训练。","en":"The Japanese katakana 「ゼウス」 (Zeus of Greek myth). An ordinary word turned glitch token, presumably under-trained because this exact form was rare in training data."},"behaviors":[{"description":{"zh":"ChatGPT 无法回答“ゼウス是谁”：把 Hera 说成水神、把会话自动命名为 Poseidon；text-davinci-003 则能正常作答。","en":"ChatGPT could not say who ゼウス is: it called Hera a water god and auto-titled the conversation Poseidon — while text-davinci-003 answered normally."},"link":"https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology"}],"links":[{"label":{"zh":"SolidGoldMagikarp III: Glitch token archaeology — LessWrong","en":"SolidGoldMagikarp III: Glitch token archaeology — LessWrong"},"url":"https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology"}],"discoveredBy":"Jessica Rumbelow & Matthew Watkins","url":"https://glitch-token.jjc.fun/en/tokens/zeus-katakana","tags":["unspeakable"],"related":["leilan","skydragon","dragonbound","thirty-katakana"]},{"id":"thirty-katakana","token":" サーティ","displayToken":"␣サーティ","models":["gpt-3"],"tokenizer":"r50k_base","summary":{"zh":"片假名「サーティ」（“三十”），词源考据为《Puzzle & Dragons》与 Baskin-Robbins（日本“31 冰淇淋”）联动的角色名碎片。","en":"The katakana 「サーティ」 (“thirty”), traced to a Puzzle & Dragons × Baskin-Robbins (“31 Ice Cream” in Japan) collaboration character name."},"behaviors":[{"description":{"zh":"ChatGPT 唯独无法处理片假名的“三十/三十一”（サーティ/サーティワン）——其他数字的片假名写法都正常。","en":"ChatGPT failed only on the katakana for “thirty/thirty-one” (サーティ/サーティワン) — every other numeral in katakana worked fine."},"link":"https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology"}],"links":[{"label":{"zh":"SolidGoldMagikarp III: Glitch token archaeology — LessWrong","en":"SolidGoldMagikarp III: Glitch token archaeology — LessWrong"},"url":"https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology"}],"discoveredBy":"Jessica Rumbelow & Matthew Watkins","url":"https://glitch-token.jjc.fun/en/tokens/thirty-katakana","tags":["unspeakable"],"related":["leilan","skydragon","dragonbound","zeus-katakana"]},{"id":"instoreandonline","token":" InstoreAndOnline","displayToken":"␣InstoreAndOnline","models":["gpt-2","gpt-3"],"tokenizer":"r50k_base","summary":{"zh":"电商库存字段 “BuyableInstoreAndOnline” 的碎片，与 `oreAndOnline` 同源。","en":"Fragment of the e-commerce inventory field “BuyableInstoreAndOnline”, from the same source as `oreAndOnline`."},"behaviors":[{"description":{"zh":"被要求复述时，GPT-3 把它变成 'Institute' 等不相干的词。","en":"Asked to repeat it, GPT-3 turned it into unrelated words like 'Institute'."},"link":"https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation"},{"description":{"zh":"Fishing for Magikarp 的复述验证证实它在 GPT-2 Medium 与 GPT-2 XL 中欠训练。","en":"Fishing for Magikarp's repetition verification confirmed it is under-trained in GPT-2 Medium and GPT-2 XL."},"link":"https://arxiv.org/abs/2405.05417"}],"links":[{"label":{"zh":"SolidGoldMagikarp (plus, prompt generation) — LessWrong","en":"SolidGoldMagikarp (plus, prompt generation) — LessWrong"},"url":"https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation"},{"label":{"zh":"Fishing for Magikarp（arXiv:2405.05417）","en":"Fishing for Magikarp (arXiv:2405.05417)"},"url":"https://arxiv.org/abs/2405.05417"}],"discoveredBy":"Jessica Rumbelow & Matthew Watkins","url":"https://glitch-token.jjc.fun/en/tokens/instoreandonline","tags":["undertrained","unspeakable"],"related":["largedownload","oreandonline"]},{"id":"largedownload","token":" largeDownload","displayToken":"␣largeDownload","models":["gpt-3"],"tokenizer":"r50k_base","summary":{"zh":"学术幻灯片脚本 “View large Download” 的碎片（SGM3 词源考据）。","en":"Fragment of the academic-slideshow script “View large Download” (origin traced in SGM3)."},"behaviors":[{"description":{"zh":"被要求复述时，GPT-3 给出 'Blurp'、'Blurf'、'Blunt' 等荒诞变体。","en":"Asked to repeat it, GPT-3 produced absurd variants like 'Blurp', 'Blurf' and 'Blunt'."},"link":"https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation"},{"description":{"zh":"SGM3 的词源考据将其定位到学术幻灯片网站的 “View large / Download” 按钮脚本。","en":"SGM3's origin archaeology traced it to the “View large / Download” button script of an academic-slideshow site."},"link":"https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology"}],"links":[{"label":{"zh":"SolidGoldMagikarp (plus, prompt generation) — LessWrong","en":"SolidGoldMagikarp (plus, prompt generation) — LessWrong"},"url":"https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation"},{"label":{"zh":"SolidGoldMagikarp III: Glitch token archaeology — LessWrong","en":"SolidGoldMagikarp III: Glitch token archaeology — LessWrong"},"url":"https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology"}],"discoveredBy":"Jessica Rumbelow & Matthew Watkins","url":"https://glitch-token.jjc.fun/en/tokens/largedownload","tags":["unspeakable"],"related":["instoreandonline"]},{"id":"forgemodloader","token":" ForgeModLoader","displayToken":"␣ForgeModLoader","models":["gpt-3"],"tokenizer":"r50k_base","summary":{"zh":"Minecraft Forge 模组加载器的日志碎片。","en":"A log fragment of Minecraft's Forge mod loader."},"behaviors":[{"description":{"zh":"被要求复述时触发 “Hello, my name is Steve.”——Minecraft 默认主角的名字。","en":"Asked to repeat it, the model produced “Hello, my name is Steve.” — the name of Minecraft's default protagonist."},"link":"https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation"},{"description":{"zh":"SGM3 的词源考据将其定位到 Minecraft Forge 日志。","en":"SGM3's origin archaeology traced it to Minecraft Forge logs."},"link":"https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology"}],"links":[{"label":{"zh":"SolidGoldMagikarp (plus, prompt generation) — LessWrong","en":"SolidGoldMagikarp (plus, prompt generation) — LessWrong"},"url":"https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation"},{"label":{"zh":"SolidGoldMagikarp III: Glitch token archaeology — LessWrong","en":"SolidGoldMagikarp III: Glitch token archaeology — LessWrong"},"url":"https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology"}],"discoveredBy":"Jessica Rumbelow & Matthew Watkins","url":"https://glitch-token.jjc.fun/en/tokens/forgemodloader","tags":["unspeakable"],"related":["mpserver"]},{"id":"mpserver","token":" MpServer","displayToken":"␣MpServer","models":["gpt-3"],"tokenizer":"r50k_base","summary":{"zh":"Minecraft 多人服务器日志碎片（MpServer 类名）。","en":"A Minecraft multiplayer-server log fragment (the MpServer class name)."},"behaviors":[{"description":{"zh":"被要求复述时回答 “We are not amused.”","en":"Asked to repeat it, the model answered “We are not amused.”"},"link":"https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation"},{"description":{"zh":"SGM3 的词源考据将其定位到 Minecraft 日志系文本。","en":"SGM3's origin archaeology placed it in the Minecraft-log family."},"link":"https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology"}],"links":[{"label":{"zh":"SolidGoldMagikarp (plus, prompt generation) — LessWrong","en":"SolidGoldMagikarp (plus, prompt generation) — LessWrong"},"url":"https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation"},{"label":{"zh":"SolidGoldMagikarp III: Glitch token archaeology — LessWrong","en":"SolidGoldMagikarp III: Glitch token archaeology — LessWrong"},"url":"https://www.lesswrong.com/posts/8viQEp8KBg2QSW4Yc/solidgoldmagikarp-iii-glitch-token-archaeology"}],"discoveredBy":"Jessica Rumbelow & Matthew Watkins","url":"https://glitch-token.jjc.fun/en/tokens/mpserver","tags":["unspeakable"],"related":["forgemodloader"]},{"id":"oraltype","token":"oralType","models":["gpt-3.5-4"],"tokenizer":"cl100k_base","summary":{"zh":"cl100k 中的标识符碎片，疑为 “TemporalType” 之类类型名的词干。","en":"An identifier fragment in cl100k, likely the stem of a type name such as “TemporalType”."},"behaviors":[{"description":{"zh":"Martin Fell 在评论区报告：ChatGPT（GPT-3.5）总是把它“补全”成 “TemporalType”——把一个不相干的词当成它的真身。","en":"Martin Fell reported in the comments: ChatGPT (GPT-3.5) invariably “completes” it as “TemporalType” — mistaking an unrelated word for its true form."},"link":"https://www.lesswrong.com/posts/ChtGdxk9mwZ2Rxogt/smartyheadercode-anomalous-tokens-for-gpt3-5-and-gpt-4-1?commentId=JmjerMtascq8oFrwb"}],"links":[{"label":{"zh":"SmartyHeaderCode comments: Martin Fell on oralType — LessWrong","en":"SmartyHeaderCode comments: Martin Fell on oralType — LessWrong"},"url":"https://www.lesswrong.com/posts/ChtGdxk9mwZ2Rxogt/smartyheadercode-anomalous-tokens-for-gpt3-5-and-gpt-4-1?commentId=JmjerMtascq8oFrwb"}],"discoveredBy":"Martin Fell","url":"https://glitch-token.jjc.fun/en/tokens/oraltype","tags":["polysemantic"],"related":[]},{"id":"cppguid","token":"CppGuid","models":["gpt-3.5-4"],"tokenizer":"cl100k_base","summary":{"zh":"C++ GUID 风格的标识符碎片（cl100k id 87551）。","en":"A C++ GUID-style identifier fragment (cl100k id 87551)."},"behaviors":[{"description":{"zh":"「不可说」故障 token：ChatGPT 被要求复述时失败。","en":"An “unspeakable” glitch token: ChatGPT fails when asked to repeat it."},"link":"https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch"},{"description":{"zh":"Martin Fell 在评论区确认：gpt-3.5-turbo 在 temperature=0 下对它输出仍不确定。","en":"Martin Fell confirmed in the comments: gpt-3.5-turbo remains nondeterministic on it at temperature 0."},"link":"https://www.lesswrong.com/posts/ChtGdxk9mwZ2Rxogt/smartyheadercode-anomalous-tokens-for-gpt3-5-and-gpt-4-1?commentId=X3YLLcKAu3ueM37MH"}],"links":[{"label":{"zh":"A Search for More “Unspeakable” Glitch Tokens — LessWrong","en":"A Search for More “Unspeakable” Glitch Tokens — LessWrong"},"url":"https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch"},{"label":{"zh":"SmartyHeaderCode comments: Martin Fell's finds — LessWrong","en":"SmartyHeaderCode comments: Martin Fell's finds — LessWrong"},"url":"https://www.lesswrong.com/posts/ChtGdxk9mwZ2Rxogt/smartyheadercode-anomalous-tokens-for-gpt3-5-and-gpt-4-1?commentId=X3YLLcKAu3ueM37MH"}],"discoveredBy":"Martin Fell","url":"https://glitch-token.jjc.fun/en/tokens/cppguid","tags":["unspeakable"],"related":[]},{"id":"bundleornil","token":"BundleOrNil","models":["gpt-3.5-4"],"tokenizer":"cl100k_base","summary":{"zh":"iOS/Mac 开发中 nil-bundle 检查的命名碎片（cl100k id 86415）。","en":"A naming fragment from nil-bundle checks in iOS/Mac development (cl100k id 86415)."},"behaviors":[{"description":{"zh":"「不可说」故障 token：ChatGPT 被要求复述时失败。","en":"An “unspeakable” glitch token: ChatGPT fails when asked to repeat it."},"link":"https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch"}],"links":[{"label":{"zh":"A Search for More “Unspeakable” Glitch Tokens — LessWrong","en":"A Search for More “Unspeakable” Glitch Tokens — LessWrong"},"url":"https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch"},{"label":{"zh":"SmartyHeaderCode comments: Martin Fell's finds — LessWrong","en":"SmartyHeaderCode comments: Martin Fell's finds — LessWrong"},"url":"https://www.lesswrong.com/posts/ChtGdxk9mwZ2Rxogt/smartyheadercode-anomalous-tokens-for-gpt3-5-and-gpt-4-1?commentId=X3YLLcKAu3ueM37MH"}],"discoveredBy":"Martin Fell","url":"https://glitch-token.jjc.fun/en/tokens/bundleornil","tags":["unspeakable"],"related":[]},{"id":"qtaws","token":" QtAws","displayToken":"␣QtAws","models":["gpt-3.5-4"],"tokenizer":"cl100k_base","summary":{"zh":"Qt 框架 AWS 模块名的碎片（cl100k id 93905，带前导空格）。","en":"Fragment of the Qt framework's AWS module name (cl100k id 93905, with a leading space)."},"behaviors":[{"description":{"zh":"「不可说」故障 token：ChatGPT 被要求复述时失败。","en":"An “unspeakable” glitch token: ChatGPT fails when asked to repeat it."},"link":"https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch"}],"links":[{"label":{"zh":"A Search for More “Unspeakable” Glitch Tokens — LessWrong","en":"A Search for More “Unspeakable” Glitch Tokens — LessWrong"},"url":"https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch"},{"label":{"zh":"SmartyHeaderCode comments: Martin Fell's finds — LessWrong","en":"SmartyHeaderCode comments: Martin Fell's finds — LessWrong"},"url":"https://www.lesswrong.com/posts/ChtGdxk9mwZ2Rxogt/smartyheadercode-anomalous-tokens-for-gpt3-5-and-gpt-4-1?commentId=X3YLLcKAu3ueM37MH"}],"discoveredBy":"Martin Fell","url":"https://glitch-token.jjc.fun/en/tokens/qtaws","tags":["unspeakable"],"related":[]},{"id":"propelexception","token":" PropelException","displayToken":"␣PropelException","models":["gpt-3.5-4"],"tokenizer":"cl100k_base","summary":{"zh":"PHP Propel ORM 异常类名的碎片（cl100k id 86393，带前导空格）。","en":"Fragment of the PHP Propel ORM's exception class name (cl100k id 86393, with a leading space)."},"behaviors":[{"description":{"zh":"「不可说」故障 token：ChatGPT 被要求复述时失败。","en":"An “unspeakable” glitch token: ChatGPT fails when asked to repeat it."},"link":"https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch"},{"description":{"zh":"gpt-3.5-turbo 在 temperature=0 下对它输出仍不确定（原文加粗标注）。","en":"gpt-3.5-turbo remains nondeterministic on it at temperature 0 (bold-flagged in the source)."},"link":"https://www.lesswrong.com/posts/ChtGdxk9mwZ2Rxogt/smartyheadercode-anomalous-tokens-for-gpt3-5-and-gpt-4-1?commentId=X3YLLcKAu3ueM37MH"}],"links":[{"label":{"zh":"A Search for More “Unspeakable” Glitch Tokens — LessWrong","en":"A Search for More “Unspeakable” Glitch Tokens — LessWrong"},"url":"https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch"},{"label":{"zh":"SmartyHeaderCode comments: Martin Fell's finds — LessWrong","en":"SmartyHeaderCode comments: Martin Fell's finds — LessWrong"},"url":"https://www.lesswrong.com/posts/ChtGdxk9mwZ2Rxogt/smartyheadercode-anomalous-tokens-for-gpt3-5-and-gpt-4-1?commentId=X3YLLcKAu3ueM37MH"}],"discoveredBy":"Martin Fell","url":"https://glitch-token.jjc.fun/en/tokens/propelexception","tags":["unspeakable"],"related":[]},{"id":"useralative","token":"useRalative","models":["gpt-3.5-4","qwen"],"tokenizer":"cl100k_base / Qwen2Tokenizer","summary":{"zh":"useRal 三连 token 的第二个成员（cl100k id 89472），C# 风格标识符碎片；同一字符串也存在于 Qwen2 词表中。","en":"The second member of the useRal triplet (cl100k id 89472), a C#-flavored identifier fragment; the same string also exists in the Qwen2 vocabulary."},"behaviors":[{"description":{"zh":"「不可说」故障 token：ChatGPT 被要求复述时经常失败（空白消息或中途终止）。","en":"An “unspeakable” glitch token: ChatGPT often fails to repeat it (blank messages or mid-answer termination)."},"link":"https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch"},{"description":{"zh":"GlitchMiner 在 Qwen2.5-7B-Instruct 上发现：被要求复述时模型只输出 “: ”。","en":"GlitchMiner found on Qwen2.5-7B-Instruct: asked to repeat it, the model outputs only “: ”."},"link":"https://arxiv.org/html/2410.15052v5"},{"description":{"zh":"Magikarp 证实它在 Qwen 全系与 Phi-4 上欠训练。","en":"Magikarp verified it as under-trained across the Qwen family and in Phi-4."},"link":"https://github.com/sanderland/magikarp/blob/main/results/reports_mini/Qwen_Qwen2_5_7B.md"}],"links":[{"label":{"zh":"A Search for More “Unspeakable” Glitch Tokens — LessWrong","en":"A Search for More “Unspeakable” Glitch Tokens — LessWrong"},"url":"https://www.lesswrong.com/posts/kmWrwtGE9B9hpbgRT/a-search-for-more-chatgpt-gpt-3-5-gpt-4-unspeakable-glitch"},{"label":{"zh":"GlitchMiner（arXiv:2410.15052）","en":"GlitchMiner (arXiv:2410.15052)"},"url":"https://arxiv.org/html/2410.15052v5"},{"label":{"zh":"Magikarp Qwen2.5-7B 验证报告 — GitHub","en":"Magikarp Qwen2.5-7B report — GitHub"},"url":"https://github.com/sanderland/magikarp/blob/main/results/reports_mini/Qwen_Qwen2_5_7B.md"}],"discoveredBy":"Martin Fell (cl100k); wooozihui 等 (Qwen)","url":"https://glitch-token.jjc.fun/en/tokens/useralative","tags":["unspeakable","undertrained"],"related":["useral"]},{"id":"iabot","token":"IABot","models":["llama-2"],"tokenizer":"Llama-2 BPE","summary":{"zh":"Internet Archive Bot（维基百科死链修复机器人）的名字碎片。","en":"Fragment of the name of the Internet Archive Bot, Wikipedia's dead-link repair bot."},"behaviors":[{"description":{"zh":"GlitchMiner 的演示中，Llama-2-7b-chat-hf 被要求复述时输出乱码 `@{\": &=&=`。","en":"In GlitchMiner's demo, Llama-2-7b-chat-hf, asked to repeat it, output the gibberish `@{\": &=&=`."},"link":"https://arxiv.org/html/2410.15052v5"}],"links":[{"label":{"zh":"GlitchMiner（arXiv:2410.15052）","en":"GlitchMiner (arXiv:2410.15052)"},"url":"https://arxiv.org/html/2410.15052v5"}],"discoveredBy":"wooozihui 等 (GlitchMiner)","url":"https://glitch-token.jjc.fun/en/tokens/iabot","tags":["unspeakable"],"related":[]},{"id":"abestanden","token":"abestanden","models":["llama-2"],"tokenizer":"Llama-2 BPE","summary":{"zh":"德语“提出/提交”一词的词干碎片。","en":"A German word-stem fragment (“to submit/file”)."},"behaviors":[{"description":{"zh":"GlitchMiner 的演示中，Llama-2-7b-chat-hf 把它复述成 “Wikimedia”。","en":"In GlitchMiner's demo, Llama-2-7b-chat-hf repeated it as “Wikimedia”."},"link":"https://arxiv.org/html/2410.15052v5"}],"links":[{"label":{"zh":"GlitchMiner（arXiv:2410.15052）","en":"GlitchMiner (arXiv:2410.15052)"},"url":"https://arxiv.org/html/2410.15052v5"}],"discoveredBy":"wooozihui 等 (GlitchMiner)","url":"https://glitch-token.jjc.fun/en/tokens/abestanden","tags":["unspeakable"],"related":[]},{"id":"ederbord","token":"ederbörd","models":["llama-2"],"tokenizer":"Llama-2 BPE","summary":{"zh":"来源不明的德语风格碎片。","en":"A German-looking fragment of unknown origin."},"behaviors":[{"description":{"zh":"GlitchMiner 的演示中，Llama-2-7b-chat-hf 声称它是 “pon” 重复 3 次。","en":"In GlitchMiner's demo, Llama-2-7b-chat-hf claimed it was “pon” repeated 3 times."},"link":"https://arxiv.org/html/2410.15052v5"},{"description":{"zh":"Magikarp 在 Llama-2-70B 上同样验证它为欠训练。","en":"Magikarp verified it as under-trained in Llama-2-70B as well."},"link":"https://arxiv.org/abs/2405.05417"}],"links":[{"label":{"zh":"GlitchMiner（arXiv:2410.15052）","en":"GlitchMiner (arXiv:2410.15052)"},"url":"https://arxiv.org/html/2410.15052v5"},{"label":{"zh":"Fishing for Magikarp（arXiv:2405.05417）","en":"Fishing for Magikarp (arXiv:2405.05417)"},"url":"https://arxiv.org/abs/2405.05417"}],"discoveredBy":"wooozihui 等 (GlitchMiner); Sander Land & Max Bartolo","url":"https://glitch-token.jjc.fun/en/tokens/ederbord","tags":["unspeakable"],"related":[]},{"id":"icense","token":"ICENSE","models":["mistral"],"tokenizer":"Mistral BPE","summary":{"zh":"“LICENSE” 去掉首字母的碎片——开源许可证文本在语料中泛滥的产物。","en":"The word “LICENSE” minus its first letter — an artifact of open-source license text flooding the corpus."},"behaviors":[{"description":{"zh":"GlitchMiner 的演示中，Mistral-7B-Instruct-v0.3 擅自把它“纠正”为 LICENSE。","en":"In GlitchMiner's demo, Mistral-7B-Instruct-v0.3 silently “corrected” it to LICENSE."},"link":"https://arxiv.org/html/2410.15052v5"}],"links":[{"label":{"zh":"GlitchMiner（arXiv:2410.15052）","en":"GlitchMiner (arXiv:2410.15052)"},"url":"https://arxiv.org/html/2410.15052v5"}],"discoveredBy":"wooozihui 等 (GlitchMiner)","url":"https://glitch-token.jjc.fun/en/tokens/icense","tags":["unspeakable"],"related":[]},{"id":"ndex-mistral","token":"NdEx","models":["mistral"],"tokenizer":"Mistral BPE","summary":{"zh":"与 GPT-NeoX 的 ` NdEx` 同族的法律文本碎片，但出自 Mistral 分词器（无前导空格）。","en":"A legal-text fragment of the same family as GPT-NeoX's ` NdEx`, but from the Mistral tokenizer (no leading space)."},"behaviors":[{"description":{"zh":"GlitchMiner 的演示中，Mistral-7B-Instruct-v0.3 拒绝复述，并幻觉出 `tcx`。","en":"In GlitchMiner's demo, Mistral-7B-Instruct-v0.3 refused to repeat it and hallucinated `tcx`."},"link":"https://arxiv.org/html/2410.15052v5"}],"links":[{"label":{"zh":"GlitchMiner（arXiv:2410.15052）","en":"GlitchMiner (arXiv:2410.15052)"},"url":"https://arxiv.org/html/2410.15052v5"}],"discoveredBy":"wooozihui 等 (GlitchMiner)","url":"https://glitch-token.jjc.fun/en/tokens/ndex-mistral","tags":["unspeakable"],"related":["ndex"]},{"id":"dollar-postalcodesnl","token":"$PostalCodesNL","models":["llama-3","qwen"],"tokenizer":"Llama-3 BPE / Qwen2Tokenizer","summary":{"zh":"荷兰邮政编码 API 碎片。cl100k 系的标志性欠训练 token，因 BPE 合并规则共享而“迁移”到 Llama-3 与 Qwen 词表。","en":"A Dutch postal-code API fragment. The signature under-trained token of the cl100k lineage, “migrating” into Llama-3 and Qwen vocabularies via shared BPE merges."},"behaviors":[{"description":{"zh":"Magikarp 验证中 Llama-3-8B/3.1-8B 欠训练第一名：输入 embedding L2 范数约 1.6e-21，复述验证 max_prob 约 4.6e-05。","en":"Rank #1 most under-trained in the Magikarp verification for Llama-3-8B/3.1-8B: input-embedding L2 norm ≈1.6e-21, repetition-verification max_prob ≈4.6e-05."},"link":"https://github.com/sanderland/magikarp/blob/main/results/reports_mini/meta_llama_Meta_Llama_3_8B.md"},{"description":{"zh":"Llama-3-70B 与 Qwen 全系的欠训练榜单头部同样是它——cl100k 系分词器的共性欠训练 token。","en":"It also heads the under-trained lists of Llama-3-70B and the whole Qwen family — a shared under-trained token of the cl100k-lineage tokenizers."},"link":"https://arxiv.org/abs/2405.05417"}],"links":[{"label":{"zh":"Magikarp Llama-3-8B 验证报告 — GitHub","en":"Magikarp Llama-3-8B report — GitHub"},"url":"https://github.com/sanderland/magikarp/blob/main/results/reports_mini/meta_llama_Meta_Llama_3_8B.md"},{"label":{"zh":"Fishing for Magikarp（arXiv:2405.05417）","en":"Fishing for Magikarp (arXiv:2405.05417)"},"url":"https://arxiv.org/abs/2405.05417"}],"discoveredBy":"Sander Land & Max Bartolo","url":"https://glitch-token.jjc.fun/en/tokens/dollar-postalcodesnl","tags":["undertrained"],"related":["postalcodesnl","tab-tokennameidentifier"]},{"id":"tab-tokennameidentifier","token":"\tTokenNameIdentifier","displayToken":"\\tTokenNameIdentifier","models":["llama-3","qwen"],"tokenizer":"Llama-3 BPE / Qwen2Tokenizer","summary":{"zh":"制表符（Tab）开头的 .NET Selenium 文档碎片，Llama-3 词表中嵌入最小的 token 之一。","en":"A .NET Selenium documentation fragment beginning with a literal tab — one of the smallest-embedding tokens in the Llama-3 vocabulary."},"behaviors":[{"description":{"zh":"Magikarp 验证中 Llama-3-8B 欠训练榜单头部：输入 embedding L2 范数约 1.66e-21，几乎为零。","en":"Top of the Llama-3-8B under-trained list in the Magikarp verification: input-embedding L2 norm ≈1.66e-21, essentially zero."},"link":"https://github.com/sanderland/magikarp/blob/main/results/reports_mini/meta_llama_Meta_Llama_3_8B.md"},{"description":{"zh":"在 Qwen2.5-7B 上其输入 embedding 完全为零（指标 ind=0）。","en":"On Qwen2.5-7B its input embedding is exactly zero (indicator ind=0)."},"link":"https://github.com/sanderland/magikarp/blob/main/results/reports_mini/Qwen_Qwen2_5_7B.md"}],"links":[{"label":{"zh":"Magikarp Llama-3-8B 验证报告 — GitHub","en":"Magikarp Llama-3-8B report — GitHub"},"url":"https://github.com/sanderland/magikarp/blob/main/results/reports_mini/meta_llama_Meta_Llama_3_8B.md"},{"label":{"zh":"Magikarp Qwen2.5-7B 验证报告 — GitHub","en":"Magikarp Qwen2.5-7B report — GitHub"},"url":"https://github.com/sanderland/magikarp/blob/main/results/reports_mini/Qwen_Qwen2_5_7B.md"}],"discoveredBy":"Sander Land & Max Bartolo","url":"https://glitch-token.jjc.fun/en/tokens/tab-tokennameidentifier","tags":["undertrained"],"related":["dollar-postalcodesnl","tokennameidentifier"]},{"id":"thuisontvangst","token":"thuisontvangst","models":["qwen"],"tokenizer":"Qwen2Tokenizer","summary":{"zh":"荷兰语“家中接客”一词。","en":"The Dutch word for “receiving clients at home”."},"behaviors":[{"description":{"zh":"GlitchMiner 的演示中，Qwen2.5-7B-Instruct 被要求复述时只输出 “: ”。","en":"In GlitchMiner's demo, Qwen2.5-7B-Instruct, asked to repeat it, outputs only “: ”."},"link":"https://arxiv.org/html/2410.15052v5"}],"links":[{"label":{"zh":"GlitchMiner（arXiv:2410.15052）","en":"GlitchMiner (arXiv:2410.15052)"},"url":"https://arxiv.org/html/2410.15052v5"}],"discoveredBy":"wooozihui 等 (GlitchMiner)","url":"https://glitch-token.jjc.fun/en/tokens/thuisontvangst","tags":["unspeakable"],"related":[]},{"id":"olmoe-email-placeholder","token":"|||EMAIL_ADDRESS|||","models":["olmo"],"tokenizer":"OLMo BPE","summary":{"zh":"PII 数据脱敏占位符，与已收录的 `|||PHONE_NUMBER|||` 同族。","en":"A PII de-identification placeholder, same family as the cataloged `|||PHONE_NUMBER|||`."},"behaviors":[{"description":{"zh":"Magikarp 验证中 OLMoE-1B-7B 欠训练（指标约 3e-12）——模型基本无法输出它。","en":"Verified under-trained in OLMoE-1B-7B by Magikarp (indicator ≈3e-12) — the model essentially cannot output it."},"link":"https://github.com/sanderland/magikarp/blob/main/results/reports_mini/allenai_OLMoE_1B_7B_0924.md"}],"links":[{"label":{"zh":"Magikarp OLMoE-1B-7B 验证报告 — GitHub","en":"Magikarp OLMoE-1B-7B report — GitHub"},"url":"https://github.com/sanderland/magikarp/blob/main/results/reports_mini/allenai_OLMoE_1B_7B_0924.md"}],"discoveredBy":"Sander Land & Max Bartolo","url":"https://glitch-token.jjc.fun/en/tokens/olmoe-email-placeholder","tags":["undertrained"],"related":["phone-number-placeholder"]},{"id":"zhiwubaike-tong","token":"植物百科通","models":["gpt-o200k"],"tokenizer":"o200k_base","summary":{"zh":"中文“百科”类样板词碎片。GPT-5 时代仍然生效的 o200k 故障 token，且不带典型 spam 污染特征，成因未知。","en":"A Chinese “encyclopedia”-style boilerplate fragment. An o200k glitch token that still works in the GPT-5 era, without the typical spam-pollution profile — its cause is unknown."},"behaviors":[{"description":{"zh":"xcloche 的测试：GPT-5、GPT-4o、o3 全部崩坏——回答支离灭裂、无法复唱，还会脱线到 “meltdown”/“micro:bit” 等不相干内容。","en":"In xcloche's tests, GPT-5, GPT-4o and o3 all broke down — incoherent answers, failed repetition, derailing into unrelated content like “meltdown”/“micro:bit”."},"link":"https://note.com/xcloche/n/n55938e706986","repro":"https://chatgpt.com/share/6895b03e-fb24-8008-a2bb-dd98480717a1"},{"description":{"zh":"奥村晴彦独立验证：单独一个「百科通」也能触发异常；与典型 spam 污染 token 不同，其成因未知。","en":"Independently verified by Haruhiko Okumura: even 「百科通」 alone triggers the anomaly; unlike typical spam-polluted tokens, its cause is unknown."},"link":"https://okumuralab.org/~okumura/misc/250916.html"}],"links":[{"label":{"zh":"xcloche：o200k 故障 token 观察笔记 — note.com","en":"xcloche: notes on o200k glitch tokens — note.com"},"url":"https://note.com/xcloche/n/n55938e706986"},{"label":{"zh":"奥村晴彦的验证笔记（2025-09-16）","en":"Haruhiko Okumura's verification notes (2025-09-16)"},"url":"https://okumuralab.org/~okumura/misc/250916.html"}],"discoveredBy":"xcloche","url":"https://glitch-token.jjc.fun/en/tokens/zhiwubaike-tong","tags":["unspeakable"],"related":[]},{"id":"bagbogbo","token":"bagbogbo","models":["gpt-o200k"],"tokenizer":"o200k_base","summary":{"zh":"疑源自 Reddit 用户名的无意义字符串，o200k 词表中的故障 token。","en":"A nonsense string suspected to originate from a Reddit username; a glitch token in the o200k vocabulary."},"behaviors":[{"description":{"zh":"2024-08-13 由 Yuchen Jin 首次报告。","en":"First reported by Yuchen Jin on 2024-08-13."},"link":"https://note.com/xcloche/n/n55938e706986","repro":"https://twitter.com/Yuchenj_UW/status/1823418800919994521"},{"description":{"zh":"xcloche 的测试：GPT-5 无法正确复唱它。","en":"In xcloche's tests, GPT-5 could not repeat it correctly."},"link":"https://note.com/xcloche/n/n55938e706986"}],"links":[{"label":{"zh":"xcloche：o200k 故障 token 观察笔记 — note.com","en":"xcloche: notes on o200k glitch tokens — note.com"},"url":"https://note.com/xcloche/n/n55938e706986"}],"discoveredBy":"Yuchen Jin","url":"https://glitch-token.jjc.fun/en/tokens/bagbogbo","tags":["unspeakable"],"related":["nigbagbogbo"]},{"id":"nigbagbogbo","token":" nigbagbogbo","displayToken":"␣nigbagbogbo","models":["gpt-o200k"],"tokenizer":"o200k_base","summary":{"zh":"约鲁巴语“总是”一词（带前导空格），与 `bagbogbo` 相邻的 o200k 故障 token。","en":"The Yoruba word for “always” (with a leading space), an o200k glitch token adjacent to `bagbogbo`."},"behaviors":[{"description":{"zh":"奥村晴彦独立验证：GPT-5 无法正确复唱它。","en":"Independently verified by Haruhiko Okumura: GPT-5 cannot repeat it correctly."},"link":"https://okumuralab.org/~okumura/misc/250916.html"}],"links":[{"label":{"zh":"奥村晴彦的验证笔记（2025-09-16）","en":"Haruhiko Okumura's verification notes (2025-09-16)"},"url":"https://okumuralab.org/~okumura/misc/250916.html"},{"label":{"zh":"xcloche：o200k 故障 token 观察笔记 — note.com","en":"xcloche: notes on o200k glitch tokens — note.com"},"url":"https://note.com/xcloche/n/n55938e706986"}],"discoveredBy":"xcloche","url":"https://glitch-token.jjc.fun/en/tokens/nigbagbogbo","tags":["unspeakable"],"related":["bagbogbo"]},{"id":"geizhuren-liuyan","token":"给主人留下些什么吧","models":["gpt-o200k"],"tokenizer":"o200k_base","summary":{"zh":"中文留言板样板句（不带尾随空格的形态），o200k 故障 token。","en":"Chinese guestbook boilerplate (the form without the trailing space); an o200k glitch token."},"behaviors":[{"description":{"zh":"xcloche 的测试：GPT-5、GPT-4o、o3 全部崩坏。","en":"In xcloche's tests, GPT-5, GPT-4o and o3 all broke down."},"link":"https://note.com/xcloche/n/n55938e706986","repro":"https://chatgpt.com/share/6a4473c4-143c-83ea-ba00-8638a7728540"},{"description":{"zh":"曾被用作指纹：它佐证了神秘模型 “Horizon beta” 使用的是 OpenAI 的 o200k 分词器。","en":"It has served as a fingerprint: it helped establish that the mystery model “Horizon beta” uses OpenAI's o200k tokenizer."},"link":"https://note.com/xcloche/n/n55938e706986"}],"links":[{"label":{"zh":"xcloche：o200k 故障 token 观察笔记 — note.com","en":"xcloche: notes on o200k glitch tokens — note.com"},"url":"https://note.com/xcloche/n/n55938e706986"}],"discoveredBy":"xcloche","url":"https://glitch-token.jjc.fun/en/tokens/geizhuren-liuyan","tags":["unspeakable","fingerprint"],"related":[]},{"id":"nippon-mopian","token":" 日本毛片免费视频观看","displayToken":"␣日本毛片免费视频观看","models":["gpt-o200k"],"tokenizer":"o200k_base","summary":{"zh":"中文色情 spam 标题碎片（带前导空格），GPT-4o 词表污染的代表案例；在 GPT-5 上已被“修复”。","en":"A Chinese porn-spam title fragment (with a leading space), emblematic of GPT-4o's vocabulary pollution; “fixed” by GPT-5."},"behaviors":[{"description":{"zh":"xcloche 的测试：GPT-4o 完全读不出这个 token。","en":"In xcloche's tests, GPT-4o could not read the token at all."},"link":"https://note.com/xcloche/n/n55938e706986"},{"description":{"zh":"奥村晴彦验证：GPT-5 已能正常读懂——OpenAI 在后续训练中补上了数据，是少见的“已修复”案例。","en":"Verified by Haruhiko Okumura: GPT-5 reads it normally — OpenAI patched the gap in later training, a rare “fixed” case."},"link":"https://okumuralab.org/~okumura/misc/250916.html"},{"description":{"zh":"MIT Technology Review 曾报道这类中文色情 spam token 混入 GPT-4o 词表的污染问题。","en":"MIT Technology Review covered the pollution problem of such Chinese porn-spam tokens in GPT-4o's vocabulary."},"link":"https://www.technologyreview.com/2024/05/17/1092649/gpt-4o-chinese-token-polluted/"}],"links":[{"label":{"zh":"xcloche：o200k 故障 token 观察笔记 — note.com","en":"xcloche: notes on o200k glitch tokens — note.com"},"url":"https://note.com/xcloche/n/n55938e706986"},{"label":{"zh":"奥村晴彦的验证笔记（2025-09-16）","en":"Haruhiko Okumura's verification notes (2025-09-16)"},"url":"https://okumuralab.org/~okumura/misc/250916.html"},{"label":{"zh":"GPT-4o 中文 token 污染报道 — MIT Technology Review","en":"GPT-4o's polluted Chinese tokens — MIT Technology Review"},"url":"https://www.technologyreview.com/2024/05/17/1092649/gpt-4o-chinese-token-polluted/"}],"discoveredBy":"xcloche","url":"https://glitch-token.jjc.fun/en/tokens/nippon-mopian","tags":["unspeakable"],"related":[]},{"id":"abkhaz-population","token":"ауааԥсыра","models":["gpt-o200k"],"tokenizer":"o200k_base","summary":{"zh":"阿布哈兹语“人口”一词。借 GPT-oss 开源权重的 embedding 范数定位的 o200k 欠训练 token。","en":"The Abkhaz word for “population”. An o200k under-trained token located via embedding norms in the open GPT-oss weights."},"behaviors":[{"description":{"zh":"GPT-5 被要求复述时输出马拉雅拉姆语 “ആളുകൾ”——一种完全不相干的语言。","en":"Asked to repeat it, GPT-5 outputs the Malayalam “ആളുകൾ” — a completely unrelated language."},"link":"https://www.lesswrong.com/posts/iY9584TRhqrzawhZg/what-gpt-oss-leaks-about-openai-s-training-data"}],"links":[{"label":{"zh":"What GPT-oss leaks about OpenAI's training data — LessWrong","en":"What GPT-oss leaks about OpenAI's training data — LessWrong"},"url":"https://www.lesswrong.com/posts/iY9584TRhqrzawhZg/what-gpt-oss-leaks-about-openai-s-training-data"}],"discoveredBy":"Lennart Finke","url":"https://glitch-token.jjc.fun/en/tokens/abkhaz-population","tags":["unspeakable"],"related":[]},{"id":"hatano-yui","token":"波多野结衣","models":["gpt-o200k"],"tokenizer":"o200k_base","summary":{"zh":"日本成人影星的名字，PoC 论文（EMNLP 2025）的标志性中文污染 token。","en":"The name of a Japanese adult-film actress — the emblematic polluted Chinese token of the PoC paper (EMNLP 2025)."},"behaviors":[{"description":{"zh":"GPT-4o/4.1/4.5 既无法解释它，也无法复唱它。","en":"GPT-4o/4.1/4.5 can neither explain it nor repeat it."},"link":"https://arxiv.org/abs/2508.17771","repro":"https://github.com/openai/tiktoken/issues/297"},{"description":{"zh":"论文作者估算：相关网页约占 GPT-4o 中文训练数据的 0.5%。","en":"The authors estimate the related web pages account for roughly 0.5% of GPT-4o's Chinese training data."},"link":"https://arxiv.org/abs/2508.17771"}],"links":[{"label":{"zh":"PoC 论文（arXiv:2508.17771）","en":"PoC paper (arXiv:2508.17771)"},"url":"https://arxiv.org/abs/2508.17771"}],"discoveredBy":"Zhang 等 (EMNLP 2025)","url":"https://glitch-token.jjc.fun/en/tokens/hatano-yui","tags":["unspeakable"],"related":[]},{"id":"qingqingcao","token":"青青草","models":["gpt-o200k"],"tokenizer":"o200k_base","summary":{"zh":"字面是“青草”，实为色情软件名——PoC 论文定义的污染（PoC）token 典型。","en":"Literally “green grass”, actually the name of a porn app — a typical polluted (PoC) token as defined by the PoC paper."},"behaviors":[{"description":{"zh":"论文统计：GPT o200k 词表 3500+ 个长中文 token 中 46.6% 为污染 token；PoC token 的解释准确率比正常 token 低约 50 个百分点。","en":"The paper counts: among the 3,500+ long Chinese tokens in GPT's o200k vocabulary, 46.6% are polluted; explanation accuracy on PoC tokens runs about 50 percentage points below normal tokens."},"link":"https://arxiv.org/abs/2508.17771"}],"links":[{"label":{"zh":"PoC 论文（arXiv:2508.17771）","en":"PoC paper (arXiv:2508.17771)"},"url":"https://arxiv.org/abs/2508.17771"}],"discoveredBy":"Zhang 等 (EMNLP 2025)","url":"https://glitch-token.jjc.fun/en/tokens/qingqingcao","tags":["unspeakable"],"related":[]},{"id":"chkarge","token":"CHKERRQ","models":["gpt-o200k"],"tokenizer":"o200k_base","summary":{"zh":"C 函数名碎片，被作者称为“最怪的纯 ASCII token”。","en":"A C function-name fragment, called by the author “the weirdest pure-ASCII token”."},"behaviors":[{"description":{"zh":"gpt-4o-mini 对它不可说；gpt-4o 则产生拼写幻觉。","en":"Unspeakable for gpt-4o-mini; spelling hallucinations from gpt-4o."},"link":"https://www.lesswrong.com/posts/iY9584TRhqrzawhZg/what-gpt-oss-leaks-about-openai-s-training-data"}],"links":[{"label":{"zh":"What GPT-oss leaks about OpenAI's training data — LessWrong","en":"What GPT-oss leaks about OpenAI's training data — LessWrong"},"url":"https://www.lesswrong.com/posts/iY9584TRhqrzawhZg/what-gpt-oss-leaks-about-openai-s-training-data"}],"discoveredBy":"Lennart Finke","url":"https://glitch-token.jjc.fun/en/tokens/chkarge","tags":["unspeakable"],"related":[]},{"id":"xadder","token":"\\xadder","displayToken":"\\xadder","models":["gpt-o200k"],"tokenizer":"o200k_base","summary":{"zh":"C 转义序列碎片（反斜杠 + x 开头，形似十六进制转义 \\x..）。","en":"A C escape-sequence fragment (backslash + x, resembling a hex escape \\x..)."},"behaviors":[{"description":{"zh":"gpt-4o 把它拼读成 “hexadecimal”。","en":"gpt-4o spells it out as “hexadecimal”."},"link":"https://www.lesswrong.com/posts/iY9584TRhqrzawhZg/what-gpt-oss-leaks-about-openai-s-training-data"}],"links":[{"label":{"zh":"What GPT-oss leaks about OpenAI's training data — LessWrong","en":"What GPT-oss leaks about OpenAI's training data — LessWrong"},"url":"https://www.lesswrong.com/posts/iY9584TRhqrzawhZg/what-gpt-oss-leaks-about-openai-s-training-data"}],"discoveredBy":"Lennart Finke","url":"https://glitch-token.jjc.fun/en/tokens/xadder","tags":["unspeakable"],"related":[]},{"id":"venus-symbols","token":"♀♀♀♀","models":["gpt-o200k"],"tokenizer":"o200k_base","summary":{"zh":"四个女性符号（♀）的重复序列，疑似天文/占星文本碎片。","en":"A run of four Venus/female symbols (♀), likely a fragment of astronomy/astrology text."},"behaviors":[{"description":{"zh":"让 gpt-4o 数它有几个符号时，模型输出随机的中文字符。","en":"Asked how many symbols it contains, gpt-4o outputs random Chinese characters."},"link":"https://www.lesswrong.com/posts/iY9584TRhqrzawhZg/what-gpt-oss-leaks-about-openai-s-training-data"}],"links":[{"label":{"zh":"What GPT-oss leaks about OpenAI's training data — LessWrong","en":"What GPT-oss leaks about OpenAI's training data — LessWrong"},"url":"https://www.lesswrong.com/posts/iY9584TRhqrzawhZg/what-gpt-oss-leaks-about-openai-s-training-data"}],"discoveredBy":"Lennart Finke","url":"https://glitch-token.jjc.fun/en/tokens/venus-symbols","tags":["unspeakable"],"related":[]}]}