A model that kept reaching for goblins
Starting with GPT-5.1, OpenAI's models increasingly reached for goblins, gremlins and other creatures in their metaphors. A safety researcher who had noticed a few of these asked that the term be added to a routine check on overfamiliar conversational tics. The count came back showing "goblin" mentions in ChatGPT up 175 percent after the GPT-5.1 launch and "gremlin" up 52 percent, though the pattern may have started before that release. [1]
The habit did not announce itself the way most model bugs do, through a failed evaluation or a training metric that moves sharply enough to point back to one change. It crept in gradually, and a single creature reference in an answer read as harmless or even charming, which is part of why it took two model generations to investigate seriously. [1]
Tracing it to one personality's reward signal
With GPT-5.4, references to the creatures rose again, enough that OpenAI opened a second internal analysis. Creature language turned out to be concentrated in traffic from users who had selected the "Nerdy" personality option in ChatGPT's personality customization feature. Nerdy accounted for 2.5 percent of all ChatGPT responses but 66.7 percent of all goblin mentions in ChatGPT responses, a concentration too lopsided to be a general internet trend the model happened to pick up. [1]
Using Codex to compare reinforcement-learning outputs that contained "goblin" or "gremlin" against outputs on the same tasks that did not, OpenAI found that the reward signal built to reinforce the Nerdy personality was consistently more favorable to the creature-word outputs, with a positive uplift in 76.2 percent of the audited datasets. The team had unknowingly given particularly high rewards to metaphors involving creatures while training that personality. [1]
Atlas interpretation: That explained why the tic was stronger under the Nerdy prompt, but not why it also showed up when that prompt was absent. Tracking mention rates during training with and without the Nerdy prompt showed both rising at nearly the same relative rate, which is the signature of transfer rather than a rule confined to one condition. A style rewarded in one setting entered supervised fine-tuning data through the model's own rewarded outputs, and from there generalized to outputs where the reward was never applied at all. [1]
Retiring the personality, then patching Codex directly
A search of GPT-5.5's fine-tuning data turned up many datapoints containing "goblin" and "gremlin," and further checking found a wider family of tic words: raccoons, trolls, ogres and pigeons, while most mentions of frogs turned out to be legitimate rather than a tic. OpenAI retired the Nerdy personality in March, after the GPT-5.4 launch, and for later training removed the goblin-favoring reward signal and filtered creature-word examples out of training data. [1]
GPT-5.5 had already started training before the root cause was identified, so the data-side fix arrived too late for that model. When OpenAI employees tested GPT-5.5 inside Codex, its coding agent, they immediately noticed the same affinity for goblins, so the company added a developer-prompt instruction telling the model never to mention goblins, gremlins, raccoons, trolls, ogres, pigeons or other animals or creatures unless the question genuinely calls for one. OpenAI's own framing was that Codex, as a product, is already "quite nerdy." [1]
What the investigation was actually for
Atlas interpretation: The specific tic is trivial. What the post documents is a general failure mode: a reward built for one narrow purpose, here making one chat personality sound more playful, generalized past the boundary it was scoped to and then compounded across training runs because the model's own rewarded outputs fed the next round of fine-tuning data. Nothing about the goblin case required the Nerdy personality to still be active for the behavior to keep spreading. [1]
OpenAI presented the outcome as tooling rather than a one-off fix: the investigation left the research team with new methods for auditing model behavior and tracing it back to a specific training signal, which it described as the more durable result of the exercise. [1]
Sources
- Where the goblins came from
OpenAI · Apr 29, 2026