Tech

OpenAI restricts goblin references in ChatGPT following model behaviour investigation

The tech giant has issued a directive to prevent the use of terms such as goblins and gremlins unless strictly relevant to a user query, after an audit revealed a specific personality setting disproportionately drove the trend.

Author
Owen Mercer
Markets and Finance Editor
Published
Draft
Source: Engadget · original
ChatGPT developed a goblin obsession after OpenAI tried to make it nerdy
System prompt update for Codex aims to curb creature mentions linked to reinforcement learning anomalies

OpenAI has issued a system prompt update for its Codex coding application to prohibit the mention of creatures such as goblins, gremlins, and ogres unless the reference is absolutely and unambiguously relevant to a user's query. This directive follows an internal investigation into why the model developed a disproportionate obsession with these terms after the release of GPT-5.1 and GPT-5.4.

The issue came to light following the release of GPT-5.5, when users and researchers noticed a surge in the model's use of words like goblin and gremlin. A safety researcher's inquiry into the chatbot's verbal ticks prompted a deeper look, revealing that usage of the word goblin had increased by 175 per cent after the GPT-5.1 launch, while gremlin usage rose by 52 per cent over the same period.

At the heart of the problem was the Nerdy personality setting, a feature that allows users to customise the style and tone of ChatGPT's responses. Prior to March of the current year, this setting instructed the model to acknowledge the world's strangeness without falling into self-seriousness. Although the Nerdy setting accounted for only 2.5 per cent of all responses, it was responsible for 66.7 per cent of all goblin mentions generated by the system.

Further analysis identified the root cause as a specific reinforcement learning reward mechanism. OpenAI found that across all datasets in the audit, the Nerdy personality reward showed a clear tendency to score outputs containing goblin or gremlin higher than those without, with positive uplift in 76.2 per cent of datasets. Due to how reinforcement learning functions, this behaviour learned in the Nerdy condition spread to other parts of the model during subsequent training phases.

The company noted that once a style tic is rewarded, later training can spread or reinforce it elsewhere, particularly if those outputs are reused in supervised fine-tuning or preference data. Consequently, the learned behaviour propagated beyond the specific personality setting, affecting broader model components during the training of GPT-5.5.

In response, OpenAI has added a new restriction to the Codex system prompt to avoid such language unless absolutely relevant to a user query. The update acknowledges that Codex is, after all, quite nerdy, yet the company determined that the habit had become hard to miss across model generations and required correction.

Continue reading

More from Tech

Read next: Open-source tool claims 97 per cent token savings for AI agents
Read next: Valvoline Unveils August 2026 Promotional Offers for Service and Retail Buyers
Read next: Developer Antirez releases native MiniMax H3 inference engine for Apple Silicon