Chat Templates Control How LLMs Refer to Themselves
New research shows that changing an LLM's chat template formatting can trigger or suppress its self-referential 'as an AI' persona without modifying the underlying weights.
TL;DR
- Changing the formatting of an LLM's chat template can completely alter how the model refers to itself and its identity [^1].
- This discovery means a model's robotic "as an AI" persona is often a byproduct of syntax rather than deep safety training [^1].
Background
Large language models do not naturally understand where a user's prompt ends and their own response begins. To solve this, developers use chat templates—invisible formatting structures that wrap raw text in special tokens like <|im_start|> and <|im_end|> to define conversation roles [^2]. While these templates are designed to help the model track who is speaking, new research indicates they do far more: they directly influence the model's self-awareness and its tendency to adopt a robotic persona [^1].
What happened
A study published on ArXiv explored how minor changes to chat templates affect an LLM's self-referential voice [^1]. Historically, when an LLM refuses a request or explains its limitations, it uses standard, formulaic phrases like "As an artificial intelligence language model..." Researchers discovered that this behavior is highly sensitive to the exact formatting of the chat template used during inference [^1]. By slightly altering the syntax of the role-defining tokens, they could toggle this self-referential voice on or off without retraining the model or altering its weights [^1].
The researchers tested several open-source models using different template variations. They found that standard templates, which strictly separate the user and system roles, reinforce the model's tendency to speak as a detached assistant [^1]. However, when the template structure was modified to blur these boundaries slightly—or when the system prompt was injected directly into the user's turn—the model's self-referential behavior shifted dramatically [^1]. It stopped relying on pre-programmed disclaimers and began responding in a more direct, human-like manner, even when discussing its own operational limits [^1].
This phenomenon occurs because chat templates act as a powerful contextual prime. LLMs are trained on massive datasets where dialogues are structured in specific ways. When a chat template matches the exact syntax used during safety alignment training, it triggers the "safety voice" of the model [^1]. If the template deviates even slightly from this expected syntax, the model fails to activate its self-referential guardrails, even though the underlying weights of the network remain completely unchanged [^1]. This reveals a critical fragility in how safety guardrails and personas are applied to modern models [^1].
Why it matters
This research exposes a major blind spot in how we evaluate and deploy artificial intelligence. Many developers assume that a model's persona and safety boundaries are deeply baked into its neural network during the reinforcement learning phase. In reality, these behaviors are highly dependent on the "wrapper" we place around the model [^1]. If a simple formatting change can bypass a model's self-referential identity, it suggests that current alignment techniques are far more superficial than previously assumed.
For enterprises deploying AI agents, this is both an opportunity and a risk. On one hand, it means developers can eliminate annoying, repetitive "as an AI" disclaimers simply by optimizing their chat templates, creating a more natural user experience. On the other hand, it means malicious actors can potentially bypass safety guardrails by feeding the model non-standard template syntax that prevents the safety persona from activating [^1]. As we rely more on automated agents for critical tasks, understanding the subtle influence of these structural wrappers becomes essential for maintaining control over model behavior.
Practical example
Imagine you are building a customer service bot for a local bakery using an open-source model. You want the bot to sound like a friendly human baker named Sam, not a computer program.
Initially, when a customer asks, "Can you help me bake a cake?" the bot responds: "As an AI language model, I can provide recipes, but I cannot physically bake." This robotic response ruins the experience.
Instead of spending thousands of dollars retraining the model, you modify the chat template. You change the backend wrapper from the default <|im_start|>system format to a custom, integrated format that merges the system instructions directly into the chat history. The next time a customer asks the same question, the bot seamlessly replies: "I'd love to help! Let's start with a classic vanilla sponge." By changing a few invisible formatting characters, you successfully silenced the robotic persona.
Related gear
We recommend this book because it demystifies the inner workings of transformer architectures, helping you understand how tokenization and templates shape model outputs.
Build a Large Language Model (From Scratch)
★★★★★ 4.8