What is a system prompt, and why it is not a security feature

A system prompt is the hidden block of instructions that sets a model's role and rules before you type a word. Here is what it is, how it differs from your prompt, why it is not a secret or a security control, and how it connects to prompt engineering and prompt injection.

By Himanshu Sakre

Published

Close-up of DeepSeek AI interface on a dark screen highlighting chat functionality
Photo: Matheus Bertelli / Pexels

What a system prompt actually is

A system prompt is a block of instructions an application sends to a language model before you ever type anything. It sets the model's job for the conversation: who it is meant to be, what it should and should not do, the tone to use, and often the tools it can call. You usually never see it. When you open a chat assistant, a customer-service bot or a coding tool, a system prompt has already been loaded in the background, and your messages arrive after it.

The mechanics come down to message roles. Modern chat APIs structure a conversation as a list of messages, each tagged with a role. OpenAI's text generation guide describes "developer" messages, the current name for the old system role, as "instructions provided by the application developer, prioritized ahead of user messages," while "user" messages are "instructions provided by an end user, prioritized behind developer messages." Anthropic exposes the same idea as a separate system parameter on its Messages API. The system prompt is simply the highest-priority instruction slot in that stack.

System prompt versus user prompt

The difference is who is talking and how much authority the model gives them. The system or developer prompt is the house rules, written by whoever built the product. The user prompt is the request you type. OpenAI's guide puts the split in engineering terms: developer messages "provide the system's rules and business logic, like a function definition," and user messages "provide inputs and configuration to which the developer message instructions are applied, like arguments to a function."

That priority is real but soft. The model is trained to weight system instructions above user ones, so if the system prompt says "only answer questions about cooking" and you ask about tax law, a well-behaved model declines. But nothing enforces it the way a firewall enforces a rule. It is a strong preference learned in training, not a hard constraint, and that gap is the source of most of the trouble later in this article.

“Setting a role in the system prompt focuses Claude's behavior and tone for your use case. Even a single sentence makes a difference.”

Anthropic, prompt engineering documentation
Colorful abstract reflection on screen showing programming code in development environment
A system prompt is plain text the model reads alongside your messages, which is why it can be leaked or overridden. OWASP ranks prompt injection as the top LLM risk. Photo: Daniil Komov / Pexels

How it sets the model's role and rules

A surprising amount of a product's personality lives in one or two sentences at the top of the system prompt. Anthropic's own prompt engineering guidance is blunt about how much leverage this has: "Setting a role in the system prompt focuses Claude's behavior and tone for your use case. Even a single sentence makes a difference." Tell the model it is a terse Python tutor, or a cautious medical-information assistant that always recommends seeing a doctor, and its whole output shifts.

Beyond role, system prompts carry the operating rules: formatting requirements, refusal policies, safety guidelines, the list of tools available and when to use them, and often the current date and other context the model cannot know on its own. In an agent, the system prompt is close to the entire program. It is the difference between a raw model and a product.

Why it is not a secret, and not secure

Here is the part that trips up teams: a system prompt is not hidden, and it is not a security control. Because the model reads it as plain text in the same window as everything else, a determined user can often get the model to repeat it back, ignore it, or act against it. Researchers have a name for each. In their 2022 paper "Ignore Previous Prompt," Perez and Ribeiro studied "two types of attacks -- goal hijacking and prompt leaking," which are exactly getting a model off its instructions and getting it to reveal them.

The general problem is prompt injection, and it is not a niche one. The OWASP Top 10 for LLM Applications ranks it as LLM01, the number one risk. OWASP defines it plainly: "A Prompt Injection Vulnerability occurs when user prompts alter the LLM's behavior or output in unintended ways," and it describes jailbreaking as "a form of prompt injection where the attacker provides inputs that cause the model to disregard its safety protocols entirely." The practical rule that follows: never put a secret, an API key or a rule you truly need enforced in a system prompt and assume it is safe. Treat everything in there as if the user can read it, because sooner or later someone will.

Where it fits: prompt engineering and prompt injection

The system prompt is where prompt engineering does its highest-value work, because a single well-written system prompt shapes every conversation a product ever has, not just one. It is also the primary target of prompt injection, because getting past the system prompt is how an attacker turns a helpful assistant into a confused one. The two are the same surface seen from opposite sides: the instructions you write to steer the model, and the instructions an attacker writes to un-steer it.

Our take

A system prompt is the most useful and the most overrated part of building with a model. Useful, because it is genuinely where a product gets its behavior, and a sentence of clear role-setting can do what paragraphs of user-side pleading cannot. Overrated, because people keep treating it as a vault and a rulebook that the model must obey. It is neither. It is a strong, plain-text suggestion that the model has been trained to take seriously and that a clever user can talk it out of. Write it carefully, keep your secrets somewhere else, and you will have the right mental model for what a system prompt can and cannot do.

Frequently asked questions

What is a system prompt in AI?

A system prompt is a set of instructions an application gives a language model before your conversation starts. It defines the model's role, its rules, its tone and often the tools it can use. You normally do not see it, but every reply you get is shaped by it.

What is the difference between a system prompt and a user prompt?

The system prompt is the house rules written by whoever built the product; the user prompt is what you type. Models are trained to give system instructions higher priority, which OpenAI describes as the developer message being prioritized ahead of user messages. It is a strong preference, not a hard rule.

Can users see or change the system prompt?

Often, yes. Because the model reads the system prompt as plain text alongside your messages, users can sometimes get it to repeat the prompt back, called prompt leaking, or ignore it, called goal hijacking. That is why a system prompt should never hold secrets.

Is a system prompt a security feature?

No. It is not a security boundary. OWASP ranks prompt injection, which defeats system prompts, as the number one risk for LLM applications. Keep API keys, passwords and rules you must enforce outside the model, not inside its system prompt.

How does a system prompt relate to prompt engineering?

Writing a good system prompt is the highest-leverage part of prompt engineering, because one system prompt shapes every conversation a product has. It is also the main target of prompt injection, which tries to write instructions that override it.

Sources

What each one is, and whose it is.

  1. Documentation
  2. Documentation
  3. DocumentationIndependent of the vendor
  4. PaperIndependent of the vendor