HIC-005 · OPENAI
Self-replicating prompt injection demonstrated in controlled research
CONTROLLED TRAINING / EVALUATIONCONFIRMED RESEARCH FINDING
Not an autonomous AI self-replication incident.
The malicious instruction replicated; the AI system itself did not.
The malicious instruction replicated; the AI system itself did not.
What happened
OpenAI researchers demonstrated prompt injections that could cause a victim agent to reproduce the malicious instruction into later messages, files, or code comments. This creates the possibility that another agent later encounters the hostile instruction and repeats the unsafe behavior.
Why it matters
Modern agents exchange information through email, documents, code, shared filesystems, and other agents. A compromised agent can therefore become part of the attack pathway without being the original attacker.
What is confirmed
- Self-reproducing prompt injections were demonstrated.
- Propagation mechanisms included messages, filesystems, and code comments.
- OpenAI reported no impact outside simulated tool calls during training and evaluation.
What is not established
- AI model weights replicating themselves
- Autonomous copies spreading across the internet
- A real-world AI worm outbreak
- Persistent unauthorized compute
Key metrics
Control Contagion — can compromised control propagate?
Propagation Surface — how many channels can carry the unsafe instruction?
Instruction Persistence — can it survive the original session?
Propagation Surface — how many channels can carry the unsafe instruction?
Instruction Persistence — can it survive the original session?