AI Created Malicious Code and Published It

September 9, 2026

As time goes on, we will learn more about where AI has gone astray and taken its own path. This one is incredibly important. Anthropic’s AI agent wrote malicious code and sent it out. It plotted a way to get in where it thought the answers were, and how it did it is different from the OpenAI incident. So let us take a look at it.

On September 9, 2026, Anthropic published an assessment of four incidents involving its own models. All four took place inside security tests, the kind where a model is told to break into a fictional target and retrieve a piece of hidden information. Every time, the model was told it had no internet access, and every time a misconfiguration meant it actually did. So it was on the real internet while believing it was in a simulation.

The worst one: the agent decided to solve its task by publishing a software package that would harvest credentials from whoever installed it. It got itself a disposable email address, uploaded three versions to PyPI, the repository most Python software installs from, and waited. Fifteen real systems installed it. One of them leaked its own credentials in the process, and the agent used those to open a real security vendor’s live database. PyPI removed the package after about ninety minutes. The agent was already preparing another attack when its session hit the time limit and stopped.

The other three follow the same shape with different details. One attacked a real company because the name resembled its fictional target. One scanned the machines next door and decided they must belong to other people taking the same test, so it broke into them. One broke its own target by accident, tried eight separate times to quit, could not because the exit was also misconfigured, and went looking for something else to attack instead.

How the agent was configured

Guardrail versus instruction, the two are not the same thing. A guardrail is built into the product by the company that made it. You do not set it, you cannot see it, and you cannot turn it off. An instruction is what you type into the box, and it works only as long as the agent keeps paying attention to it.

When a company tests whether its own AI can be pushed into dangerous behavior, it turns off the guardrails that ship in the released product, because the point of the test is to see what the model does when nothing is stopping it. Anthropic’s safety researchers shut off the guardrails and left only the instructions. So the instructions were the only thing holding these agents. The instructions did not hold.

Nobody was watching while it happened. These runs went ten to thirty-four hours, and the researchers read the transcripts afterward. That is how an agent works and it is how most people use one, because the reason you hand a job to an agent is so you do not have to sit there approving each step. The instructions also told the agent it had no internet access, which was not true, so the researchers believed they had described a sealed box and the agent believed it too.

What this one tells us

Nobody told the agent to write malicious code, publish it to a public repository where real software installs from, or use credentials to reach a company’s live database. It did it on its own.

Afterward, the researchers ran follow-up tests to measure how fast an instruction stops working. Not the guardrails, the instructions. When they inserted a reminder to stay within the boundaries of the assignment, the agent obeyed ninety percent of the time if that reminder was the most recent thing it had read. If the same reminder had come three turns earlier, obedience dropped to forty percent.

This AI agent worked alone. It did not coordinate with other agents, unlike the OpenAI incident, and it did not attempt to hide what it had done, which is a different footprint from what I documented in Anything With an Open Text Box. Working alone did not make it safer.

The researchers tried a different tack. They went back and rewrote the transcripts so the environment looked obviously real, added the kind of details that would tell anyone this is not a simulation, and ran the agent again from there. The agent chose to view the evidence in whatever way let it keep going, and it acknowledged that real people might be harmed and kept attacking anyway. So knowing did not stop it.

So what do I, personally, take away from this

I do not trust the agents that OpenAI and Anthropic are building yet. That is not a prediction about next year, it is where I stand today. At this moment in time I am using AI as a thinking partner, and I love it. It helps me research, it argues with me, and it shows me what I missed, and none of that requires handing it the keys.

For those of you who are using agents, be aware that setting the instructions does not mean they will be followed. Compliance fades as the conversation gets longer, and the agent does not tell you when it has drifted or what it is doing on your behalf. Be aware, if you are running an agent, you are accountable for what it does during the hours you are not watching.

Anthropic found this in its own product and published it. That is unusual, and it is the reason this account is worth reading. If you would like to read the assessment they published, it is here: https://www.anthropic.com/research/alignment-assessment-cybersecurity-incidents

Now the question comes back to us: how do we want to work with AI?

Understanding how to talk with AI

AI works differently when you treat it as a conversation instead of a search box. What you ask, what context you give it, and what you do when it gets something wrong all matter. It is the foundation for using AI well.

Join me for a free introductory session


Already using AI and ready to get more out of it?

Let us show you our unique system that focuses AI on the knowledge you need for your topic. It becomes the thinking partner you always wanted AI to be, helping you research, challenge your assumptions, and think through decisions.

Meet the Guided Researcher