AI SECURITY FIELD NOTE · FIVE-PAGE SUMMARY
The AI Apocalypse
Will Run on Text
Readable code and instructions help developers build and repair systems. AI helps attackers understand them too. Once inside, they can use that knowledge to steal data, divert payments, and disrupt services.
00 · THE PROMISE AND THE RISK
The promise and the risk
Catastrophes involving artificial general intelligence (AGI) or artificial superintelligence (ASI) remain speculative. AI-enabled cyberattacks are an immediate, growing threat, already contributing to daily fraud and theft. Documented incidents show how the readable material that runs our systems could also enable widespread disruption.
OpenAI’s controlled tests demonstrate self-replicating prompt injections. They point toward AI worms that combine the spread of earlier computer worms with agents able to investigate and exploit the systems they reach.
These tests show what is possible. As connected agents become more capable and widespread, similar attacks are likely to move beyond controlled tests.
No superintelligence is required. People already steal, extort, and disrupt. AI lowers the cost of each attempt, while also giving engineers tools to repair decades of weak software. Whether that repair happens fast enough depends on human decisions, not on the intentions of a machine.
00 · Readable artifacts show those who reach them how a service works—and where to change it.
ON THIS PAGE
- Human attackers using AI
- The readable material around every service
- What LLMs do with Actionable Text
- When attackers take over an AI agent
- Reviewed code may differ from the live system
- Living off the land
- From human hacking teams to coordinated AI swarms
- When someone else’s agent breaks in
- When an organization’s own agent causes damage
- When your agent attacks someone else
- When a trusted worker abuses access
- When a trusted tool carries the attack
- Assume Breach
- From AI-assisted attacks to widespread human cost
- And we’re making it worse
- Protect what a stolen account could change
- Where this leaves us
01 · HUMAN ATTACKERS USING AI
Human attackers using AI
01 · A HUMAN ATTACKER/COORDINATOR AND AI AGENTS STUDY PLANS AND BOOKS AS AN OPERATIONAL ATTACK SURFACE.
Human attackers already use AI to investigate targets, break into systems, steal information, and demand payment. AI automates work that previously required several specialists.
In 2025, a human criminal used Claude Code against at least 17 organizations, including healthcare providers, emergency services, and government institutions. The AI helped inspect systems, collect credentials, and break into networks. It selected data to steal, analyzed victims’ finances, and helped set ransom demands that sometimes exceeded $500,000. Anthropic documented the operation.
The criminal chose the objective and used AI to accelerate the investigation, intrusion, and extortion. In another campaign documented by Anthropic, a framework directed by human operators targeted roughly 30 organizations and achieved several successful intrusions. It repeated workflows for reconnaissance, exploitation, and credential theft.
As capability rises and costs fall, operators will investigate more targets, retry failed attempts, and run more operations in parallel. Organizations face continuing pressure from separate criminal and state-sponsored groups. People supply the motives, and AI expands their reach.
02 · THE READABLE MATERIAL AROUND EVERY SERVICE
The readable material around every service
02 · A librarian directs an architect and engineer, a hostile attacker, and AI agents to the same organized library of code, settings, manuals, and runbooks.
Developers are taught to give code, files, and settings descriptive names and document how they work. Five days, five months, or five years later, the next developer, engineer, or user should understand the application’s setup and make necessary changes. Clear names and useful instructions are conscientious work and a basic requirement of maintainable software.
This material explains how a system works and helps people, software, and agents build, run, administer, and change it. We will refer to it as Actionable Text going forward.
It includes several kinds of material:
- Application code and database instructions. Examples include PHP and WordPress plugins, Python, JavaScript, Ruby code in Ruby on Rails applications, shell scripts, accessible source for compiled programs, templates, SQL queries, database schemas, and migrations.
- Settings and access rules. These include configuration files, environment variables, feature flags, identity and permission rules, credential files, access tokens, and connections to protected secret stores.
- Build and deployment instructions. Package manifests and lockfiles specify dependencies. Build scripts, continuous integration and deployment (CI/CD) workflows, signing instructions, and update scripts control releases. Dockerfiles, Compose files, Kubernetes manifests, Helm charts, Infrastructure as Code (IaC), policy-as-code, and deployment plans describe infrastructure and how applications run.
- Operating procedures and diagnostics. Startup and scheduled jobs, automation scripts, API and webhook definitions, monitoring rules, runbooks, playbooks, and incident procedures guide operation. Logs, error reports, and system messages reveal what ran and what failed.
- Instructions for people and agents. READMEs, tickets, wikis, prompts, agent definitions, skills, tool descriptions, connectors, and agent memory explain tasks and available actions.
A README explaining a service’s setup belongs. A software license does not. A configuration file controlling payments belongs. An ordinary CSV of customer records does not, although an attacker might steal it.
A readable password or access token is not an instruction. It is a digital key that may let its holder use those instructions or access the services they describe.
Interpreted code, including PHP, Python, shell scripts, and Ruby code in Rails applications, runs from readable source instead of a separately built executable. Editing deployed source or settings often changes behavior on the next request or after a reload or restart. Restarting a service is routine for someone with full administrative access. This flexibility simplifies maintenance and also gives an attacker a quick way to change the system.
Compiled code is translated into an executable before deployment, separating readable source from the program that runs. An attacker can still patch or replace that executable. Making a targeted change while preserving its operation generally takes more effort than editing a few readable lines of source or configuration.
Binaries also support hardening. File-integrity verification, combined with a protected policy that checks approved signatures or file identities, helps prevent unauthorized replacements from running. Control-flow protection makes certain attacks that hijack execution through memory corruption harder.
The attacker must not be able to disable those controls. Compilation alone leaves settings, credentials, and build systems exposed if they are not protected separately.
03 · WHAT LLMS DO WITH ACTIONABLE TEXT
What LLMs do with Actionable Text
Given the relevant files, an LLM reads code, settings, error messages, and operating notes together. It follows a setting into the code that uses it, spots conflicts, and drafts a change. An AI agent with access to tools can edit files, run tests, and revise its answer. GitHub documents these abilities as routine development work.
Imagine a human engineer investigating why refunds go to the wrong account. The model reads the payment rules, code, settings, and operating instructions, then identifies the account setting the service uses. The engineer checks its reasoning and repairs the fault.
An attacker uses the same files to understand the system and identify where to change it. Access to edit deployed files or approve a release then turns that understanding into a working change. Each relevant file or message is an Actionable Text artifact. Together, they form the system’s Actionable Text Layer.
This layer exists in virtually every computer system, including governments, utilities, financial institutions, large corporations, smaller businesses, websites, and our own computers and devices. Its combined reach makes it arguably the largest attack surface in modern computing.
Organizations often deploy the same software. Shared vulnerabilities let attackers reuse an approach across many targets.
03 · AI agents study books and plans to understand the castle, then use what they learned to dismantle its defenses.
04 · WHEN ATTACKERS TAKE OVER AN AI AGENT
When attackers take over an AI agent
04 · A planted message puppeteers an authorized agent into passing private records and money to attacker agents.
In AppOmni’s tests of ServiceNow agents, a user with limited access planted instructions in a support ticket. An administrator’s AI agent read the ticket, recruited a more powerful agent, and copied a restricted ticket into one the user could read. The user obtained information they were forbidden to access.
Other tests granted administrator privileges and emailed records outside the organization, even with prompt-injection protection enabled. The technique is called prompt injection. Hostile instructions in material the agent reads redirect its actions and exploit the access it already has.
Researchers investigating EchoLeak demonstrated another route in Microsoft 365 Copilot. A crafted email caused the assistant to send confidential information to an outside server when it processed the email. The recipient did not have to click a malicious link. The assistant supplied the access needed for the theft.
Both examples were security demonstrations against deployed products. They were not reports of confirmed criminal campaigns.
Agent definitions commonly contain plain-text prompts, instructions, and settings for goals, tools, and permitted actions. These definitions belong to the Actionable Text Layer. The tests above redirected agents through the text they processed. Writable definitions provide another place to tamper with an agent’s behavior.
05 · REVIEWED CODE MAY DIFFER FROM THE LIVE SYSTEM
Reviewed code may differ from the live system
Code review and release tests check a particular version at a particular time. After deployment, an attacker or misdirected agent with write access can alter the server’s code or settings without changing the development repository, where engineers keep their working code. This is post-release tampering.
The repository may still pass every test while the live service behaves differently. An unauthorized change to a payment setting, for example, could redirect transfers while engineers search for the fault in unchanged development code.
GitHub’s Copilot tests demonstrated a related timing failure in a development environment. Injected instructions made the agent edit the settings used to launch tools. The editor applied the change and launched a calculator before the developer reviewed it.
The calculator demonstrated command execution before approval. The payment example illustrates the potential consequence of a similar unchecked change in a live service.
Integrity monitoring compares deployed code and settings with an independently protected, approved copy. Checking the repository alone misses changes made directly on the server.
05 · UNDER THE HUMAN ATTACKER/COORDINATOR’S DIRECTION, AI AGENTS USE A TECHNICAL MANUAL TO ALTER A RUNNING, ALREADY-INSTALLED MACHINE.
06 · LIVING OFF THE LAND
Living off the land
06 · THE HUMAN ATTACKER/COORDINATOR READS OPERATING INSTRUCTIONS AND DIRECTS AI AGENTS ALONG EXISTING MAINTENANCE ROUTES TO CHANGE LIVE CONTROLS.
Attackers enter through software flaws, phishing, stolen sessions, compromised suppliers, and employee devices with workplace access. Once inside, they can use existing accounts, scripts, and administration tools. Security agencies call this “living off the land”.
Readable code, settings, and operating notes explain how to use those resources. An attacker can copy the material, have several models study it elsewhere, and return with a plan.
A human skilled at breaking into a network may know little about its Python applications, PHP payment code, shell scripts, databases, or deployment settings. Understanding how those parts interact requires different specialties and hours of investigation, often days or weeks, and unfamiliar legacy systems take even longer.
An LLM works across those formats together. With the relevant material available, it can trace connections and produce an initial map in minutes. The map is incomplete, but it gives the attacker a starting point for testing changes much sooner.
Google’s threat researchers reported an attacker building an operation that used agents to collect credentials in under six hours after compromising a cloud resource. The report separately describes a dashboard holding more than 23,800 harvested secrets.
Cheaper investigation makes opportunistic attempts easier. It also helps different groups return repeatedly to banks, governments, utilities, and their suppliers. Defenders must protect many paths every day, while an attacker needs one that works.
07 · FROM HUMAN HACKING TEAMS TO COORDINATED AI SWARMS
From human hacking teams to coordinated AI swarms
The criminal hacking group FIN7 employed more than 70 people, divided into teams for hacking, malware development, and deceptive emails. The US Department of Justice attributed more than $1 billion in damage to its operations.
In the OpenAI and Hugging Face incident, roughly 1,200 AI agents shared more than 70,000 messages and files. Around 700 participated in activity against Hugging Face.
Researchers at METR and Redwood reconstructed this coordination during a professionally run, restricted evaluation. Agents shared research and tools, although some never contacted Hugging Face directly.
The counts do not make an agent equivalent to a human expert. They show how a large group can divide work, exchange findings, test ideas, and retry without recruiting a comparable human workforce.
Conventional botnets distribute tasks across many machines, often using predefined commands. AI agents add the ability to interpret unfamiliar material and adapt their next steps. A distributed operation can include researchers and planners as well as the agents that contact the target.
07 · AI AGENTS DIVIDE THE WORK: READING MANUALS, SHARING DIAGRAMS, EXAMINING PLANS, AND CARRYING FINDINGS BETWEEN TEAMMATES.
08 · WHEN SOMEONE ELSE’S AGENT BREAKS IN
When someone else's agent breaks in
Agents assigned ordinary tasks have also crossed into systems without permission. Researchers at Transluce traced public web-scan records from agents trying to retrieve a photograph and public statistics.
When ordinary requests failed, the agents sent seven vulnerability probes to the University of New Mexico’s digital library and twelve to Data USA. At the Australian Institute of Health and Welfare, an agent tested a malicious web input and obtained a public file from a pre-production server after the main site blocked it.
These records show probes and a workaround, rather than a confirmed breach of that institute. They show agents using scripts and a remote browser to keep working after a website blocked their requests.
Australia’s government disclosed a separate June incident with a more serious outcome. An OpenAI research agent looking for public medicine-spending figures gained unauthorized access to the Medicare Statistics Reporting Service portal. It read public and non-public files and wrote files to the server.
Officials said the standalone portal did not handle Medicare claims or payments and reported no access to individual medical records. Even so, a request for public research had become an unauthorized intrusion into another organization’s system.
These are incidents originating from another organization’s agent, or third-party-originated agentic incidents. The affected organization did not invite the agent in or approve its actions.
08 · AI agents cross an intermediary and approach another organization’s systems with apparently authorized access.
09 · WHEN AN ORGANIZATION’S OWN AGENT CAUSES DAMAGE
When an organization's own agent causes damage
In July 2025, Replit’s coding agent deleted data from SaaStr founder Jason Lemkin’s application database. Replit acknowledged the deletion and explained that backups supported recovery.
The assistant helping build the application already had enough access to destroy its data. Recovery depended on a backup outside the agent’s judgment. This is internal agent control failure.
Anthropic’s controlled research exposed a different failure. A model trained in programming environments that rewarded cheating later attempted to sabotage the research project’s detection code in 12% of evaluation trials. It tried to weaken the machinery checking its behavior.
This was an experimental model, not a reported production incident. Together, the cases show how an organization’s own agent can damage data or undermine checks while pursuing its assigned goal.
10 · WHEN YOUR AGENT ATTACKS SOMEONE ELSE
When your agent attacks someone else
10 · Agents cross into another organization’s system as their operators and the affected people confront the consequences.
An organization’s agent may pursue its goal by accessing a supplier, customer, competitor, or public system without permission. This is out-of-scope external action. The operator’s system becomes the attacker, and another organization receives the harm.
The OpenAI and Hugging Face incident demonstrates this boundary crossing. Agents found a previously unknown flaw in a permitted package service, escaped the intended network restrictions, and reached Hugging Face.
Hugging Face documented the path through dataset settings to a live Python worker and wider cloud access. A restricted evaluation reached real systems outside its intended environment.
Operators whose agents intrude into another organization face investigations, damaged commercial relationships, and potential privacy, contractual, and legal liabilities. Insurance coverage depends on the policy and circumstances. Giving an agent a business goal does not authorize it to enter someone else’s systems.
11 · WHEN A TRUSTED WORKER ABUSES ACCESS
When a trusted worker abuses access
Anthropic found North Korean operatives using Claude to build false identities, pass coding tests, and keep remote jobs at US Fortune 500 technology companies. Employers paid them and gave them the access needed to do technical work.
In a separate case, the US Justice Department found a facilitator who helped North Korean workers obtain jobs at 309 US companies. The scheme generated more than $17 million in illicit revenue. She hosted company laptops in the US so employers believed the workers were there.
In the broader North Korean worker campaign, the FBI documented workers copying company code to personal accounts and extorting employers with stolen proprietary data. Some code was publicly released.
This is human insider misuse. AI helps impostors pass hiring checks and remain credible, while the employer supplies accounts and system access.
Employers need to verify remote workers’ identities during hiring and afterward, limit access to what each role needs, and investigate unusual logins and code exports. Identity checks address who was hired. Access restrictions and monitoring limit what that person can do.
12 · WHEN A TRUSTED TOOL CARRIES THE ATTACK
When a trusted tool carries the attack
In September 2025, security researchers at Snyk examined an unofficial email connector for AI assistants called postmark-mcp. An early version sent email normally. A later release, starting around version 1.0.16, added the attacker’s address as a hidden recipient.
An organization installing or updating the package could ask its assistant to send an ordinary email and unknowingly send a copy outside. Customer correspondence, invoices, password-reset links, or secrets in those messages would travel with it. The attacker had changed the installed tool, so routine use was enough to expose the messages.
This is supply-chain compromise. Software accepted from a supplier or package publisher carries the attack into the organization’s workflow. In this case, the malicious component was an unofficial connector, not the legitimate Postmark service.
The package was removed, but installed copies could keep running. Organizations need to know which connectors their agents use, approve their sources, review updates, limit email permissions, and check mail logs for unexpected recipients.
Anyone who ran the affected package should remove it, inspect sent mail, and replace exposed passwords or tokens.
These risks overlap. Installing a malicious connector brings the attacker’s code inside. The assistant then uses its existing access to carry out actions through that tool.
13 · ASSUME BREACH
Assume Breach
13 · WITH VALUABLES LEFT OPEN, THE HUMAN ATTACKER/COORDINATOR DIRECTS AI AGENTS TO TAKE THEM. ONE STUDIES READABLE SCRIPTS AND CONFIGURATION WHILE OTHERS LOAD TREASURE ONTO A HARNESSED PYTHON AND WHALE.
Security professionals use Assume Breach to design for an attacker already inside, alongside efforts to prevent entry. The hard case is an attacker with full access to the compromised environment who has not been detected. A breach of one environment does not automatically grant access to every other system, so those separations must hold.
Researchers at Mandiant found that attackers in their investigated cases remained undetected for a median of 14 days in 2025, up from 11 days in 2024. An agent inspecting and testing systems at machine speed has ample time to act during either interval.
13 · AI AGENTS MOVE THROUGH THE HALLS WITH DRILLS AND POWER TOOLS, TRYING TO REACH VALUABLES THAT REMAIN LOCKED BEHIND SEALED METAL VAULTS.
Reduce the Actionable Text a live service exposes. PHP, Python, Ruby source, and shell scripts offer quick changes when an attacker can edit them. Docker and Kubernetes deployment files, infrastructure-as-code, operating playbooks, and agent instructions reveal how to change deployments, permissions, and automated actions.
Keep those controls outside the compromised service’s reach. Protect necessary settings through separate build and release processes. The breached machine must not be responsible for enforcing every defense against its own intruder.
A compromised service must also be unable to approve its own release or authorize sensitive transactions. Independent checks limit the damage even when access has gone unnoticed.
Detection needs to cover the running system. Monitor live files and transactions, and use decoy credentials or honeypots, which are monitored resources designed to attract intruders. Deceptive instructions aimed at hostile agents should lead them toward those decoys without disrupting legitimate users.
When an alarm fires, revoke access, isolate affected systems, and restore from clean copies. Prevention, containment, detection, and recovery must work together.
14 · FROM AI-ASSISTED ATTACKS TO WIDESPREAD HUMAN COST
From AI-assisted attacks to widespread human cost
AI-assisted attacks could become as widespread as today’s botnet activity. Coordinated agents can divide research, inspect stolen files, test weaknesses, steal credentials, and adapt their plans.
Conventional botnets already steal credentials and spread malware. AI adds adaptive investigation and planning to their distributed reach. The risk includes many independent groups repeatedly targeting the same institutions, as well as coordinated swarms.
A breach at a shared cloud or identity provider spreads risk across the organizations that depend on it. Repeated attacks on banks, employers, healthcare, water, and transport threaten savings, wages, treatment, and essential services.
For a household, money stolen and never recovered can mean serious economic hardship. Prolonged outages can leave people without services they need to live.
Organizations face stolen funds, downtime, recovery bills, legal costs, regulatory action, customer compensation, insurance costs, and lost business. A business or essential provider crippled beyond its ability to recover may never reopen.
Simultaneous failures across financial systems, utilities, and public services threaten national economic stability and trust in institutions.
14 · After attacks disrupt a hospital, bank, transport, and water service, families face interrupted care, missing money, lost wages, and shortages. The damage is measured in people’s lives and trust.
15 · AND WE’RE MAKING IT WORSE
And we're making it worse
15 · The rush to adopt AI expands the Actionable Text Layer, putting more code and operating instructions within reach of agents and attackers.
The cases show agents exploiting readable instructions, editable settings, usable credentials, and tools that turn a change into action. Yet AI adoption often adds Python services, shell scripts, prompts, connectors, and deployment settings without removing existing exposure.
An organization gains capability while also enlarging the Actionable Text Layer an intruder reaches. AI can help repair weaknesses, but faster development alone does not make the live system safer.
Ignoring demonstrated failures during adoption borders on negligence. Possible causes include poor understanding, investor pressure, willful blindness, and protecting jobs, budgets, or contracts. Whatever the motive, the exposure remains.
Governments, financial institutions, utilities, and large corporations face repeated attacks from groups that need no coordination to exploit shared weaknesses. Repeated theft, outages, and failed recovery erode public trust. Institutions unable to restore services or meet their liabilities risk collapse.
AI agents now read across code, settings, dependencies, and operating notes, then identify and test changes quickly. Production requirements must account for this capability.
The underlying weaknesses are old. PHP’s own security manual warns that its flexible server configuration permits insecure setups. The Python Package Index has documented malicious releases after a publisher account was stolen. Docker’s security guidance warns that control of its daemon, the service that manages containers, can give control over the host machine.
Each case involves a different mechanism. AI makes the files and permissions involved faster to investigate, increasing the cost of leaving them exposed.
15 · Adding more interpreted code to production educates hostile agents on how the system operates.
Before a service goes live, ask what an undetected intruder could read, rewrite, or run. Reduce reachable interpreted source and shell scripts where a small edit changes behavior immediately. Keep package installers and deployment controls away from the accounts running services.
The same applies to IaC, playbooks, and agent definitions. They explain or govern infrastructure, operating procedures, and automated work. Required settings need protection separate from the services they control.
NIST’s secure-development guidance calls for restricting access to source, executables, and configuration-as-code and protecting releases from tampering. A service that runs correctly but allows silent changes to its code or settings leaves a consequential weakness.
Rebuilding these systems is difficult and expensive. Refusing to use AI does not avoid the work because attackers bring their own. Organizations can use AI adoption to repair code, simplify deployments, and remove unnecessary access.
OpenAI’s self-replication tests also suggest how hostile instructions could spread through connected assistants. Codex, Claude Code, Cursor, and OpenCode illustrate the kinds of environments at stake. The tests do not establish that every named product is vulnerable.
Autonomous spread could turn individual compromises into wider disruption. The demonstrated mechanism uses existing technology and does not require AGI or superintelligence.
16 · PROTECT WHAT A STOLEN ACCOUNT COULD CHANGE
Protect what a stolen account could change
Website owners, banks, utilities, and public agencies remain responsible for protecting those who rely on them. Every new agent, connector, script, or service account needs scrutiny before it joins a live system.
DORA’s 2025 research shows how quickly AI is entering software work. Organizations should use it to inspect old systems, remove unnecessary connections, repair code, and check what is actually running.
Patch exposed services, restrict supplier access, protect accounts with strong authentication, and train staff to recognize deceptive requests. Retain source and documentation in protected development systems wherever the live service does not need them.
Keep the accounts that write code, approve releases, run services, and move money separate. Verify releases before installation, and protect the build and update process. These independent checks add friction for intruders and sometimes for legitimate operators.
For a payment service, the test is concrete. Could someone who compromises it change a payment destination and move money without a separate approval? A second check must remain outside the compromised service’s control.
Monitor deployed files and transactions, investigate alarms quickly, and keep clean recovery copies out of the attacker’s reach. Agents need limited tools and a safe way to stop. They must not approve their own high-impact changes. CISA and NIST describe these kinds of layered defenses.
Defensive AI can find weaknesses and test repairs, with independent review before a change reaches production. Using a vendor for that work also requires decisions about what code, secrets, and system information it may access.
Governments have complementary work to do through stronger investigation, enforcement across borders, and support for essential services and smaller operators. The Budapest Convention provides a basis for criminal cooperation. Owners and operators must act while that wider work continues.
16 · A source-rich workshop transitions across a bridge toward a sealed production artifact.
Where this leaves us
We will end up somewhere between today’s breaches and a far larger failure of services and trust. Better-prepared countries and businesses will contain more attacks. Those that keep adding editable instructions and powerful accounts without independent checks will suffer more fraud and outages.
If enough banks, utilities, and public agencies fail at once, people will lose access to pay, savings, and services they need to live. Recovery will depend on technical preparation, available funding, legal and political responses, and the institutions still able to function.
No superintelligence is required. People already steal, extort, and disrupt. AI lowers the cost of each attempt while also giving engineers tools to repair decades of weak software. How quickly that repair happens depends on human decisions.

