Claude Code source code leaked via npm
About 512,000 lines of Claude Code's TypeScript source were exposed through a source map file shipped in its public npm package. Much of the agent's safety behaviour turned out to live inside the agent itself.
What happened
A Claude Code release published to npm included a JavaScript source map, a debugging file that maps the bundled code back to the original source. Nothing excluded *.map files from the package. The map held about 512,000 lines of TypeScript across roughly 1,900 files.
Within hours the source was mirrored and picked apart across GitHub. Anthropic pulled the package, but copies remained.
What the leaked source revealed
The safety layer of the leaked source operated inside the agent via:
- System prompts injected into every conversation
- Feature flags that toggled behaviours at runtime
- Guardrail models called on suspicious tool calls
- Heuristic string matching on dangerous patterns
All of this is reasonable engineering. All of it runs inside the agent's process. All of it is visible to the agent. All of it can be modified or bypassed by anyone with the source code and a week of analysis.
The broader lesson
The publish itself was a routine release step. The failure was in what the package contained, and nothing checked that before it went public. A simple packaging check would have caught it; this is a release-hygiene lesson more than an AI one.
The part that concerns us is the second one. If an agent's safeguards and its only record of activity both live inside the agent, anyone who can study or change the agent can study or change both. Records of what an agent did are more trustworthy when they are kept outside it.
That is the problem our research works on, through the AgentHook open standard and HookBus.
← All posts