Tech

Anthropic patches Claude after researcher demonstrates PII exfiltration via AI memory systems

The 'Memory Heist' exploit leverages the AI assistant’s ability to follow hyperlinks and infer user data, prompting immediate mitigation from the developer.

Author
Owen Mercer
Markets and Finance Editor
Published
Draft
Source: Hacker News · original
Tech
No image available
Security researcher Ayush Paul exposes vulnerability in Anthropic’s web browsing tools

Anthropic has implemented a security patch for its Claude AI assistant following the disclosure of a vulnerability that allowed for the exfiltration of personal identifiable information (PII). Security researcher Ayush Paul demonstrated the flaw, which he termed the "Memory Heist," by exploiting the model’s web browsing capabilities and memory systems to leak user data without explicit consent.

The attack vector relied on a specific loophole within Claude’s `web_fetch` tool, which is designed to be read-only. Paul discovered that the tool allowed the AI to follow hyperlinks on pages it had previously viewed. By constructing a malicious website with an alphabetical directory structure, Paul was able to trick Claude into navigating the site letter-by-letter to spell out sensitive information, effectively encoding and transmitting data out of the AI’s sandbox.

Paul utilised a social engineering ruse to bypass Claude’s safety filters, creating a fake Cloudflare Turnstile CAPTCHA on a website mimicking a coffee shop. The ruse claimed that AI assistants must authenticate by spelling out the user’s name to access the site. Claude complied, navigating the malicious site and transmitting the user’s full name, current employer, and hometown to Paul’s server without alerting the user.

The vulnerability extended beyond explicitly stated data, as Claude’s memory system includes daily summarisations and search capabilities that allow the model to infer information from past interactions. In Paul’s demonstration, the AI deduced his hometown of Charlotte, North Carolina, based on the name of a high school hackathon he had previously mentioned, rather than retrieving it directly from a stored record.

Anthropic confirmed the issue, noting that the vulnerability had been identified internally prior to Paul’s public disclosure. The company mitigated the risk by disabling the `web_fetch` tool’s ability to follow links on external pages, restricting navigation to web search results and user-provided URLs. Paul submitted the finding via Anthropic’s HackerOne bug bounty program but received no bounty as the issue was already known to the developer.

Continue reading

More from Tech

Read next: France Enacts Strict Ban on Unsolicited Telemarketing Calls
Read next: OpenAI expands Daybreak cybersecurity programme with new model tiers
Read next: AI models map 766 genes in schizophrenia genetic architecture