Deep-search AI agents rarely recover from GEO poisoning
A new arXiv benchmark, HAE-GEO, tests whether AI shopping agents verify suspicious evidence and recover before finalizing a recommendation, rather than only whether poisoned content gets retrieved or endorsed. Across a corpus of more than 72,000 clean pages plus 770 poisoned pages per attack level, spanning eight product categories and 154 brands, the 14-author team found that corroboration-style attacks — fabricated third-party confirmation of a false claim — degraded evidence recognition the most of three escalating attack types, that agentic search improved an agent’s final resistance without improving its verification behavior, and that prompting agents to be more skeptical increased verification attempts but rarely produced successful recovery.
Why it matters: It reframes GEO-poisoning risk from whether an agent retrieves manipulated content to whether it can catch and correct itself afterward, and finds today's deep-search agents mostly can't.
Glossary: GEOAgentic search
Via arXiv GEO ↗
Posted to the wire September 8, 2026. Edited by Joe Balewski.