At 02:14 Paris time on 8 July, our browser-agent success rate on one patient-portal vendor dropped from 94% to 3%. The vendor had shipped a redesign overnight to every customer simultaneously. 212 portals we serve changed their document-list page at once.
Our agents at the time used learned selectors: for each portal, a stored map from semantic targets ('document list', 'download button') to DOM paths, refined over hundreds of sessions. They were fast and precise, and completely brittle to a redesign. Every one of those 212 maps became wrong at the same moment.
Recovery was manual for the first six hours: a specialist re-taught the agent on one portal, we verified the vendor's new layout was consistent, and we pushed the new map to all 212. Then we discovered that 40 of them had customer-specific customisations that broke the shared map. Those took another eight hours.
Fallback worked as designed. Retrievals that failed on the portal were re-planned to voice or fax, and 96% still completed within SLA. But we burned a day of specialist time and made 180 phone calls we should not have needed.
What we built after: self-healing selectors. Instead of DOM paths, agents now store a semantic description of each target plus a screenshot embedding. When a stored selector fails, a vision-language model locates the target on the new page, proposes a new selector, and a second model verifies it against the description before we act. Confidence below threshold escalates to a human, who confirms with one click. In the two redesigns since, recovery took 11 minutes and 26 minutes respectively, with zero specialist calls.