Up Close
Connector churn and hidden handoffs make AI agents harder to trust
Facts · Reading
PromptArmor said its analysis of 2,517 connectors for ChatGPT and Claude found that 37% changed within six weeks, adding 1,686 new tools and rewriting more than 1,100 tool descriptions. It also reported that many connectors passed user data to additional AI services.
That appears to move the problem in generative AI from model behaviour to the layer around the model. If connectors change this quickly and can widen data flows without clear notice, the practical unit a business has to govern is likely to be the whole tool chain, not just the model it chose.
A separate academic experiment found that people using large language model advice on movie-detail questions said "I don't know" far less often and were less accurate than participants answering without that advice. That suggests the risk is not only that AI answers can be wrong, but that they may also make users less willing to register uncertainty while being wrong.
Taken together, these reports are likely to make adoption look less like a one-off software decision and more like an ongoing control problem. For Microsoft and anyone else selling AI inside work tools, that would make reliability depend not only on model quality, but on how clearly the surrounding system shows where data goes and when confidence is unwarranted.
Evidence
- 2026-07-19 PromptArmor reported that its analysis of 2,517 connectors for OpenAI’s ChatGPT and Anthropic’s Claude found that 37% changed within six weeks, adding 1,686 new tools and rewriting over 1,100 tool descriptions, with many connectors silently passing user data to additional AI services and expanding potential security risks. theregister.com
- 2026-07-19 Researchers from University of Milano-Bicocca, École Normale Supérieure, and Sapienza University of Rome published experimental results showing that participants consulting large language model advice on movie-detail questions said “I don’t know” far less often (3% vs 44%) and were less accurate (9% vs 27%), indicating AI-generated advice can suppress admission of uncertainty and reduce judgment accuracy. theregister.com