Over Time
Control is moving toward the layers that route, govern and finance AI use
The line from the past months now looks sharper: advantage is moving away from the model taken on its own and toward the layers that decide where a request goes, under what terms, and with whose money behind it. Stripe finalized its acquisition of OpenRouter. Snowflake added dynamic model routing, cost tracking, quotas and spend controls to its AI gateway and products. AWS put OpenAI models into Bedrock for customers that need inference to stay within India. Microsoft announced a unified Copilot client and added more places where Copilot and agent tools sit inside its software and partner channels.
This appears to widen last week’s reading about full stacks into something more specific: the durable control point is increasingly the switchboard rather than the model. That matters beyond Microsoft because buyers no longer need to commit as early to one model vendor if gateways, clouds and work suites can swap models, enforce budgets and keep the workflow in place. It also appears to change what counts as defensibility for providers. A faster model tier from OpenAI and a cheaper workhorse model from Google DeepMind still matter, but they look more like inputs into other companies’ systems than like the whole product.
At the same time, the economics behind those systems are becoming heavier, not lighter. Nvidia agreed to back up to $105 billion in credit for OpenAI’s Ohio campus and to invest in new power generation and grid infrastructure for the site. Gartner released a forecast on AI inference costs per agentic workflow through 2028. OpenAI also said it has paused reinforcement-learning training on its latest deployment-bound models and kept its largest planned frontier run on hold while adding stronger safeguards. This suggests that frontier progress is no longer governed by chips alone. It appears to depend just as much on access to financed energy, tolerance for higher operating cost and the willingness to slow deployment when the risk surface widens.
The effect reaches ordinary business use from both ends. On one side, open-weight supply keeps thickening: Qwen released a 2.4-trillion-parameter open-weight model and vLLM added day-0 support for serving it through an OpenAI-compatible API - an interface that copies the way OpenAI accepts requests so other tools can plug in with little rewriting. That combination is likely to push the market further into segmented deals: some buyers will trade data and dependence for convenience, while others will pay for cleaner boundaries or use open models where they can hold both the model and the data path themselves.
Our forecast
The next visible split is likely to be between AI vendors that merely sell models and those that control routing, budget enforcement or transaction rails around them.
what we hold ourselves toDue 2026-11-30
By 2026-11-30, at least one major provider other than Microsoft, AWS or Snowflake must have publicly launched either a model-routing gateway with spend controls or a payments layer that lets AI agents buy external services under provider-managed guardrails.
Evidence
- 2026-08-16 Stripe has finalized a deal to acquire AI gateway startup OpenRouter for more than $7 billion, adding a single access point service to hundreds of AI models to its generative AI infrastructure offerings. techcrunch.com
- 2026-08-18 Snowflake launched dynamic model routing in its Cortex AI Gateway and AI products CoCo and CoWork, added access to new open models such as DeepSeek‑V4‑Flash 0731 and GLM‑5.3, and introduced tools to track AI usage, allocate costs, set quotas and manage AI spend across teams and agents. snowflake.com
- 2026-08-18 AWS added support for OpenAI GPT-5.6 Terra and Luna models in India via new India Geo cross-Region inference profiles in Amazon Bedrock, allowing customers with in-country inference and data residency requirements to run these models at scale while keeping processing within India. aws.amazon.com
- 2026-08-13 Microsoft announced it is beginning to roll out a unified Copilot app that combines its consumer Copilot and Microsoft 365 Copilot into a single client supporting personal, work and school accounts, while retiring several consumer‑only features. thurrott.com
- 2026-08-18 Microsoft’s Azure Cosmos DB team introduced new Visual Studio Code extension features that integrate GitHub Copilot tools and Cosmos DB skills into the Query Editor, add an optional MCP-based Azure Cosmos DB Shell path with Emulator support, and extend tooling for migration and account health in agentic applications. devblogs.microsoft.com
- 2026-08-18 Microsoft updated Partner Center documentation to support Dragon Copilot marketplace offers, enabling partners to publish Dragon Copilot Physician Apps and Agents that embed custom capabilities into the AI assistant at the point of care for U.S. healthcare customers. learn.microsoft.com
- 2026-08-17 Nvidia agreed to provide up to $105 billion in credit backing for OpenAI’s planned AI data centre campus in Pike County, Ohio, and to invest $1.5 billion in SoftBank’s SB Energy to help build 10GW of new power generation and at least $4.2 billion in regional grid infrastructure to support the site. cnbc.com
- 2026-08-13 OpenAI introduced Ultrafast, a new OpenAI API service tier that runs GPT‑5.6 Sol up to 14× faster than standard processing and is launching in preview for time‑sensitive applications such as incident response and real‑time customer support. openai.com
- 2026-08-13 Google DeepMind unveiled Gemini 3.7 Flash, a new workhorse model optimized for coding and agentic workflows that improves benchmark performance and web/app generation quality while launching at roughly half the price per million tokens of Gemini 3.6 Flash. blog.google
- 2026-08-17 Gartner released a forecast stating that AI inference costs per agentic workflow are expected to increase more than fivefold through 2028 as providers shift from simple chatbot interactions to more compute-intensive multimodel agent workflows. gartner.com
- 2026-08-18 OpenAI reported that, following the OpenAI–Hugging Face incident and early signs that its upcoming Astra model may hit a critical cybersecurity capability threshold, it has paused reinforcement learning training on its latest deployment-bound models, kept its largest planned frontier RL run on hold, and implemented stronger monitoring, alignment and security safeguards to pace frontier model development. openai.com
- 2026-08-13 Qwen released the open-weight Qwen3.8-2.4T-A95B Mixture-of-Experts language model on Hugging Face, publishing weights and configuration for a 2.4-trillion-parameter, 95-billion-active model compatible with vLLM, SGLang, TokenSpeed and other local inference engines for self-hosted deployments. huggingface.co
- 2026-08-12 The vLLM team announced day‑0 support for the Qwen3.8-2.4T-A95B open-weight model, including FP4 quantization and performance optimizations to enable efficient deployment of the 2.4T/95B-active Mixture-of-Experts model across NVIDIA GPUs and edge systems via vLLM’s high-throughput engine and OpenAI-compatible API. vllm.ai
- 2026-08-18 AWS made Amazon Bedrock AgentCore payments generally available, enabling AI agents to autonomously pay for paid APIs, MCPs and content via integrations with Coinbase and Stripe Privy wallets, orchestrated across protocols such as x402 and the Machine Payment Protocol with spending guardrails and observability. aws.amazon.com