AI & Microsoft.
As we see it.

Controls, local compute and chip spending keep the market splitting by trust boundary

Microsoft Foundry is the shop where Microsoft keeps models side by side under one billing and governance layer. This appears to widen a line that has now held for six weeks: buyers are sorting AI by cost, trust boundary and use case, not converging on one interface for everything.

Competition is also still moving away from the standalone chatbot toward the full stack that surrounds a model. Google folded its Antigravity agent platform into Gemini Enterprise subscriptions with one admin console for governance, billing and usage metrics, OpenAI added an admin plugin for ChatGPT Work and Codex, Amazon Web Services linked OpenSearch observability views into chat for agents, and Anthropic joined Claude chat and Claude Cowork through one persistent memory system with user controls. A persistent memory system means the assistant keeps usable context across sessions instead of starting from zero each time. This likely confirms a statement that has held for six weeks: the contest is increasingly in orchestration, control surfaces, memory, monitoring and workflow fit, because those are the parts that determine whether agents can be run inside ordinary business software rather than beside it.

The infrastructure layer still behaves as if high-end AI compute demand will persist. Alibaba followed strong AI cloud revenue with an HK$80 billion share placement whose proceeds it said will go to full-stack AI capabilities and infrastructure, Nvidia put its Groq 3 LPX inference accelerator into full production, OpenAI published first benchmark results for its Jalapeño inference chip, and Apple added new desktop chips and machines pitched for demanding AI workloads. Bloomberg also reported Nvidia warning customers of about 15% higher prices on Blackwell- and Vera Rubin-based AI servers because of memory costs. That does not refute the line that has held for six weeks; it appears to sharpen it. Demand is no longer showing up only as new model launches, but as financing, production commitments and rising component costs that spread through cloud budgets and local-machine plans alike.

Our forecast

The next visible enterprise split is likely to be between providers that can offer one governed surface across cloud and local AI, and providers that only supply models into someone else’s control layer.

what we hold ourselves toDue 2026-12-15

By 2026-12-15, at least one of Microsoft, Google, Amazon Web Services, OpenAI or Anthropic must announce a generally available enterprise offer that combines centrally administered AI governance and usage controls with an officially supported on-device or local model runtime under the same named product family.

Evidence
  • 2026-08-19 Microsoft expanded the Microsoft Foundry model catalog by adding the DeepSeek‑V4‑Flash‑0731 model as a Direct from Azure option and the NVIDIA Nemotron 3.5 Lightning model via Fireworks and Hugging Face, making them available to developers for agentic and coding workloads under Azure-based governance and billing. techcommunity.microsoft.com
  • 2026-08-21 Microsoft expanded the Azure AI Foundry model router to 28 Azure regions and refreshed its model pool, adding the GPT‑5.6 family and Anthropic’s Claude Opus 4.8 while removing several end‑of‑life GPT‑5 and DeepSeek variants behind a stable endpoint. techcommunity.microsoft.com
  • 2026-08-19 xAI launched its Grok 4.6 model on Amazon Bedrock, making the 500k-context, configurable-reasoning flagship model generally available to developers for long-running agents and interactive and visual workloads with published usage-based pricing in supported AWS regions. x.ai
  • 2026-08-21 xAI made its Grok 4.6 flagship model available via Google Enterprise Agent Platform’s Model Garden with a 500k‑token context window, configurable reasoning levels, and published token‑based pricing for input, cached input, and output. x.ai
  • 2026-08-19 LiquidAI released QAD Q4_0 GGUF checkpoints for four LFM2.5 models, using quantization-aware distillation to retain about 97% of BF16 accuracy at 4-bit precision and demonstrating laptop, mini-PC, smartphone and Raspberry Pi benchmarks where the 4-bit models match higher-precision quality while improving throughput for edge deployment via llama.cpp and other GGUF runtimes. huggingface.co
  • 2026-08-20 LiquidAI introduced DSpark draft model checkpoints for three LFM2.5 models, adding speculative decoding to achieve up to 3.18× higher throughput on GPUs and up to 2.87× faster on-device inference, including significantly lower function-calling latency, with integrations open-sourced for llama.cpp and SGLang. huggingface.co
  • 2026-08-20 Ornith released the Ornith-1.5 family of self-improving mixture-of-experts coding and agentic models under an MIT license, providing 9B, 35B and 397B checkpoints along with GGUF, FP8 and MLX variants and documentation for running them locally via tools such as llama.cpp, vLLM, MLX, LM Studio, Hermes Agent and OpenClaw. huggingface.co
  • 2026-08-25 Apple introduced a new Mac Studio featuring M5 Max and M5 Ultra chips with up to 36 CPU cores, 80 GPU cores with Neural Accelerators and up to 512GB of unified memory, positioning it as a desktop capable of running very large language models entirely on-device and supporting clustered Mac Studios over Thunderbolt 5 for faster distributed AI inference. apple.com
  • 2026-08-25 Apple unveiled a new Mac mini powered by M6 and M5 Pro chips, positioning the desktop as an always-on agentic computing device capable of running advanced AI models on-device alongside macOS 27 and the next generation of Apple Intelligence. apple.com
  • 2026-08-20 Google Cloud announced that its Antigravity AI agent platform is now integrated into eligible Gemini Enterprise subscriptions, adding IDE extensions such as VS Code support and consolidating governance, security, billing and usage metrics for enterprise developers in the Gemini Enterprise admin console. cloud.google.com
  • 2026-08-25 OpenAI introduced an Admin plugin for ChatGPT Work and Codex that allows workspace administrators to manage members, permissions, usage analytics, and credit and spending limits via chat, and to automate workflows such as routing usage requests to Slack or Microsoft Teams. openai.com
  • 2026-08-25 Amazon Web Services announced that Amazon OpenSearch Service now supports Model Context Protocol Apps, enabling AI agents to receive interactive observability visualizations such as trace waterfalls, service maps, and log views directly within chat via a locally run MCP server. aws.amazon.com
  • 2026-08-25 Anthropic rolled out a unified memory system for Claude chat and Claude Cowork so that both share a single persistent memory across conversations, with interfaces for users to view, edit, or delete stored information and default safeguards against retaining sensitive personal data unless explicitly enabled. techcrunch.com
  • 2026-08-20 Alibaba Group reported quarterly revenue of about US$39.6 billion, including a 45% year-over-year increase in AI Cloud and Compute Services revenue to US$7.1 billion and US$1.8 billion from AI-related products with sustained triple-digit growth driven by integration of its T-Head chips and foundation models, broad adoption of the Zhenwu M890 AI processor by over 650 customers, and ecosystem momentum around its Qwen models and QwenWork workforce agent. alibabagroup.com
  • 2026-08-23 Alibaba announced the pricing and planned placement of HK$80 billion worth of newly issued ordinary shares to non-U.S. investors, stating that all net proceeds will be invested in expanding its full-stack AI capabilities and AI infrastructure. alibabagroup.com
  • 2026-08-24 TechCrunch reports that early testers and security experts are raising privacy and security concerns about Instinct’s AI personal assistant, highlighting terms of service that grant broad, perpetual access to users’ emails, messages, audio, location, screen activity, keyboard input and other data for model training, and allow the assistant to enter binding agreements or transactions on users’ behalf. techcrunch.com
  • 2026-08-24 Nvidia announced that its Groq 3 LPX interactive AI inference accelerator, an extension of the Vera Rubin NVL72 platform designed to boost token generation throughput for agentic AI inference, has entered full production, with Nebius named as the first AI cloud provider to adopt Groq 3 LPX systems. nvidianews.nvidia.com
  • 2026-08-24 Bloomberg reported that Nvidia is warning major customers to expect roughly 15% price increases on AI server systems built with Blackwell and Vera Rubin chips due to sharply higher high-bandwidth memory costs, as OEMs relay revised pricing for upcoming deployments. bloomberg.com
  • 2026-08-25 OpenAI published benchmark results for its Jalapeño custom inference chip, reporting 1.5–1.9× more AI work per watt and 1.7–3.6× lower end-to-end latency than comparison hardware on several large models including GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T. openai.com
  • 2026-08-25 Apple introduced its M6 and M5 Ultra chips, presenting M6 as its first 2-nanometer processor and M5 Ultra as a desktop-class SoC built from two M5 Max dies with high unified memory bandwidth to support demanding AI and professional workloads in new Mac mini and Mac Studio systems. apple.com