KI & Microsoft.
Aus unserer Sicht.

Kontrollen, lokale Rechenleistung und Chipausgaben halten den Markt entlang der Vertrauensgrenze gespalten

Microsoft Foundry ist der Ort, an dem Microsoft Modelle unter einer gemeinsamen Abrechnungs- und Governance-Ebene nebeneinander bereithält. Das scheint eine Linie weiter zu verbreitern, die nun seit sechs Wochen hält: Käufer ordnen KI nach Kosten, Vertrauensgrenze und Anwendungsfall, statt für alles auf eine einzige Schnittstelle zuzulaufen.

Auch der Wettbewerb bewegt sich weiter weg vom eigenständigen Chatbot hin zum vollständigen Stack, der ein Modell umgibt. Google integrierte seine Antigravity-Agentenplattform in Gemini-Enterprise-Abonnements mit einer einzigen Administrationskonsole für Governance, Abrechnung und Nutzungsmetriken, OpenAI ergänzte ein Admin-Plugin für ChatGPT Work und Codex, Amazon Web Services verknüpfte OpenSearch-Observability-Ansichten mit dem Chat für Agenten, und Anthropic verband Claude Chat und Claude Cowork über ein einziges persistentes Speichersystem mit Nutzerkontrollen. Ein persistentes Speichersystem bedeutet, dass der Assistent über Sitzungen hinweg nutzbaren Kontext behält, statt jedes Mal bei null zu beginnen. Das dürfte eine Aussage bestätigen, die seit sechs Wochen hält: Der Wettbewerb liegt zunehmend in Orchestrierung, Kontrolloberflächen, Speicher, Monitoring und Workflow-Passung, weil diese Teile bestimmen, ob Agenten innerhalb gewöhnlicher Unternehmenssoftware betrieben werden können statt daneben.

Die Infrastrukturebene verhält sich weiterhin so, als werde die Nachfrage nach hochwertiger KI-Rechenleistung anhalten. Alibaba liess auf starke KI-Cloud-Umsätze eine Aktienplatzierung über 80 Milliarden HK$ folgen, deren Erlös nach eigenen Angaben in Full-Stack-KI-Fähigkeiten und Infrastruktur fliessen soll, Nvidia überführte seinen Groq 3 LPX Inferenzbeschleuniger in die volle Produktion, OpenAI veröffentlichte erste Benchmark-Ergebnisse für seinen Jalapeño-Inferenzchip, und Apple ergänzte neue Desktop-Chips und Maschinen, die für anspruchsvolle KI-Workloads positioniert sind. Bloomberg berichtete zudem, Nvidia habe Kunden wegen Speicherkosten vor um etwa 15% höheren Preisen für KI-Server auf Basis von Blackwell und Vera Rubin gewarnt. Das widerlegt die Linie, die seit sechs Wochen hält, nicht; es scheint sie eher zu schärfen. Nachfrage zeigt sich nicht mehr nur in neuen Modellstarts, sondern ebenso in Finanzierungen, Produktionszusagen und steigenden Komponentenkosten, die sich gleichermassen durch Cloud-Budgets und Pläne für lokale Maschinen ausbreiten.

Unsere Prognose

Die nächste sichtbare Aufspaltung im Unternehmensmarkt dürfte zwischen Anbietern verlaufen, die eine einzige regulierte Oberfläche über Cloud und lokale KI hinweg anbieten können, und Anbietern, die nur Modelle in die Kontrollebene eines anderen liefern.

woran wir uns messen lassenFällig 2026-12-15

Bis 2026-12-15 muss mindestens eines von Microsoft, Google, Amazon Web Services, OpenAI oder Anthropic ein allgemein verfügbares Unternehmensangebot ankündigen, das zentral verwaltete KI-Governance und Nutzungskontrollen mit einer offiziell unterstützten On-Device- oder lokalen Modelllaufzeit unter derselben benannten Produktfamilie kombiniert.

Belege
  • 2026-08-19 Microsoft expanded the Microsoft Foundry model catalog by adding the DeepSeek‑V4‑Flash‑0731 model as a Direct from Azure option and the NVIDIA Nemotron 3.5 Lightning model via Fireworks and Hugging Face, making them available to developers for agentic and coding workloads under Azure-based governance and billing. techcommunity.microsoft.com
  • 2026-08-21 Microsoft expanded the Azure AI Foundry model router to 28 Azure regions and refreshed its model pool, adding the GPT‑5.6 family and Anthropic’s Claude Opus 4.8 while removing several end‑of‑life GPT‑5 and DeepSeek variants behind a stable endpoint. techcommunity.microsoft.com
  • 2026-08-19 xAI launched its Grok 4.6 model on Amazon Bedrock, making the 500k-context, configurable-reasoning flagship model generally available to developers for long-running agents and interactive and visual workloads with published usage-based pricing in supported AWS regions. x.ai
  • 2026-08-21 xAI made its Grok 4.6 flagship model available via Google Enterprise Agent Platform’s Model Garden with a 500k‑token context window, configurable reasoning levels, and published token‑based pricing for input, cached input, and output. x.ai
  • 2026-08-19 LiquidAI released QAD Q4_0 GGUF checkpoints for four LFM2.5 models, using quantization-aware distillation to retain about 97% of BF16 accuracy at 4-bit precision and demonstrating laptop, mini-PC, smartphone and Raspberry Pi benchmarks where the 4-bit models match higher-precision quality while improving throughput for edge deployment via llama.cpp and other GGUF runtimes. huggingface.co
  • 2026-08-20 LiquidAI introduced DSpark draft model checkpoints for three LFM2.5 models, adding speculative decoding to achieve up to 3.18× higher throughput on GPUs and up to 2.87× faster on-device inference, including significantly lower function-calling latency, with integrations open-sourced for llama.cpp and SGLang. huggingface.co
  • 2026-08-20 Ornith released the Ornith-1.5 family of self-improving mixture-of-experts coding and agentic models under an MIT license, providing 9B, 35B and 397B checkpoints along with GGUF, FP8 and MLX variants and documentation for running them locally via tools such as llama.cpp, vLLM, MLX, LM Studio, Hermes Agent and OpenClaw. huggingface.co
  • 2026-08-25 Apple introduced a new Mac Studio featuring M5 Max and M5 Ultra chips with up to 36 CPU cores, 80 GPU cores with Neural Accelerators and up to 512GB of unified memory, positioning it as a desktop capable of running very large language models entirely on-device and supporting clustered Mac Studios over Thunderbolt 5 for faster distributed AI inference. apple.com
  • 2026-08-25 Apple unveiled a new Mac mini powered by M6 and M5 Pro chips, positioning the desktop as an always-on agentic computing device capable of running advanced AI models on-device alongside macOS 27 and the next generation of Apple Intelligence. apple.com
  • 2026-08-20 Google Cloud announced that its Antigravity AI agent platform is now integrated into eligible Gemini Enterprise subscriptions, adding IDE extensions such as VS Code support and consolidating governance, security, billing and usage metrics for enterprise developers in the Gemini Enterprise admin console. cloud.google.com
  • 2026-08-25 OpenAI introduced an Admin plugin for ChatGPT Work and Codex that allows workspace administrators to manage members, permissions, usage analytics, and credit and spending limits via chat, and to automate workflows such as routing usage requests to Slack or Microsoft Teams. openai.com
  • 2026-08-25 Amazon Web Services announced that Amazon OpenSearch Service now supports Model Context Protocol Apps, enabling AI agents to receive interactive observability visualizations such as trace waterfalls, service maps, and log views directly within chat via a locally run MCP server. aws.amazon.com
  • 2026-08-25 Anthropic rolled out a unified memory system for Claude chat and Claude Cowork so that both share a single persistent memory across conversations, with interfaces for users to view, edit, or delete stored information and default safeguards against retaining sensitive personal data unless explicitly enabled. techcrunch.com
  • 2026-08-20 Alibaba Group reported quarterly revenue of about US$39.6 billion, including a 45% year-over-year increase in AI Cloud and Compute Services revenue to US$7.1 billion and US$1.8 billion from AI-related products with sustained triple-digit growth driven by integration of its T-Head chips and foundation models, broad adoption of the Zhenwu M890 AI processor by over 650 customers, and ecosystem momentum around its Qwen models and QwenWork workforce agent. alibabagroup.com
  • 2026-08-23 Alibaba announced the pricing and planned placement of HK$80 billion worth of newly issued ordinary shares to non-U.S. investors, stating that all net proceeds will be invested in expanding its full-stack AI capabilities and AI infrastructure. alibabagroup.com
  • 2026-08-24 TechCrunch reports that early testers and security experts are raising privacy and security concerns about Instinct’s AI personal assistant, highlighting terms of service that grant broad, perpetual access to users’ emails, messages, audio, location, screen activity, keyboard input and other data for model training, and allow the assistant to enter binding agreements or transactions on users’ behalf. techcrunch.com
  • 2026-08-24 Nvidia announced that its Groq 3 LPX interactive AI inference accelerator, an extension of the Vera Rubin NVL72 platform designed to boost token generation throughput for agentic AI inference, has entered full production, with Nebius named as the first AI cloud provider to adopt Groq 3 LPX systems. nvidianews.nvidia.com
  • 2026-08-24 Bloomberg reported that Nvidia is warning major customers to expect roughly 15% price increases on AI server systems built with Blackwell and Vera Rubin chips due to sharply higher high-bandwidth memory costs, as OEMs relay revised pricing for upcoming deployments. bloomberg.com
  • 2026-08-25 OpenAI published benchmark results for its Jalapeño custom inference chip, reporting 1.5–1.9× more AI work per watt and 1.7–3.6× lower end-to-end latency than comparison hardware on several large models including GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T. openai.com
  • 2026-08-25 Apple introduced its M6 and M5 Ultra chips, presenting M6 as its first 2-nanometer processor and M5 Ultra as a desktop-class SoC built from two M5 Max dies with high unified memory bandwidth to support demanding AI and professional workloads in new Mac mini and Mac Studio systems. apple.com