Generative KI & Microsoft.
Aus unserer Sicht.

Über die Zeit

Microsoft wird dazu gedrängt, den Stack selbst zu besitzen, während Vertrauen und Kosten die Bedingungen der KI-Nutzung verschärfen

Diese Woche deutet darauf hin, dass Microsoft versucht, weniger von einem einzelnen Modellanbieter abhängig zu werden und zugleich nützlicher als der Ort zu werden, an dem Unternehmens-KI gesteuert, bereitgestellt und bezahlt wird. Microsoft führte MAI-Bild-, Sprach- und Cyber-Modelle in Foundry ein und begann, Workloads von OpenAI-Modellen auf seine eigenen MAI-Modelle zu verlagern, und zwar in Produkten einschliesslich Bing Image Creator, PowerPoint-Bildfunktionen, OneDrive-Bildbearbeitung, Dynamics 365 Contact Center und Azure Voice Live. Ausserdem fügte das Unternehmen Anthropics Claude Opus 5 sowohl Microsoft 365 Copilot als auch Microsoft Foundry hinzu, während Manulife eine ausgeweitete fünfjährige Partnerschaft unterzeichnete, die Microsoft 365 E7 Frontier Suite, Agent 365 Governance und Copilot für mehr als 30'000 Mitarbeitende umfasst. Dies scheint die Linie der vergangenen Woche zu verlängern, wonach sich der Wettbewerb von einem einzelnen Chatbot oder Basismodell hin zum vollständigen Stack aus Modellauswahl, Governance, Speicher, Sicherheit und Workflow-Bereitstellung verschiebt. Es deutet auch darauf hin, dass für Microsoft die Kontrolle über die Kundenumgebung wichtiger sein dürfte als die ausschliessliche Kontrolle über das zugrunde liegende Modell.

Dieselbe Woche zeigte auch, warum dieser Stack eher um Beschränkungen herum als allein um Neuheit herum aufgebaut wird. SAP entschied, die Ausgaben für KI-Token zu begrenzen, nachdem die interne Nutzung die Kosten erhöht hatte, und Alibaba Cloud beschrieb eine Agentenoptimierung, die darauf abzielt, den kontextbezogenen Token-Verbrauch zu senken. OpenAI führte Presence für gesteuerte Unternehmensagenten ein, während Microsoft einen auf Cosmos DB basierenden Speicheranbieter für sein Agent Framework veröffentlichte und AWS AgentCore Gateway auf die neueste MCP-Spezifikation mit Autorisierung und Lebenszykluskontrollen für Unternehmen aktualisierte. Das dürfte bedeuten, dass sich die nächste Phase des Einkaufs von Unternehmens-KI weniger darum drehen wird, ob ein Unternehmen überhaupt Agenten will, und stärker darum, ob diese Agenten innerhalb von Budgets, Berechtigungen und Prüfregeln betrieben werden können.

Gleichzeitig steigt auch der Vertrauensdruck. uniVersa berichtete, dass ein mit OpenAI verknüpfter Crawler während einer IT-Migration auf persönliche Kundendaten zugriff, 404 Media berichtete, dass öffentliche Claude-Share-Links von Google indexiert wurden, Codeberg änderte seine Bedingungen, um KI-Trainingsnutzungen und stark automatisch erzeugte Projekte zu blockieren, und der EU AI Omnibus trat in Kraft. NIST startete ausserdem ein abgeschottetes Evaluationsprogramm, und die Reserve Bank of India veröffentlichte einen Leitlinienentwurf zum Modellrisiko für regulierte Finanzunternehmen. Wir lesen das als Beleg dafür, dass KI tiefer in normale institutionelle Kontrollsysteme eindringt: Datenschutzvorfälle, Herkunftssignale, Blindtests und formales Risikomanagement sind keine Nebenthemen mehr, sondern Teil der Bedingungen, unter denen Bereitstellung erlaubt wird.

Unter all dem weist der Ausbau der Infrastruktur weiterhin in dieselbe Richtung. Alphabet meldete ein schnelleres Wachstum von Google Cloud, das mit KI-Infrastruktur und der Nachfrage nach Unternehmens-KI verknüpft war, AMD führte sein Rack-Scale-System Helios ein und kündigte eine Partnerschaft mit Anthropic für künftige GPU-Bereitstellung an, Intel erhöhte die Investitionsausgaben, nachdem die Nachfrage nach KI-Rechenzentrumschips den Umsatz angehoben hatte, und von NVIDIA unterstützte Projekte in Korea umrissen einen weiteren grossen Ausbau von KI-Fabriken. Das dürfte die Linie bekräftigen, dass sich die KI-Nachfrage nicht wie ein kurzer Schub rund um eine Modellgeneration verhält. Die unmittelbarere Folge für Microsoft könnte jedoch einfacher sein: Wenn Rechenleistung teuer bleibt und sich ein reichliches Modellangebot weiter über Clouds ausbreitet, dürfte der dauerhafte Vorteil bei demjenigen liegen, der Modelle in gesteuerte, bepreiste und vertrauenswürdige Betriebssysteme für die Arbeit verpacken kann.

Unsere Prognose

Microsoft dürfte die Modellauswahl innerhalb seiner eigenen Unternehmensoberflächen weiter ausweiten und zugleich zuerst eigene Modelle dort einsetzen, wo sie die Bereitstellungskosten in von Microsoft betriebenen Workloads senken.

fällig 30. November 2026 · woran wir uns messen lassen

Bis zum 2026-11-30 muss Microsoft sowohl (1) mindestens ein weiteres Nicht-OpenAI-Drittmodell nach Claude Opus 5 entweder Microsoft 365 Copilot oder Microsoft Foundry hinzufügen als auch (2) mindestens eine weitere Migration eines namentlich genannten Microsoft-Produkts oder Workloads von einem externen Modell zu einem namentlich genannten MAI-Modell ankündigen.

Belege
  • 2026-07-23 Microsoft launched its in-house generative AI models MAI-Image-2.5-Pro and MAI-Voice-2-Flash on Microsoft Foundry and began migrating workloads such as Bing Image Creator, PowerPoint image features, OneDrive image editing, Dynamics 365 Contact Center and Azure Voice Live from OpenAI models to these MAI models, reporting GPU cost reductions of up to 84% versus GPT-Image-2. techcommunity.microsoft.com
  • 2026-07-24 Microsoft added Anthropic’s Claude Opus 5 as a selectable model in Microsoft 365 Copilot across core apps, with rollout via the Copilot model selector depending on region and tenant configuration. techcommunity.microsoft.com
  • 2026-07-24 Microsoft made Anthropic’s Claude Opus 5 model available in Microsoft Foundry for building long‑running enterprise workflows and agentic applications, integrated with Foundry’s evaluation, security, governance and deployment tools. techcommunity.microsoft.com
  • 2026-07-22 Manulife signed a renewed and expanded five‑year partnership with Microsoft under which it will adopt the Microsoft 365 E7 Frontier Suite, deploy Microsoft Agent 365 for AI agent governance, and expand Microsoft 365 Copilot to more than 30,000 employees using Azure and Microsoft Foundry. news.microsoft.com
  • 2026-07-24 Microsoft released a CosmosMemoryContextProvider that uses Azure Cosmos DB as a native memory store for the Microsoft Agent Framework, providing durable cross‑session agent memory with automatic summarization and profiling, currently in Python preview. devblogs.microsoft.com
  • 2026-07-27 At a security event in San Francisco, Microsoft launched its MAI-Cyber-1-Flash cybersecurity AI model and the Project Perception agentic security platform, integrating the model into its MDASH multi-agent harness and claiming around 96% performance on the CyberGym benchmark at about half the cost of prior configurations. techcrunch.com
  • 2026-07-24 SAP decided to significantly restrict its spending on AI tokens after company‑wide AI usage drove token costs sharply higher, introducing a strict three‑stage controlling system to manage and limit AI token consumption. golem.de
  • 2026-07-28 Alibaba Cloud described an optimization to Qoder’s Quest 1.0 autonomous coding agent that reduces unnecessary MCP tool definitions in prompts, cutting context‑related token usage by over 10% in internal experiments while aiming to preserve agent capabilities. alibabacloud.com
  • 2026-07-22 OpenAI launched Presence, a managed enterprise platform for building, deploying, and operating governed AI agents for high‑volume, high‑stakes workflows, currently available only through limited managed deployments. help.openai.com
  • 2026-07-28 AWS announced that Amazon Bedrock AgentCore’s AgentCore Gateway now supports the Model Context Protocol (MCP) 2026‑07‑28 specification, introducing a stateless design, governed extensions, stronger authorization aligned with enterprise OAuth 2.0/OpenID Connect, and lifecycle guarantees configurable via UpdateGateway. aws.amazon.com
  • 2026-07-23 uniVersa insurance group reported a data protection incident in which an AI web crawler linked to OpenAI accessed a server during an IT migration and reached personal customer data including names, addresses, insurance numbers, tariff information and some bank details, after which the server was shut down, forensic experts were engaged and the case was reported to the Bavarian data protection authority. heise.de
  • 2026-07-27 404 Media reported that public share links from Anthropic’s Claude chatbot and its Artifacts feature were being indexed by Google, making users’ shared conversations and creations—including some sensitive health and company information—searchable on the open web. 404media.co
  • 2026-07-23 Codeberg e.V. updated its terms of use to prohibit using its services and hosted code as training data for generative AI or large language models and to allow exclusion of repositories that are largely autogenerated by AI agents or consume excessive resources, citing heavy load from AI crawlers and the energy and hardware demands of LLMs. heise.de
  • 2026-07-27 The EU’s AI Omnibus regulation entered into force on 27 July 2026, extending timelines and simplifying compliance for high-risk AI systems and SMEs under the AI Act, strengthening the AI Office’s enforcement powers, banning nudification apps that generate non-consensual sexual or intimate content and child sexual abuse material, and permitting certain sensitive-data processing to detect and correct bias in AI systems. digital-strategy.ec.europa.eu
  • 2026-07-27 NIST launched the Artificial Intelligence Technology Evaluation (AITE) program to provide a sequestered testbed for assessing AI model performance on blind data, initially focusing on image-analysis tasks using large vision-language models in domains such as quantum science, genomics and public safety. nist.gov
  • 2026-07-22 The Reserve Bank of India released for consultation a draft "Guidance on Regulatory Principles for Model Risk Management, 2026" setting governance and risk management requirements for traditional and AI/ML models used by regulated financial entities, including AI‑specific provisions on explainability, bias testing, hallucination controls, drift monitoring, human oversight, cybersecurity, customer disclosure and accountability for third‑party models. bloomberg.com
  • 2026-07-24 Google Cloud released Open Knowledge Format v0.2, adding optional frontmatter fields that provide provenance, trust, freshness, lifecycle and attestation signals so AI agents can assess and distinguish verified from unverified knowledge artifacts before consuming them. cloud.google.com
  • 2026-07-22 Alphabet reported Q2 2026 results showing 24% year‑over‑year revenue growth to $119.8 billion and 82% growth in Google Cloud revenue to $24.8 billion, attributing the cloud acceleration to demand for AI infrastructure and enterprise AI solutions. sec.gov
  • 2026-07-23 AMD launched its Helios AI rack-scale system and Instinct MI400 Series GPUs, alongside 6th Gen EPYC CPUs and Pensando networking, as a full-stack AI compute portfolio at its Advancing AI 2026 event. ir.amd.com
  • 2026-07-22 AMD and Anthropic announced a strategic partnership under which Anthropic plans to deploy up to 2 gigawatts of AMD Instinct MI450 Series GPUs in AMD Helios rack-scale solutions starting in 2027, while the companies collaborate on optimizing Claude workloads for AMD GPUs and AMD broadly adopts Claude internally. newsroom.amd.com
  • 2026-07-24 Intel reported second-quarter 2026 revenue of $16.1 billion, up 25% year on year and its fastest growth in about 15 years, driven largely by demand for chips used in AI data centres, and increased its capital spending plans by $2 billion. cn.ft.com
  • 2026-07-25 NAVER, NVIDIA and Brookfield announced plans to expand NAVER’s NVIDIA DSX-based AI factory at the GAK Sejong data centre in Korea from 55 megawatts to 200 megawatts by 2028 as part of gigawatt-scale multi-tenant AI cloud infrastructure investments, with NVIDIA to invest about $1 billion in NAVER and Brookfield to arrange up to $9 billion in project financing. investor.nvidia.com
  • 2026-07-25 SK Group and NVIDIA announced an expanded strategic partnership worth more than $500 billion that includes SK Telecom’s plan to build a 2-gigawatt NVIDIA Vera Rubin DSX AI factory in Korea and a long-term agreement for SK hynix and NVIDIA to co-develop and secure supply of future AI memory, including HBM. investor.nvidia.com