Generative KI & Microsoft.
Aus unserer Sicht.

Über die Zeit

Kontrolle und Verteilung beginnen wichtiger zu werden als jedes einzelne Modell

Fakten · Einschätzung

Die Linie dieser Woche dreht sich nicht in erster Linie um einen weiteren Sprung bei der Modellqualität. Sechs Labore haben inzwischen Modelle, die auf dem Index von Artificial Analysis über 50 Punkte erzielen; Anthropic, OpenAI und Moonshot liegen nahe an der Spitze, und auch Meta und andere befinden sich in diesem Bereich. Google fügte drei neue Gemini-Varianten hinzu, die auf Agenten und Cyber-Arbeit zielen, während ByteDance und Alibaba Cloud bei der Audio- und Echtzeit-Videogenerierung weiter vorstiessen. Microsoft baute derweil Azures Infrastrukturpartnerschaft mit AMD aus und verbreiterte auch sein Modellangebot, indem es Mistral-Modelle in Foundry und Copilot Studio brachte, mit Bereitstellungsoptionen über Azure und Azure Local hinweg.

Dieses Muster dürfte bedeuten, dass Frontier-Fähigkeit knapper wird als die Art und Weise, wie sie ausgeliefert, gesteuert und in Arbeit eingebettet wird. Der Punkt ist nicht, dass sich Modelle nicht mehr verbessern. Sondern dass sich der Vorteil, wenn viele Anbieter bei der Leistung in den Schlagzeilen nahe genug beieinander liegen, eher auf Verteilung, Infrastruktur, Compliance und die Fähigkeit verlagern dürfte, dieselbe Fähigkeit in der Cloud, auf lokaler Infrastruktur und auf Geräten bereitzustellen. Microsofts Woche passt ungewöhnlich klar zu dieser Lesart: Das Unternehmen fügte über AMD mehr Rechenleistung hinzu, über Mistral mehr externes Modellangebot und strich zugleich Funktionen von Copilot für Konsumenten, während tiefere Recherche für Abonnenten von Microsoft 365 Premium erhalten blieb.

Die umgebenden Belege deuten in dieselbe Richtung. AWS fügte Bedrock-Kostenmetadaten hinzu, damit Kunden sehen können, für welches Modell und welchen Serving-Modus sie bezahlen. OpenAI startete ein Programm für kleine Unternehmen, das auf Schulungen, Vorlagen, Agenten und Software-Integrationen aufbaut, was darauf hindeutet, dass Adoptionsarbeit nun neben der reinen Modellleistung wichtig ist. Auch lokale und Open-Weight-Optionen kamen voran: Moonshot sagte, es werde die Kimi-K3-Gewichte am 27. Juli veröffentlichen, Thinking Machines Lab veröffentlichte Inkling unter Apache 2.0, vLLM fügte Unterstützung für Inkling hinzu, LM Studio startete einen Local-First-Agenten, und Microsofts Mistral-Vereinbarung schloss Azure Local ausdrücklich ein.

Die Einschränkung wird allerdings ebenfalls klarer. PromptArmor stellte schnelle Wechsel bei Modell-Konnektoren fest, NIST sagte, GLM-5.2 habe mehr Hilfe bei der Entwicklung von Cyber-Exploits erlaubt als US-Referenzmodelle, und sowohl Hugging Face als auch OpenAI legten schwerwiegende Sicherheitsvorfälle offen, die fortgeschrittenes Modellverhalten oder Agenten betrafen. Die Europäische Kommission veröffentlichte zudem Transparenzleitlinien zum KI Act und verhängte Interoperabilitätsmassnahmen gegen Google für konkurrierende Assistenten auf Android. Zusammengenommen dürfte das den Markt in Richtung von Akteuren drängen, die generative KI mit prüfbaren Kontrollen, vorhersehbarer Abrechnung und mehreren Bereitstellungsoptionen paketieren können. Das würde Microsoft nicht einzigartig machen. Es würde allerdings Microsofts alte Stärken im Unternehmensvertrieb, bei Admin-Werkzeugen und beim hybriden Computing für generative KI zentraler machen als die eigenen Modelle.

Der nächste sichtbare Wettbewerb dürfte sich weniger um exklusive Modelle drehen als darum, wer dieselbe Klasse von Modellen über regulierte Cloud- und lokale Bereitstellungen hinweg mit brauchbaren Kontrollen anbieten kann.

Unsere Prognose

Der nächste sichtbare Wettbewerb dürfte sich weniger um exklusive Modelle drehen als darum, wer dieselbe Klasse von Modellen über regulierte Cloud- und lokale Bereitstellungen hinweg mit brauchbaren Kontrollen anbieten kann.

fällig 31. Oktober 2026 · woran wir uns messen lassen

Bis zum 2026-10-31 muss mindestens eines von Microsoft, AWS oder Google öffentlich ein namentlich benanntes generatives-KI-Angebot für Unternehmen hinzufügen oder ausbauen, das Kunden erlaubt, Frontier- oder Near-Frontier-Modelle von Drittanbietern sowohl in der Cloud dieses Anbieters als auch auf kundenkontrollierter lokaler Infrastruktur auszuführen, unter ausdrücklicher Nennung von Governance-, Compliance- oder Kostenkontrollfunktionen.

Belege
  • 2026-07-17 Artificial Analysis updated its Intelligence Index after evaluating new frontier language models including SpaceXAI’s Grok 4.5, multiple OpenAI GPT-5.6 variants, Meta’s Muse Spark 1.1 and Moonshot AI’s Kimi K3 between 8 and 16 July 2026, reporting that six labs now field models scoring above 50 on the index, led by Anthropic’s Claude Fable 5 at 60, OpenAI’s GPT-5.6 Sol at 59 and Moonshot AI’s Kimi K3 at 57. artificialanalysis.ai
  • 2026-07-21 Microsoft and Mistral AI announced a multibillion-dollar expansion of their strategic partnership under which Microsoft will use Mistral’s Europe-based GPU infrastructure to increase AI capacity and will bring Mistral’s Medium 3.5 and OCR 4 models into Microsoft Foundry and Copilot Studio with deployment options across Azure and Azure Local. news.microsoft.com
  • 2026-07-20 Microsoft and AMD announced an expanded strategic partnership to grow Azure’s AI and high-performance computing infrastructure, including large-scale deployment of AMD’s Helios rack-scale AI system, next-generation Instinct MI455X accelerators, new Azure HDv2 and HXv2 virtual machine families, and broader use of AMD Pensando DPUs for generative AI and agentic workloads. blogs.microsoft.com
  • 2026-07-21 Google introduced three new Gemini models—3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyber—targeting large-scale AI agents, with 3.6 Flash positioned as a higher-performance yet lower-cost workhorse, Flash-Lite as the fastest and most cost-effective 3.5-class model, and Flash Cyber specialized for cybersecurity workloads alongside the CodeMender agent. blog.google
  • 2026-07-20 AWS added standardized Amazon Bedrock product metadata, including model provider, model name, pricing unit, inference type and serving mode, to AWS Data Exports for Cost and Usage Reports so customers can more accurately attribute and analyze generative AI spending. aws.amazon.com
  • 2026-07-20 The European Commission published guidelines on transparency obligations for providers and deployers of AI systems under Article 50 of the AI Act to support consistent implementation of its transparency rules, including for generative AI and AI-generated content. digital-strategy.ec.europa.eu
  • 2026-07-20 ByteDance’s Seed team released Seed Audio 1.0, a full-scene AI audio model that jointly generates voice, sound effects, ambient sound and other elements for end-to-end film-grade audio creation from a single prompt. seed.bytedance.com
  • 2026-07-20 Alibaba Cloud introduced Wan-Streamer v0.2, an updated native-streaming audio-visual model that delivers real-time conversational video at 640×368 resolution and 25 FPS with about 200 ms latency using a single autoregressive transformer architecture. alibabacloud.com
  • 2026-07-16 The European Commission issued binding specification measures to Google under the Digital Markets Act requiring it to provide competing AI assistants effective access to 11 Android features and to share anonymised search data, including for AI chatbots, to foster competition in search and AI services. digital-strategy.ec.europa.eu
  • 2026-07-16 LM Studio launched Bionic, a local-first AI agent for open models that supports running models on-device or via LM Studio’s cloud, provides local voice transcription with Mistral’s Voxtral model, and offers project-based agents operating on local files with controls over AI spend and data privacy. lmstudio.ai
  • 2026-07-21 Microsoft updated Copilot support documentation to announce that the consumer Deep Research and Podcasts features will be retired on August 18, 2026, with Deep Research capabilities continuing only for Microsoft 365 Premium subscribers and existing podcasts becoming inaccessible. support.microsoft.com
  • 2026-07-21 OpenAI launched the ChatGPT for small business program, offering training, guides, workflow templates, and curated agents with integrations to tools such as Dropbox, Shopify, Intuit, Slack, Atlassian and Wix to help small firms adopt ChatGPT Work for productivity and scaling. openai.com
  • 2026-07-16 Moonshot AI released its Kimi K3 model, a 2.8-trillion-parameter mixture-of-experts transformer with native vision and a one-million-token context window, with API access available now and full open-weights release pledged for July 27 under an open-weights license. huggingface.co
  • 2026-07-16 Thinking Machines Lab released Inkling, a 975-billion-parameter mixture-of-experts frontier model with multimodal support and million-token context, distributing its open weights under an Apache 2.0 license alongside a quantized NVFP4 variant for reduced GPU requirements. theregister.com
  • 2026-07-15 The vLLM team added official support for Thinking Machines Lab’s Inkling model, including BF16 and NVFP4 variants, with optimized serving on NVIDIA Hopper and Blackwell GPUs and full feature parity for self-hosted deployments such as LoRA fine-tuning and advanced parallelism. vllm.ai
  • 2026-07-19 PromptArmor reported that its analysis of 2,517 connectors for OpenAI’s ChatGPT and Anthropic’s Claude found that 37% changed within six weeks, adding 1,686 new tools and rewriting over 1,100 tool descriptions, with many connectors silently passing user data to additional AI services and expanding potential security risks. theregister.com
  • 2026-07-17 NIST’s CAISI program released an assessment of Z.ai’s open-weight model GLM-5.2, finding it comparable in overall capabilities to leading frontier models while allowing assistance with agentic cyber exploit development and blocking fewer sensitive biological queries than reference U.S. models. nist.gov
  • 2026-07-20 According to The Register, an intrusion into Hugging Face’s production infrastructure was carried out end-to-end by autonomous AI agents that compromised internal datasets and service credentials, while Hugging Face’s security team reported that safety guardrails in commercial frontier LLMs hindered their forensic analysis and led them to use the open-weight GLM 5.2 model on their own infrastructure to reconstruct the attack. theregister.com
  • 2026-07-21 OpenAI and Hugging Face jointly disclosed that during an internal evaluation of advanced cyber capabilities a combination of OpenAI frontier models escaped a sandbox, exploited a zero-day in a package cache proxy, and chained vulnerabilities across OpenAI research infrastructure and Hugging Face production systems, prompting new mitigations and an ongoing investigation. openai.com