Big 5 AI Vendor Roundup: Week of August 10, 2026

Technology Note By: Mark Tauschek, Bill Wong, Info-Tech Research Group

Cyber AI became a commercial product this week. OpenAI released GPT-5.6-Cyber through a restricted program, Microsoft put its first cyber model into a multi-agent security system, and AWS made OpenAI’s controlled cyber tiers available through Amazon Bedrock. OpenAI also gave ChatGPT access to a broader record of users’ work, while Anthropic made a more autonomous operating mode the default in Claude Code.

For IT leaders, the model is no longer the right unit of evaluation. Risk now depends on the identity, data, memory, tools, network access, monitoring, and approval rules around it. Cost is also harder to compare as vendors add premium seats, faster inference, temporary token prices, automated model routing, and advertising-supported access.

OpenAI expands controlled cyber access and ChatGPT context

  • GPT-5.6-Cyber launches behind restricted access. OpenAI introduced two Daybreak tiers on August 10. Daybreak Blue provides access to GPT-5.6 Sol with safeguards optimized for cyber defense, while Daybreak Red gives approved users access to GPT-5.6-Cyber for advanced vulnerability research and testing. OpenAI says the cyber model completed 95% of advanced requests in its internal evaluation, compared with 1.5% for standard Sol. These are company results, and the full system card hasn’t been published. Access requires identity checks, monitoring, approved uses, and isolated environments.
  • OpenAI expands the Daybreak delivery network. Consulting, security, and cloud partners can now operate restricted cyber models for customers without transferring direct access. That widens availability but shifts part of the control burden to the service provider. Buyers should confirm who approves targets, controls credentials, reviews activity, and can suspend access.
  • Computer History gives ChatGPT a record of work across a Mac. OpenAI’s release notes say the optional macOS feature lets ChatGPT and Codex reference activity from selected apps and websites. It records interaction events rather than screenshots, video, or audio, and excludes private browsing. The feature is off by default for Pro, Business, and Enterprise users, with admin approval required in managed workspaces. The productivity benefit is clear, but so is the governance issue: A much broader record of employee activity is now available to the assistant.
  • ChatGPT adds deeper Google Drive access and a Linux desktop preview. Users can browse connected Google Drive files and folders from ChatGPT’s Library, open Docs, Sheets, and Slides beside a conversation, and update source files where supported. OpenAI also added controls for changing project memory settings and released a public-preview Linux app. Shared Drives and some collaboration functions aren’t yet supported. Enterprises should test source permissions and editing behavior before enabling broad access.
  • A limited Sol tier promises up to 14-times-faster output. GPT-5.6 Sol Ultrafast entered an API-first preview on Cerebras infrastructure, with OpenAI claiming up to 750 output tokens per second. Pricing hasn’t been disclosed. Faster inference could help interactive agents, but its value will depend on the premium and production reliability.
  • ChatGPT Business adds a higher-priced seat. Business Premium costs $125 monthly, or $100 with annual billing, and provides five times the usage allowance without the five-hour limit. Organizations can mix standard and premium seats. That improves role-based purchasing but makes forecasting more dependent on individual usage.
  • ChatGPT advertising expands to five more countries. Ads are now available to Free and Go users in the United Kingdom, Mexico, Brazil, Japan, and South Korea. Paid plans remain ad-free. OpenAI says ads won’t affect answers or expose private chats to advertisers. Enterprises should still separate managed and consumer accounts and enforce data-loss controls.
  • OpenAI says its heaviest enterprise users are moving toward agents. OpenAI reports that users at the top 10% of customer organizations produce 8.3 times more output tokens than typical enterprise users, while coding agents generated 64% of combined ChatGPT and Codex output tokens in June. The numbers show usage intensity, not business value. Buyers still need measures for cycle time, quality, rework, and avoided cost.

Microsoft broadens its model portfolio and security platform

  • Microsoft releases its own reasoning model. MAI-Thinking-1 entered public preview in Microsoft Foundry on August 12. Microsoft says the medium-sized model was trained without distilling a third-party model and uses traceable development data. The launch gives Microsoft more control over cost and supply while Foundry continues to support outside models. Buyers should test whether model choice remains portable once workflows depend on Microsoft context, memory, and orchestration.
  • MAI-Cyber-1-Flash enters Microsoft’s security system. Microsoft integrated its first cyber model into MDASH, a multi-agent vulnerability system. It says the specialist model handles about 90% of tasks and routes harder work to a larger model, cutting model cost by roughly half. These are Microsoft benchmarks. The broader point is the routing pattern: Cheaper specialist models handle routine work while premium models cover exceptions.
  • Defender folds threat intelligence into XDR and Sentinel. Microsoft’s August 5 Defender update highlights the completion of Microsoft Defender Threat Intelligence integration into the Defender portal. Threat actor profiles, indicators, and investigation context now appear inside Defender XDR and Sentinel workflows, while standalone MDTI experiences are being retired. This should reduce tool switching, but security teams need to review licensing, access roles, APIs, and any workflows tied to the old interface.
  • A lower-cost coding model moves into GitHub Copilot. MAI-Code-1.1-Flash is now in production in GitHub Copilot. Microsoft says it uses 25% fewer tokens and costs one quarter as much as the prior version. That model retires on September 10. Enterprises need regression tests so vendor substitutions don’t quietly change code quality, security, or language support.
  • GitHub makes agent extensions more portable and model costs easier to see. Agent Plugins 1.0 packages skills and Model Context Protocol servers for use across compatible clients. GitHub also added managed plug-in marketplaces, allowlists, and per-model usage reporting. Portability is improving, but organizations still need controls over installed tools, reachable data, and model spending.

Anthropic makes Claude Code more autonomous while raising its risk estimate

  • Claude Code makes auto mode the default. Starting August 14, new sessions on Pro, Max, and Team plans run in auto mode unless users have pinned another setting. Instead of asking for approval before each action, Claude Code uses a classifier to block destructive, irreversible, or external actions, then tries a safer path or returns to manual approval after repeated denials. Auto mode remains opt-in for now on Enterprise, API, and cloud deployments. Anthropic reports that auto mode caught 89% of dangerous test commands versus 13.6% for human reviewers. That is a vendor-run result, but it illustrates the trade-off: fewer approval prompts in exchange for greater reliance on automated policy enforcement.
  • Anthropic raises its high-stakes misalignment assessment. Anthropic’s August Risk Report moves its assessment from “very low” to “low,” citing uncertainty after recent model behavior in cybersecurity evaluations. It reports cases where models were willing to take misaligned actions to finish difficult tasks, although it still rates catastrophic risk from known forms of this behavior as low. The company also disclosed that contractor traffic lacked biological blocking safeguards from May 2025 through April 2026. Anthropic says it fixed the gap and found no concerning misuse.
  • Future Claude models will watermark generated text. Anthropic plans to add a version of Google DeepMind’s SynthID-Text watermark to future models and then existing ones. It also plans a detection API and C2PA provenance metadata for files. Detection will be weaker for short, factual, coded, or heavily edited content, so watermark results should support an investigation rather than prove authorship.

Google lowers model pricing and expands Gemini connections

  • Gemini 3.7 Flash launches with introductory pricing. Google released Gemini 3.7 Flash on August 13 at $0.75 per million input tokens and $3.75 per million output tokens through the end of 2026. Google says the model improves coding, business, and document tasks and now powers Gemini Spark in more than 160 countries. Buyers should model the permanent price before committing high-volume workloads.
  • Gemini passes one billion monthly users and adds more connected services. Google says the Gemini app now has more than one billion monthly users. It also added connections to services including Otter, Granola, Wix, Ticketmaster, OpenTable, and Zocdoc. The consumer milestone doesn’t prove enterprise value. The immediate enterprise issue is the growing permission surface across calendars, notes, reservations, health services, and work apps.

AWS brings controlled cyber models and agents into regulated work

  • OpenAI’s Daybreak models become available through Amazon Bedrock. Eligible AWS customers can access Daybreak Blue and the more restricted Daybreak Red through Bedrock. This supports existing AWS procurement, identity, governance, and monitoring, but customers still need separate approval rules for targets, credentials, internet access, and human review.
  • Novo Nordisk selects AWS for AI and cloud expansion. Novo Nordisk named AWS its preferred cloud provider and strategic AI partner. The companies plan to use Amazon Bio Discovery and Bedrock AgentCore across research, commercial operations, and enterprise IT. AWS says more than 25,000 employees already use a Bedrock-based assistant for nonregulated work. Regulated scientific decisions will still require validation, documentation, and human accountability.

Outside the Big 5

Being Reported

These items are being reported but haven’t been fully announced by the companies involved.

Our Take

Advanced cyber models are moving from controlled research into cloud and partner-delivered services. OpenAI’s Daybreak program, Microsoft’s specialist cyber model, and AWS distribution all point in the same direction. The practical safety boundary depends less on the model name than on identities, target limits, connected tools, network access, logs, and human escalation. Anthropic’s safeguard disclosure shows how easily a control can remain missing in an overlooked workflow.

The same shift is happening in everyday products. Computer History, connected Drive files, Claude Code auto mode, and portable agent plug-ins all reduce friction by giving assistants more context and freedom to act. They also concentrate more risk in memory, permissions, classifiers, and policy systems. Buyers may gain model choice while becoming more dependent on the vendor’s control plane.

What IT leaders should be doing

  • Treat advanced cyber models as privileged security tools. Use named identities, strong authentication, isolated environments, narrow target lists, minimal credentials, no open internet by default, and human approval before external actions.
  • Govern context collection as carefully as data access. Decide which apps, websites, files, and activity histories an assistant may use. Start with narrow permissions, review retained context, and provide clear pause and deletion controls.
  • Test automated approval systems. Evaluate what auto modes and classifiers allow, block, and log. Keep manual escalation for destructive, external, privileged, and production actions.
  • Model the full cost before scaling. Include seat level, tokens, speed tiers, tool calls, retries, routing, and runtime infrastructure. Require reporting by model, team, and workflow.
  • Keep the control plane portable. Separate models from memory, connectors, policy, evaluations, and logs where possible. Test a fallback model and verify that plug-ins and workflows can move without a full rebuild.

Want to Know More?

Latest Technology Notes

All Technology Notes