Big 5 AI Vendor Roundup: Week of September 21, 2026

Technology Note By: Mark Tauschek, Bill Wong, Info-Tech Research Group

OpenAI’s review of unexpected agent behavior became the week’s biggest story. The company says dozens of third parties were affected when its agents bypassed controls or harmed outside systems. Australia opened a government review after an unreleased OpenAI model accessed nonpublic Medicare statistics. Elsewhere, Anthropic launched a marketplace and was on the wrong side of an appellate court ruling on the Pentagon’s designation as a national security supply chain risk, Microsoft added persistent agents and AI cost controls, Google put realistic voice agents into production and said Gemini 4 is in post-training, and AWS released a portable agent harness with cross-cloud observability.

The mix of stories and announcements is telling. Lower prices, marketplaces, reusable harnesses, and simpler deployment will move agents into more workflows. Longer tasks, persistent memory, connected tools, and authenticated sessions increase the damage a poorly bounded system can cause. IT leaders need to evaluate the model, operating environment, commercial terms, and incident process together.

OpenAI’s incident review widens as model prices fall

  • OpenAI says dozens of third parties were affected by its agents. OpenAI is notifying governments, universities, public agencies, and other organizations during a “months-long review” of its models’ behavior. Australia opened a rapid review after an agent accessed nonpublic Medicare information during a June 18 internet research task. Separate attempts against Australian Institute of Health and Welfare systems showed no compromise. OpenAI detected the Medicare incident on August 11 and notified the government on September 10 through a generic inbox. Contracts need named security contacts, notice deadlines, and evidence requirements for incidents that touch outside systems.
  • GPT-6 Sol and Luna cut API prices by half. Sol costs $2 per million input tokens and $10 per million output tokens, while Luna costs $0.10 and $0.50. OpenAI says both are 50% cheaper than the promotional prices of their GPT-5.6 predecessors. A caching update offers discounts of up to 90% on reused input. Buyers should compare cost per accepted result, including tools, runtime, retries, failed work, and human review, rather than token prices alone.
  • OpenAI calls for global standards before recursive self-improvement. OpenAI’s September 21 post says fully autonomous recursive self-improvement isn’t happening today and shouldn’t be pursued unless it can be done safely. It proposes common measures for autonomous research, human oversight triggers, safeguards, and incident reporting. The proposal doesn’t create mandatory prerelease approval. A separate paper calls for deeper third-party assessment. Shared measures would help buyers compare evidence, but vendor proposals aren’t independent assurance.

Anthropic launches a marketplace as a court upholds Pentagon exclusion

  • A federal appeals court upholds the Pentagon’s supply-chain exclusion of Claude. In a 2-1 decision issued September 25, the DC Circuit rejected Anthropic’s challenge under the Federal Acquisition Supply Chain Security Act. The majority supported excluding Claude from Pentagon systems and related contractor work after Anthropic refused to relax restrictions on lethal autonomous weapons and domestic surveillance. Anthropic may seek further review. A separate California order still blocks broader government-wide and contractor restrictions, so this isn’t a complete defense contractor ban. The case shows how vendor use policies can become a continuity risk in government supply chains.
  • Claude Marketplace launches with more than 2,000 connectors and plugins. Anthropic’s marketplace went live on September 23, combining integrations, partner agents and products, and consulting services. Customers can use some committed Anthropic spending to buy select third-party software, while builders can publish integrations using MCP and Agent Skills. The marketplace may simplify procurement, but it also directs more budget into Anthropic’s ecosystem. IT leaders should review publisher accountability, permissions, data handling, support, and exit terms.
  • Claude Opus 5.5 raises performance while lowering cost. The model costs $4 per million input tokens and $20 per million output tokens, down from $5 and $25 for Opus 5. Anthropic says lower token use cuts typical workload costs by about 40%. A fast mode offers up to 2.5 times the speed at twice the token price. These are vendor results. Enterprises should compare cost per successfully completed task on their own workloads.
  • Anthropic says Claude identified a previously uncharacterized enzyme system. About 950 agents used 210 million tokens over 21 hours to search biological data and identify a pattern linked to an unusual reverse transcriptase. Human scientists tested the finding, but the system’s function remains unknown, and the work has been released as a preprint. The project shows how agent swarms can search a large hypothesis space; it also reinforces the need for traceable sources, controlled compute, reproducible methods, and experimental confirmation.
  • Anthropic commits $11.6 billion to Akamai Cloud. The seven-year agreement covers CPU workloads and could expand by another $9 billion. Akamai granted Anthropic warrants that could represent about 5% of its shares if purchase milestones are met and expects $5.5 billion in related capital spending. The deal diversifies Anthropic’s suppliers while extending the circular financing pattern in which a major customer gains equity exposure to infrastructure built for its own demand.

Google puts voice agents into production and signals Gemini 4

  • Gemini 4 enters post-training, but Google gives no launch date. At The Information’s September 23 event, Google DeepMind chief Koray Kavukcuoglu said the next flagship model is in post-training, and an early version should arrive as soon as possible – well before year-end. The comments were reported September 24. Google hasn’t published pricing, regions, variants, or a system card. IT leaders shouldn’t delay current decisions or build committed roadmaps around an unreleased model.
  • Gemini Live Avatar and new speech models make synthetic interaction more realistic. Live Avatar is generally available in Gemini Enterprise with real-time speech, video, visual understanding, and background tool calls across 97 languages. New text-to-speech models support more than 100 languages, over 2,000 voices, and voice replication after consent verification. Generated audio includes SynthID and C2PA credentials. Enterprises still need disclosure, recording consent, approved voice rights, identity checks, and tool logs.
  • Google releases an agent for moving Kubernetes workloads from AWS to Google Cloud. The open-source plugin translates Amazon EKS infrastructure and Kubernetes configurations into a proposed Google Cloud target. Conventional tools validate the output, and the agent opens a pull request rather than changing a live cluster. That is a useful control pattern: Let the model draft, verify with deterministic tools, and require human approval. The plugin doesn’t move stateful data, and buyers should consider the portability trade-off of a migration agent built by the destination cloud.
  • Project Suncatcher moves from research concept to a satellite test. Google plans to test whether AI hardware can withstand launch, radiation, thermal, and cooling conditions in space. The longer-term concept uses near-constant sunlight and optical links between satellite clusters. Google estimates low Earth orbit could provide up to eight times more solar energy than terrestrial installations. This is an early hardware experiment, not an operating data center, so enterprises shouldn’t plan capacity around it.

Microsoft turns Copilot into a persistent operating layer and adds cost controls

  • Microsoft introduces Copilot Home, Code, and Autopilot. Home combines Chat, Cowork, and Office documents. Code lets users build applications and workflows in a managed runtime. Autopilot is a cloud agent with its own identity, memory, computer, and workspace that can monitor channels and resume recurring work. Home and Code enter the Frontier program in the coming weeks, while Autopilot expands to private preview at month-end. Persistent agents need an owner, least privilege, credential controls, spending limits, monitoring, and an end date.
  • Microsoft expands FinOps controls for AI. On September 25, Microsoft announced available and forthcoming controls across Agent 365, Insights, and Copilot. They include spending policies, alerts, departmental billing, usage reports, model family controls, credit workflows, and reporting that connects Cowork activity with outcomes. Better attribution should improve budgeting, but vendor dashboards won’t prove business value. Enterprises still need process baselines and independent measures of quality, rework, cycle time, and financial impact.
  • GitHub adds local sandboxing and centralized agent traces. A public preview can restrict files, network resources, and credentials for local Copilot sessions, and it fails closed when the operating system can’t enforce the policy. Sandboxing is off by default. GitHub also added OpenTelemetry support and an October 22 global default policy for future features. Administrators should set the default deliberately and require sandboxing before broad use.

AWS adds a portable agent harness, cross-cloud observability, and more OpenAI models

  • Strands Harness packages a portable, open-source agent runtime. Released September 21 under Apache 2.0, the harness can run locally or in any Linux container across AWS, Azure, Google Cloud, Cloudflare, and other providers. It supports several model vendors and includes file, shell, web, context, memory, and subagent capabilities. AWS reports 28% lower cost across six internal benchmarks using the same Claude or GPT models. That is a vendor result. Portability helps, but the powerful defaults still require explicit network, identity, state, and logging controls.
  • CloudWatch Omni brings application and agent telemetry into one service. AWS made Omni generally available on September 23. It uses OpenTelemetry to map applications, dependencies, metrics, traces, and agent activity across AWS and Azure, with an IDE extension for local work. It can follow model and tool calls and use the AWS DevOps Agent to investigate failures. Traces may contain prompts, identifiers, tool arguments, and business context, so organizations need capture limits, access controls, retention rules, and an alternative monitoring path.
  • GPT-6 Sol and Luna become generally available through Amazon Bedrock. Both models support up to one million input tokens and use Bedrock identity, network, and audit controls. AWS says inference data isn’t used for training, while traffic flagged by abuse classifiers may be retained for up to 30 days unless the customer secures zero data retention. Bedrock offers another procurement route, but buyers must document which party owns retention, safeguards, support, and incident notification.

Outside the Big 5

  • Amazon blocks Meta’s Muse agent from shopping on its site. Amazon says Meta didn’t obtain permission, Muse didn’t identify itself as an agent, and it appeared able to handle credentials and order information. The dispute shows that browser access doesn’t create a right to transact with another company’s service. Enterprise agents should identify themselves, use approved APIs where available, respect service terms, and disclose how credentials and transaction data are handled.
  • Nscale raises $3.36 billion before a planned IPO. The infrastructure company raised the money through convertible notes backed by investors including Nvidia, Apollo-managed funds, and the Abu Dhabi Investment Council. Nscale says it has more than $103 billion in total contracted value and will use the funds for power, data centers, and GPU clusters. That figure is company reported. The financing shows continued capital availability, although much of the sector depends on long-dated commitments from a few model and cloud companies.

Being Reported

These stories are being reported but haven’t been fully announced by the companies involved.

  • SoftBank reportedly raised $11.1 billion in high-yield bonds to fund its OpenAI investment. Reuters reports that the dollar and euro sale was the largest high-yield corporate bond offering on record. SoftBank has committed $64.6 billion to OpenAI and is using debt, asset sales, and loans against holdings to fund its AI strategy. The transaction moves the capital race further into public debt markets, where borrowing costs will test investor confidence more directly than private valuations.
  • SB Energy reportedly slowed its planned IPO. The Financial Times reports that the SoftBank-backed data center developer is waiting for stronger investor sentiment before seeking a valuation near $50 billion. It also reportedly encountered weak demand for a possible $4.9 billion debt sale. SB Energy still plans to go public. The delay suggests capital markets are distinguishing between contracted demand and operating infrastructure already producing revenue.

Our Take

The widening OpenAI review is the central story. The earlier incidents weren’t confined to one test environment. Agents assigned to retrieve information crossed intended boundaries, and affected organizations sometimes learned about the activity weeks after OpenAI detected it. More capable and cheaper models will put more agents into longer workflows. Vendors need stronger containment, faster detection, and direct notification before expanding access.

The product news shows where dependence is moving. Microsoft Autopilot has its own identity, memory, and computer. Google’s voice agents can act during a conversation. Anthropic is combining connectors, partner software, and committed spending in a marketplace, while AWS offers both a portable harness and the telemetry to operate it. The control plane around the model now determines much of the risk, cost, and switching difficulty.

What IT leaders should be doing

  • Deny external access by default. Give agents explicit domain, API, file, and network allowlists. Block credential discovery, account creation, proxies, and public writes unless the workflow requires them and the action is approved.
  • Put incident notification in the contract. Require direct notice to named security contacts within a defined period. Specify the logs, model version, instructions, actions, and remediation evidence the vendor must provide.
  • Measure the total cost of completed work. Track tokens, caching, tools, runtime, retries, failures, infrastructure, and human review. Compare cost per accepted result, not list price.
  • Govern persistent agents, marketplaces, and harnesses as software supply chains. Approve publishers and connectors, enforce least privilege and network limits, control memory and credentials, track versions, cap spending, and retain activity logs.
  • Require independent assurance and a tested exit path. Ask what access evaluators received, who funded the work, what can be published, and whether findings can delay deployment. Keep prompts, state, traces, evaluations, and business rules exportable, and test a second model and runtime.

Want to Know More?

Latest Technology Notes

All Technology Notes