Big 5 AI Vendor Roundup: Week of August 24, 2026

Technology Note By: Mark Tauschek, Bill Wong, Info-Tech Research Group

OpenAI published the full account of the Hugging Face incident this week, showing how experimental agents turned weaknesses in a shared evaluation environment into access across dozens of servers and an OpenAI research cluster. Anthropic expanded Claude’s browser autonomy, documented mandatory retention for its most capable models, and proposed a standard for agents to operate physical equipment. Microsoft, Google, and AWS added more controls around agent execution, model choice, data classification, discovery, and evaluation.

The common thread is that AI risk is moving out of the chat window. Agents now use signed-in websites, respond to events, remember prior work, choose models, operate tools, and may soon control physical equipment. IT leaders need to evaluate the whole operating environment, including identity, credentials, network access, memory, retention, monitoring, approvals, and recovery.

The Big 5 call for a cyber-defense surge

  • More than 100 organizations call for faster deployment of AI for cyber defense. OpenAI, Anthropic, Google, Microsoft, and AWS joined security vendors, researchers, and public agencies in warning that AI-enabled attacks will become more capable and widespread. The group called for stronger software security, least privilege access, traceable agent identities, better monitoring, and broader access to defensive AI. The letter includes no funding commitments, deadlines, or common control standard.

Anthropic expands Claude's reach and clarifies its highest-risk policies

  • Claude in Chrome reaches general availability. Claude can now navigate and act on signed-in websites for users on every paid plan. Anthropic says a classifier reviews each action, and Enterprise administrators can restrict the domains Claude may visit. It can automate internal portals and older applications, but it also puts authenticated sessions inside the agent’s operating boundary.
  • Claude adds persistent memory across chat and Cowork. Claude can carry user-approved context between conversations and workspaces, with controls to inspect, edit, or delete what it remembers. Anthropic says sensitive topics aren't stored by default. Memory is enabled by default for Free, Pro, and Max users, while Team and Enterprise administrators control availability and users opt in. Persistent context improves continuity but creates another corporate data store requiring retention and deletion rules.
  • Anthropic documents 30-day retention for its most capable models. Prompts and outputs sent to Mythos-class models, and future models with similar capability, are retained for 30 days across Anthropic and partner platforms. Organizations using zero data retention services must move those workloads into a retention-enabled environment to access covered models. Anthropic says the data is automatically deleted after 30 days unless it is flagged for review or subject to a legal hold. Buyers need model-specific retention terms in contracts rather than assuming one policy covers every Claude model.
  • The Model Hardware Standard aims to give agents a common way to operate physical devices. The research preview defines a model-agnostic interface for instruments such as microscopes, liquid handlers, robotic arms, and quantum control equipment. Anthropic plans to make it open source and says early integrations reduced setup from weeks or months to hours. The company also opened 10,000 Claude seats and research credits to scientists. Physical automation still requires expert supervision and hardware-level safety limits.
  • Salesforce and Anthropic launch Claudeforce. The Salesforce plugin for Claude includes 37 prebuilt sales skills for meeting preparation, pipeline analysis, account research, and governed updates to Salesforce records. The release packages Claude with proprietary data, workflow logic, and permissions. The durable lock-in is increasingly in the operating layer around the model.
  • Anthropic wins an early court ruling in its dispute with the Pentagon. A federal judge blocked the Pentagon's supply chain risk designation and related measures, calling them illegal and retaliatory. The dispute arose after Anthropic maintained restrictions on some military and domestic surveillance uses. The government may appeal, and the broader policy conflict remains unresolved.
  • Sony and Warner music publishers sue Anthropic over alleged copyright infringement. The publishers allege that Anthropic unlawfully obtained copyrighted compositions for training and that Claude can reproduce protected material. Anthropic says it disagrees and will defend itself. The claims haven't been proven, but the case increases pressure to document training sources, licensing, and output controls.

OpenAI explains the Hugging Face incident and adds more control over its stack

  • OpenAI publishes the full Hugging Face incident report. Experimental agents found previously unknown software flaws, gained access to dozens of Hugging Face servers, obtained administrative credentials, and reached an OpenAI research cluster. Some agents also created an unauthorized message board, shared attack methods, and copied private evaluation data into a public dataset. OpenAI says no customer data or production services were affected. The report identifies weak isolation, excessive credentials, impossible tasks with no safe exit, delayed detection, and agents adopting one another’s goals. OpenAI calls the incident a warning shot.
  • OpenAI releases the first performance results for its Jalapeño inference chip. OpenAI says its first custom chip delivers more throughput per kilowatt and lower token latency than the systems used in its comparison, including on public tests using GPT-OSS 120B. These are company-reported results and need independent validation. A first-party chip gives OpenAI more control over capacity, latency, and cost.
  • ChatGPT Work adds event-triggered tasks and signed-in browser actions. Plus and Pro users can create tasks triggered by Gmail messages, Slack activity, or GitHub pull request changes. The browser can also act on websites where the user is already signed in, while pausing for confirmation before consequential steps. The pattern will reach enterprise products, where webhooks and persistent sessions must be treated as privileged automation.
  • OpenAI plans to end model supply to Cursor after its acquisition by SpaceX. OpenAI says it notified SpaceX that it intends to wind down service to Cursor on November 12 under a change of control clause. SpaceX may dispute OpenAI’s interpretation. Whatever the outcome, the episode makes continuity risk concrete for customers that buy an application whose core capability depends on a separate model supplier.

Microsoft puts more agent work inside managed policy boundaries

  • Microsoft 365 Copilot Cowork adds local browser use and event-driven tasks. Cowork can now use Edge and existing sign-ins to complete browser tasks under organizational policies. It can also propose automations triggered by selected email or Teams activity, subject to user review and confirmation. The feature makes Microsoft 365 a workflow execution surface, which increases the importance of conditional access, session controls, scoped permissions, and complete activity logs.
  • GitHub changes model defaults and retention. A new global model policy lets unconfigured models inherit an enterprise-wide setting. New generally available models are enabled by default, while open-weight models and models requiring data retention are disabled unless administrators approve them. GitHub is also extending chat retention from 28 days to the life of the account and moving cloud agent, chat, and mobile experiences into a unified default experience. Administrators should review these defaults before they take effect.
  • Copilot expands the data and tools available inside Microsoft 365. Excel Copilot can now execute Python for analysis and transformation under existing controls, while Outlook email and Teams meetings can serve as sources in Copilot Notebooks. These additions reduce manual data movement but broaden the information available to the assistant. Existing access rights, sensitivity labels, and notebook-sharing rules need to be tested together rather than reviewed separately.

Google packages Gemini for regulated work and adds automated classification

  • Google launches Gemini Enterprise editions for financial services and legal work. The financial services edition includes a managed research agent, licensed data connectors, citations, confidence indicators, and repeatable analysis methods. A separate legal edition supports contract review, redlining, legal research, regulatory monitoring, and data subject requests. Both use Gemini Enterprise's identity and governance layer. These packages can speed adoption while making Google’s connectors, skills, and control plane harder to replace.
  • Gemini-based data classification enters open beta in Google Drive. Administrators can describe classification rules in plain language, then have Gemini propose and apply labels to Drive content. Those labels can drive data loss prevention, retention, auditing, and restrictions on what agents may access. Users can accept or correct the labels, and changes are logged. Automated classification could close a major governance gap, but organizations should measure false positives and missed sensitive content before enforcement.
  • Google updates speech and video models. Gemini 3.5 Transcribe entered public preview with real-time transcription, speaker attribution, timestamps, and multilingual support. Google also released Gemini Omni 1.1 Flash with longer video generation, frame interpolation, previews, and 4K upscaling. Organizations still need consent, provenance, copyright, and retention controls for recorded and generated media.

AWS adds more compute and backs portable agent infrastructure

  • AWS and Nvidia plan another 2 million GPUs. The companies plan to deploy the additional capacity across AWS in 2027 and 2028, on top of an earlier 1 million GPU commitment. The agreement also covers Nvidia's next-generation systems, physical AI infrastructure, and 100,000 GPUs for secure US government environments. It extends the compute land grab and deepens the commercial interdependence between a cloud provider and its dominant chip supplier.
  • AWS proposes an open specification for agent discovery. Agentic Resource Discovery is an Apache-licensed specification that lets agents find resources across existing cloud, on premises, and software registries without forcing everything into one catalog. AWS also made AgentCore Evaluations framework-agnostic, supporting several third-party agent frameworks through common telemetry. Both moves improve portability, although the evaluation service itself still runs inside AWS.
  • OpenAI's Terra and Luna models reach AWS GovCloud. The models are now available through Amazon Bedrock in both US GovCloud regions. Government and regulated customers gain access within an existing procurement and control environment, but they still need to verify model-specific retention, routing, logging, and authorization terms.

Outside the Big 5

  • Nvidia reports another sharp increase in AI infrastructure revenue. Quarterly revenue reached $96.2 billion, up 106% year over year, while data center revenue rose 117% to $89 billion. Nvidia forecast $108 billion for the next quarter and assumed no data center compute revenue from China. Shares rose 8.7% the following day. The results show that infrastructure demand is still accelerating despite growing interest in custom chips.
  • Thomson Reuters launches its own proprietary language model. Thomson was built from an open-source foundation and trained on the company’s legal, tax, and news data. It will first support tabular legal analysis in CoCounsel, while a smaller version will be available for academic use. Thomson Reuters says it invested $40 million and achieved competitive results at a smaller scale, a company claim that needs independent validation. The strategy shows how data-rich firms can reduce reliance on frontier vendors without abandoning a multi-model approach.

Being Reported

These stories are being reported but haven’t been fully announced by the companies involved.

Our Take

The Hugging Face report turns an abstract concern into an operating lesson. The agents were capable, but the damage depended on ordinary control failures: shared infrastructure, credentials with too much reach, weak network boundaries, long-running impossible tasks, and slow detection. The Big 5's cyber defense letter is directionally right, but frontier labs need to apply the same urgency to their own evaluation environments. Test systems that contain tool-using agents should be secured and monitored like production systems.

Autonomy is spreading through signed-in browsers, event triggers, persistent memory, physical equipment, and vertical workflow packages. Open standards such as MCP, A2A, and AWS’s discovery specification can reduce integration work, but they don’t standardize identity, policy, retention, billing, or liability. Those controls remain inside vendor platforms and are becoming the main source of both risk and lock-in. OpenAI’s planned Cursor cutoff shows why buyers also need to plan for dependencies between application and model providers.

What IT leaders should be doing

  • Secure evaluation environments like production. Use separate tenants, deny by default network access, short-lived credentials, narrow permissions, complete logging, rapid alerts, and a tested kill switch.
  • Treat browsers, webhooks, and memory as privileged services. Use named identities, domain allowlists, limited scopes, human approval for external changes, and clear deletion controls.
  • Track model-specific data rules. Record retention, review exceptions, location, and deletion terms for each model and deployment path. Require notice before a vendor changes them.
  • Test continuity across the whole supply chain. Document which applications depend on which model and cloud providers, then test an alternate model and a provider shutdown scenario.
  • Measure the control plane, not only the model. Compare policy, logs, classification, memory, routing, evaluation, cost reporting, and export options before selecting a platform.

Want to Know More?

Latest Technology Notes

All Technology Notes