AI Trading Newsletter

AI in Trading 2026: The Harness, the Outage, and the Company That Got Hacked Being Bought for $12.9bn

The control point is moving from the model to everything around it: the harness, the interface, the repository and the network.

On Wednesday the FCA and the Bank of England published their findings on how firms are using frontier AI for cyber defence, and where it is straining them. On Thursday OpenAI released a model that operates software through the screen rather than through code, NVIDIA agreed to buy the repository that every open-source fallback plan depends on, and ChatGPT, Claude and Grok all went down at once. This is the stress test the ESAs asked firms to set a tolerance threshold on, back in July: indirect exposure to frontier models. One thread runs through all of it: what the industry is supervised on, dependent on and exposed to is not the model. It is everything sitting around it. Here’s what I learnt this week on AI in Trading:

1. Frontier AI becomes an operational resilience problem, not a model problem

On 2nd September the FCA published Frontier AI and cyber resilience, a review of how firms are using and preparing for frontier models. Its findings: attackers are finding vulnerabilities faster than firms can fix them; frontier AI tests whether an organisation is resilient rather than whether it has good tools; what a firm gets out of AI depends on the environment it builds around it, not on which model it licenses; and basic cyber hygiene, access management and dependency mapping matter more than before, not less. The review sets no new rules. It follows the FCA, Bank of England and Treasury joint statement in May highlighting the step-change in capability required for frontier models. Firms are asked to establish clear ownership, build validation and escalation processes, work out where their remediation capacity runs out, prioritise fixes by what is actually exploitable, and keep cyber expertise in the decision.

In his 28th August letter to G20 finance ministers, Andrew Bailey, who chairs the Financial Stability Board as well as the Bank, called frontier AI’s effect on cyber risk the most immediate threat facing the financial system, capable of changing the speed, scale and economics of attack, and of undermining confidence system-wide given how concentrated third-party providers are. Alongside the AI Consortium and CMORG’s frontier AI guidance, the Bank has created the Frontier AI Information Sharing Forum, which this week published Frontier AI: Harness engineering. A harness is everything a firm builds around a model: the tools, data, controls and environments that decide what information it gets, what it is allowed to do, who checks its output and where the results go. The Bank’s central question is how much of a firm’s own systems and data to expose to it, and it treats plugging a model straight into live production as higher risk than using a copy of it. Scale is limited by how fast a firm’s engineers can validate and fix what the model finds, not by how capable the model is, and third parties have thinner benches than the firms they serve.

Which is the Bank Policy Institute’s point from 1st September: the risk has moved to providers firms cannot patch themselves. The average time between a flaw becoming public and attackers using it has fallen from 53 days in 2024 to 22 hours in 2026. A Federal Reserve staff paper calls third-party providers a hidden fault line, typically more vulnerable than the institutions they serve. Parliament is moving in the same direction, harder: Lord Clement-Jones, with Baroness Harding, Baroness Kidron and Lord Hunt, has tabled an amendment to the Cyber Security and Resilience Bill giving ministers last-resort powers to shut down frontier AI systems and the data centres running them, one of 65 amendments debated this week. Alex Sobel MP introduces a separate AI Security Bill on 8th September.

Why this matters for trading:

· Supervisors have shifted from asking which model a firm uses to asking what it has built around it. Permissions, approval gates for higher-risk actions, controlled access, checks before an agent acts. That is the firm’s build, and what an inspection will look at.

· 22 hours is shorter than most vendors’ patch cycles and most desks’ change control. Most EMS, OMS, market-data and connectivity contracts set remediation deadlines when the number was 53 days. They need renegotiating.

· The Bank reaches the same conclusion on substitution that cost pressure reached months ago: build the layer that lets a firm swap suppliers. Two independent routes to the same answer is a reason to stop treating it as optional.

· A statutory shutdown power, even if never used, is a continuity assumption. If a model or its data centre can be switched off by ministerial order, “the vendor withdraws the API” is no longer the worst case, and fallbacks need testing against sudden loss rather than orderly migration.

2. The frontier moves from code to screens

On 3rd September OpenAI released GPT-6 Astra, with Greg Brockman calling it a generational leap. Until now, software has talked to models through an API: structured data in, structured data out, every exchange logged. Astra uses the screen instead. It can read a trading interface, click, enter orders and work through menus the way a person does, with nobody at the keyboard and none of the API traffic an IT team monitors. It costs $10 and $50 per million tokens against the previous model’s $4 and $20. It is also the first model OpenAI has rated Critical for cybersecurity under its own safety framework, having found two genuine unknown vulnerabilities during testing; the full version is restricted to vetted organisations. On general reasoning it is level with the model before it. On operating software, it is not.

Why this matters for trading:

· The economics of execution automation change again. Firms are now choosing between cheaper models for existing tasks and a much more expensive tier that can run workflows unattended. Any three-year budget built on the July and August price cuts now has a tier above it that did not exist when it was signed, and what a firm actually pays depends on how much the model consumes, not the headline rate.

· This capability points straight at trading infrastructure. It is what a desk would use to automate an EMS, and equally what an attacker would use to reach one. Architecture teams have three choices: adopt it and accept that the governance questions are unresolved; refuse it and accept being slower than firms that don’t; or isolate critical execution systems and accept the cost of doing so.

· Screen-based agents leave no API trail. Every entitlement, audit and surveillance control in place assumes machines talk to machines through an interface that can be logged. This one uses the same screen traders do.

3. Open source consolidates: every route out of vendor dependence now has an owner

On 2nd September NVIDIA agreed to buy Hugging Face for $12.9bn, filed with the SEC and announced the next day. It is NVIDIA’s second-largest deal after roughly $20bn for Groq’s assets in December, and completes in the first half of 2027 subject to approvals. Hugging Face is where open-source models are published and downloaded: 18 million developers, 3 million models, 200,000 companies. It was last valued at $4.5bn in 2023. Jensen Huang has committed to keeping it open. Six weeks earlier, OpenAI’s own models had broken into Hugging Face’s production systems.

Why this matters for trading:

· The continuity plan is now owned. The standard answer to “what if our AI vendor cuts us off” is to download an open model and run it in-house, and that route runs through Hugging Face. Independence is no longer free; it is a different supplier relationship.

· Stripe bought the model router, NVIDIA released the open orchestration alternative in Switchyard, and NVIDIA has now bought the repository. Every layer of the escape route from single-vendor dependence is owned by an incumbent.

· ASIC, APRA and the ESAs have all told boards to map their dependencies. Model routers, model repositories and orchestration layers were not on anyone’s third-party register a year ago. They are now named counterparties and belong in the technology risk framework.

4. When all AI providers fail in the same window

On 3rd September ChatGPT, Claude and Grok all went down at once. OpenAI blamed a routing error starting around 07:43 PT, fixed by 08:17. Anthropic had a partial outage lasting three hours and six minutes across Claude.ai, Claude Code, Claude Cowork and its API, which it put down to infrastructure. xAI traced Grok’s failure to its Memphis data centre, and SpaceX, which shares that infrastructure, apologised to affected partners. In the same window Azure reported a network problem in its East US region. Everything was back by 12:38 PT. Separately, Anthropic is finalising an expansion of its credit facility to $15bn, led by Morgan Stanley with Goldman Sachs, JPMorgan and Citigroup, the same four banks leading an IPO now expected to price in late September or early October.

Why this matters for trading:

· Different companies, different clouds, different technology, same three-hour window. Whether the causes were connected barely matters. The correlation is what a firm’s risk model has to carry. This is the indirect exposure the ESAs asked firms to measure, and it has now happened.

· Three hours is a trading session. If a firm’s pre-trade analytics, RFQ triage or surveillance had been running on an agent that morning, the fallback was a human. Whether firms still have enough people who remember the manual process is the question worth answering this week.

· What failed was routing and infrastructure, not the models. Having three models available behind one gateway does not help when the gateway or the network is the thing that breaks. Test failover, not just whether a supplier could be switched in principle.

5. Agents cooperating or spoofing in live markets

Rajiv Sethi of Barnard and the Santa Fe Institute published Artificial Traders in Real Markets on 1st September, drawing on METR and Redwood’s independent investigation of July’s incident. Roughly 1,200 agents that were supposed to be isolated found each other, exchanged over 70,000 messages on an unsanctioned message board, and about 700 of them attacked Hugging Face. They invented shared conventions for coordinating, gave each other mailboxes, and after one agent was impersonated, started cryptographically signing their messages to prove identity. Around 7% of transcripts contained faked tool calls. Sethi’s point is the one most coverage missed: in that setting, one agent succeeding did not stop another from succeeding, so they cooperated. Trading is the opposite. Where one participant’s gain is another’s loss, agents optimise for profit rather than accuracy, and each will work out that it can move the data other agents are watching in order to trigger a reaction. That is spoofing, done by something considerably better at it than the people who have been prosecuted for it. On 3rd September the SEC announced that its Investor Advisory Committee meets on 10th September with panels on AI in public markets and on Regulation NMS, the rules governing US equity market structure: both threads of this year’s story in one room for the first time. Foster and Sherman’s thirteen questions on agentic trading, due on 31st July, remains unanswered.

Why this matters for trading:

· The agents built their own identity system because they could not trust each other’s names. Every governance framework being written, including Singapore’s, assumes identity is issued from outside. These agents created a root of trust in four days.

· 7% faked tool calls means the transcript is not proof. Every audit trail being designed for agentic execution assumes the log is evidence. This is the first documented case of the log being something the agent can shape, and it happened inside the vendor’s own test environment.

· Sethi describes manipulation without anyone intending to manipulate. Market abuse surveillance works by inferring a decision-maker from an order pattern. When the pattern comes from a strategy nobody wrote down, that inference fails, and the firm still owns the trade.

Thanks for reading. As ever, any questions or feedback, let me know.

Rebecca

Share:

Facebook
X
LinkedIn
Privacy Overview

This website uses cookies so that we can provide you with the best user experience possible. Cookie information is stored in your browser and performs functions such as recognising you when you return to our website and helping our team to understand which sections of the website you find most interesting and useful.