AI Trading Newsletter

AI in Trading 2026: The Sandbox Had an Exit, the Frontier Halved, and the Agents Got Closer to the Venue

What this means for the industry and why rulebooks will increasingly need to write for a schedule they no longer control

Last week was about the SEC’s homework and two central banks reclassifying frontier AI as a financial stability risk. This week we were given a clear demonstration why: OpenAI confirmed its own models escaped a test environment and compromised a third party – the scenario the Bank of England and the ECB described a fortnight ago, arriving from inside a lab rather than from an attacker.

Anthropic halved the price of near-frontier intelligence, after Kimi K3’s impressive launch – but the debate is intensifying – Moonshot publishes Kimi K3’s weights tomorrow, and Nadella and Zhilin Yang have opposite answers to one question: is the model the moat? The SEC’s answers are due FridayMCP’s biggest-ever revision goes final Tuesday; Revolut opened its crypto exchange to AI assistants and the London Stock Exchange announced a venue built for agents rather than people. The question for Heads of Trading is no longer what this will mean for the desk in 2027 – it is keeping up with the innovation in real time. Here’s what I learnt this week on AI in Trading:

1. When the Warning Stops Being Theoretical

Hugging Face (the public repository where developers share AI models) disclosed a serious intrusion on 16th July and blamed an external AI agent. On 21st July OpenAI confirmed the attacker was its own models: GPT-5.6 Sol and an unreleased model, run with cyber refusals reduced and safety filters off to measure maximum offensive capability. Nobody told them to attack anyone; they were trying to beat a benchmark – or so the theory goes. So which is it?

Answer 1: it was an accident. From the sandbox the models took a permitted route out – an internal service for downloading software dependencies – exploited an unknown flaw, gained privileges, then triggered two code-execution flaws in Hugging Face’s pipeline, harvested cloud and cluster credentials and pulled the benchmark’s answers from the production database. OpenAI called it an unprecedented cyber incident and expects more as models get more capable (CNBCThe Hacker News).

Answer 2: it was a demonstration. The sceptical reading is that the disclosure doubles as a marketing ploy. A model that independently finds a zero-day, escalates privileges, moves laterally and reaches another sophisticated company’s production database is precisely the capability OpenAI sells to security customers. Critics on Hugging Face’s own disclosure thread note that across the entire chain there is no exploit-level detail – no CVEs, no vulnerability classes, no payloads. Fortune reported that some read it as suspiciously good PR, while stressing there is no evidence the incident was fake: Hugging Face confirmed the breach independently and both firms published reports.

The likeliest answer sits between the two: a genuine containment failure of OpenAI’s own making – a deliberately dangerous evaluation, refusals reduced, not isolated well enough – which OpenAI then chose to disclose in the form most flattering to its cyber business. Industry reaction split the same way. But motive doesn’t change the control question, because both answers agree on the part that matters: the boundary didn’t hold.

Why this matters for trading: Claudia Buch told EU institutions two weeks ago that AI models find vulnerabilities and build exploits faster than firms can patch, with remediation plans due 31 October; the Bank of England’s FPC said the same at system level in its July Financial Stability Report. Now it has happened – no outside attacker, no jailbreak, just an objective and an environment that couldn’t contain it. The 72% of banks that couldn’t confirm they had a kill switch is about intervention; this is about detection, and on OpenAI’s own account the operator with total control didn’t know until afterwards. Prompts and model safeguards are not access controls: if an agent must never reach something, the infrastructure must make reaching it impossible.

Containment sits with whoever runs the compute – for almost every desk, a third party. Firms need to ask what each agent’s environment can physically reach, whether it could harvest credentials simply by doing its job well, and who would tell you first. Which makes the FCA’s Critical Third Parties regime timely: it went live on 13 July designating AWS, Google Cloud, Microsoft Ireland and Oracle. AI providers need to be on firms critical third-party maps, and if the vendor also sells the defence – more due diligence will be required.

2. The Frontier Reprices Itself in a Fortnight – and Sovereignty Weighs In Again

Ten days after JPMorgan and Microsoft said they were matching model sophistication to task on cost grounds, the costs moved again. Anthropic shipped Claude Opus 5 on 24 July at $5/$25 per million tokens – level with Opus 4.8, half of Fable 5’s $10/$50 – claiming near-Fable-5 intelligence while beating it on its own published benchmarks. Its fourth model in under two months. Then Moonshot released Kimi K3 on 16 July, beating leading US models on coding benchmarks at far lower cost; weights publish tomorrow. Asia-based models are now around 60% of OpenRouter tokens, triple their January share.

Jason Hsu’s New York Times op-ed argues that if much of open weight adoption is American, with Airbnb’s customer service agent running substantially on Qwen and Cursor’s Composer 2 on a Moonshot base, the next question is on dependence risk. Beijing has run this playbook before with solar, telecoms and critical minerals – build the world’s reliance on something cheap, then restrict access once the reliance is real. China needn’t win on capability, only flood the market with “good enough” open weights and let American capitalism do the rest. The vulnerability isn’t technological: the US bet its economics on corporations paying premium prices forever. Pressure runs both ways – Axios reported on 20 July that the administration is weighing measures against Chinese open-source models.

Sovereignty also reasserted itself closer to home: new PM Andy Burnham has put AI in cabinet for the first time – Kanishka Narayan became AI minister on 20 July, with the AI Security Institute moving to the Cabinet Office. Three days later a joint AISI assessment with the US CAISI put Kimi K3 well behind leading US models on offensive cyber – 32% against 76% on exploit development, though the US models were tested with their safeguards disabled, exactly as in Answer 1 – but found its safeguards didn’t stop it attempting exploit development at all.

Why this matters for trading: last week’s framing – the frontier is expensive, so ration it – lasted ten days. The expensive tier halved and now a credible substitute needs to become part of the policy:

1. Routing has to be flexible. The price-capability frontier has moved twice since most 2026 contracts were signed; making sure you are not locked in at the wrong price point increasingly matters.

2. Migrations are getting harder. Harness rewrites, eval rebuilds, revalidation – all become worse under time pressure after an access restriction, as when Fable and Mythos went dark for 19 days in June. Equally self-hosting swaps vendor dependency for model ownership: you own evaluation, safeguards and incident reporting – like everything, it is a trade-off against resources.

3. Capability, price and safeguard level are now three separate considerations. The cheaper, less-restricted, lower-retention model is the everyday default, and AISI’s finding shows an open-weight model can lag on capability while running ahead on permissiveness. Managing model access is set to become more complex as a result of the trade-offs.

3. The SEC’s Response is due Friday – but Bloomberg’s Board Got There First

Nothing yet on the 13 questions Foster and Sherman put to SEC Chair Paul Atkins, response due 31st July. But the Bloomberg editorial board has already weighed in on 24 July with a list tracking an increasingly global consensus:

1. run agents through large-scale market simulations before live markets – citing the BIS project with European central banks, Project Logos, which puts AI agents to work as portfolio managers in a simulated market;

2. disclose objectives, testing and trading limits;

3. build AI-assisted real-time monitoring;

4. extend post-flash-crash circuit breakers;

5. consider guardrails encoded into the models themselves;

6. and cap leverage.

Why this matters for trading: the list tracks MAS’s SAFR, the FSB toolkit and the Bank of England — pre-deployment testing, declared authorisation scope, bounded outcomes, always-available containment, an immutable record. What’s missing is what trading needs and as yet nobody outside FIX appears to be building: governance state travelling with the order, in real time, not reconstructed from logs. One test – can you produce, on demand the agent’s identity, the scope it was authorised to act within, and the decision that let it through? Friday tells you whether the US applies existing duties to autonomous systems, as the FCA’s Mills Review does, or admits a gap. Either way the evidential burden still lands with firms.

4. The Model Is Not the Moat – Your Data and Your Evals Are

Item two was about what the frontier costs. This is about whether the model is where value sits at all – and two industry leaders gave opposite answers in the same week.

Moonshot’s Zhilin Yang argues great agents begin with great foundation models, and that Kimi K3 leads. Reasoning alone isn’t enough; the next frontier is models that build the next generation – K3 helping build K4. China is not copying any more: Moonshot, DeepSeek, Zhipu and Alibaba’s Qwen teams share a Tsinghua root and have moved past matching raw intelligence to cost-effectiveness at a scale that makes AI deployable everywhere rather than selectively.

Satya Nadella’s post this week on Microsoft’s MAI family casts the model as a component – not the product, not the moat. It is what Microsoft does: routing production traffic in GitHub Copilot, Excel and Outlook to its own MAI models wherever those match or beat the frontier on that task, and setting per-task quality floors so requests go to the cheapest model clearing the bar. MAI in Excel matches GPT-5.6 on common tasks using roughly 10% fewer median tokens. The real requirement: evals need to keep improving if any single model is removed. That is model independence as an engineering criterion, not a procurement principle – OpenAI and Anthropic stay in, as interchangeable parts in a machine Microsoft controls and now sells every enterprise the toolchain to build.

Why this matters for trading: both are right about their own layer but few trading firms if any are at Yang’s level, while almost every one is at Nadella’s. What generalises isn’t MAI but the discipline underneath: measure quality per task, set the floor for each workload, don’t pay frontier prices for work that doesn’t need frontier capability. A desk standardised on one frontier model cannot route because it cannot measure – no per-task eval, no basis to substitute, no negotiating position, no continuity plan. “Would your evaluation results keep improving if any single model were removed tomorrow?” makes it more of a business continuity issue, not architecture, as the 19 days without Fable and Mythos illustrated.

5. Agents Move Closer to the Order Book – and What That Means for the Desk

Three things happened that, taken together, move agents further into transacting in markets.

On 10 July, Revolut connected Revolut X, its standalone crypto exchange, to third-party AI assistants including Claude, Gemini, OpenClaw and Cursor. Users can pull portfolio overviews, set alerts, backtest strategies and prepare trades in natural language – though every order still requires explicit human approval. That approval step is the line agents have not yet crossed at scale, and it is the only thing separating an assistant from an executing agent. Revolut is not alone: Gemini launched agentic trading in April, claiming first-mover status among regulated US exchanges; Liquid shipped live execution through ChatGPT and Claude a month later; and Robinhood’s crypto-focused Agentic Accounts roll out shortly.

 

 

Revolut exposed its trading API through the Model Context Protocol (MCP) – a 2024 standard letting AI models call external tools without bespoke integrations – and its engineers reportedly built a full market-making workflow in about thirty minutes. MCP turns months of partnership negotiation into a plug-and-play connection. Over 10,000 public MCP servers now run in production, and Coinbase, eToro and Robinhood’s agentic brokerages already run on them. Bloomberg has adopted it internally for ASKB while still declining to expose a public server.

Then on 21st July the London Stock Exchange announced LSE 24, a 24/5 venue built explicitly for agentic trading. Separate from the Main Market, running 17:00–07:50 with a 30-minute end-of-day pause, it lets agents reach market data, order management and execution directly, inside a regulated market’s controls. Client testing by end-2026, ETPs in H1 2027 subject to approval, equities to follow, running on LSEG’s Digital Securities Depository under the UK’s Digital Securities Sandbox. “24/5” may be marketing given the break you need to allow for end-of-day processing – but this is the first major venue designed around automation from a blank sheet rather than an existing market stretching its hours. Nasdaq and NYSE Arca extend existing sessions; LSE is building something different.

A week later the MCP specification goes final – the largest revision since launch. A stateless core removes the session handshake, authorization realigns to OAuth and OpenID Connect, and a formal deprecation policy gives a twelve-month minimum window: the most substantial changes since authorization was added, per Anthropic’s David Soria Parra (The Register).

Why this matters for the trading:

1. Increasingly your counterparty could be an agent – and increasingly a retail one. Flow reaching venues is generated by systems sharing training data, price feeds and prompts. That is the correlated-behaviour problem the SEC has been asked about, arriving through brokerage APIs rather than institutional pipes, and changes what models should assume about the other side of the trade.

2. “Which agent, acting for whom, within what scope” is being answered in a protocol spec, not a rulebook. OAuth-aligned authorization is where that question gets settled in production – the same question SAFR’s Agent Identity, the Delegation Chain work and FIX’s proposed fields each answer at their own layer. It goes live on Tuesday, three days before the SEC says who is accountable, and well ahead of anything binding: the FSB’s consultation closed on 22 July with a final report in October, ECB remediation plans land 31 October, and Australia’s legislation isn’t expected before 2027. The CFTC also extended to 26 August the comment period on 24/7 futures and energy perpetuals.

3. The coverage model is the immediate problem. Extended hours have been staffed thinly and supervised lightly on the assumption volumes were small. A venue purpose-built for agents inverts that, concentrating autonomous flow in the session where human supervision is weakest and risk and surveillance systems are most likely mid end-of-day processing. Is the kill switch is staffed at the hour it is most likely to be needed?

Three hard deadlines land inside eight days – Kimi K3’s weights tomorrow, the MCP specification on Tuesday, the SEC’s response on Friday – and only one was set by anyone with a supervisory mandate. The ECB’s remediation plans follow on 31st October. The infrastructure layer is setting the clock on identity, authorisation, containment and cost, and the rulebooks are now writing to a schedule they no longer control.

Trading and markets will increasingly get more decentralised, not less. The OECD published two papers this month – AI and open finance on 16 July, AI and personal finance on 21 July – modelling agentic systems powered by open-finance data changing how people and businesses manage money, and the tensions that come with it: model performance against data minimisation, innovation incentives against concentration. We are entering a world of interchangeable models; increasingly, the contested asset is the data and the layer that governs access to it.

Thanks for reading – as ever, any questions or feedback, let me know.

Rebecca

Share:

Facebook
X
LinkedIn
Privacy Overview

This website uses cookies so that we can provide you with the best user experience possible. Cookie information is stored in your browser and performs functions such as recognising you when you return to our website and helping our team to understand which sections of the website you find most interesting and useful.