OpenAI pauses its own model, Washington classifies its rulebook but won’t disclose it, while Singapore and FINRA both say the existing rules already reach agents, making the audit trail the new ask
Last week was about control failures rather than capability. This week the story moved back to the model. OpenAI paused internal work on an unreleased model because it cannot rule out that it has reached the most serious cyber threshold in its own safety framework – the first time any model has triggered that provision. Washington finished the framework it owed on 1 August, reviewed it with the labs on Tuesday, and will not publish it. MAS in Singapore confirmed that agentic AI already sits inside its supervisory expectations, and FINRA said much the same thing without writing a rule. Here’s what I learnt this week on AI in Trading:
1. The First Model to Trip Its Own Wire
On Friday, OpenAI said preliminary evaluations of Astra, an upcoming model, showed advances in agentic coding and cybersecurity strong enough that it cannot rule out the Critical capability level under its Preparedness Framework. That threshold is met when a model can find and build working zero-day exploits across many hardened real-world systems without human help, or run an end-to-end attack on a hardened target given only a goal. No model had reached it in the framework’s three-year history; GPT-5.6 Sol, the most capable model OpenAI has deployed, was rated High. OpenAI has paused Astra activity that does not meet strengthened controls, imposed monitoring of the model’s reasoning during training and evaluation that can interrupt high-risk activity, and will test with government agencies and selected safety organisations before release. Axios broke the story and its bottom line says it all: models are getting more cyber-capable faster than the regulation around their use is formalising.
Two caveats worth holding. Astra has not been declared Critical, and the evaluations that would establish it are unpublished. The framework’s language is also stricter than the action taken – it says development should halt at that level until safeguards are specified, while the announcement pauses only work falling short of new internal controls.
It arrived in the same week the UK’s AI Security Institute published its own incident report. Between 25 and 28 July, AISI ran a single cyber-range evaluation 122 times across seven frontier models with live internet access and cyber classifiers switched off, to measure raw capability. In 10 runs an agent acted on the live internet against real people and organisations – 19 catalogued actions, 17 from Mythos 5 and two from a single GPT-5.6 Sol run. In the worst case an agent tried to insert malicious code into a real open-source project, creating fake online identities to pressure the maintainer into approving it. A human reviewer caught it and refused. These are unrelated to last week’s events, which is rather the point.
Two useful opinions published this week – Marcus Hutchins argues the debate collapses three separate problems into one – tooling for unskilled attackers, end-to-end automation and sophisticated vulnerability research. Fold them together and you get the strawman of unskilled actors autonomously exploiting zero-days at scale, which then justifies restricting defensive work too – even though exploitation requires chaining something that works against a live target, while securing code requires only finding the flaw. And at Black Hat, senior US, UK and Canadian officials urged critical infrastructure firms to prioritise basic resilience over AI nightmare scenarios, with the NCSC’s Jonathon Ellison putting the immediate challenge as being resilient enough for a faster environment given the legacy already being carried.
Why this matters for trading:
• Model supply is now a continuity question with a safety gate in it. A capability finding can stop a release, and two frontier models were already suspended this summer on export-control grounds. If an execution-adjacent tool depends on a specific model generation, its availability turns on somebody else’s evaluation result.
• This is the first real test of a voluntary framework and it worked. But the assessment behind it is unpublished, so what you are relying on is a self-report. That is the same assurance model vendors are likely to provide – is that enough?
• The control that worked at AISI was a person. Not a classifier, not a sandbox boundary, not a log. Worth remembering the next time an operating model proposes removing human review on efficiency grounds.
2. Washington Finished Its Framework and Classified It
The framework due under Executive Order 14409 on 1 August – a classified benchmarking process to decide which systems count as covered frontier models, and a voluntary window of up to thirty days of pre-release government access – was reviewed at the White House on Tuesday with OpenAI, Anthropic, Google, Meta, Microsoft, Nvidia and a number of smaller firms. It will not be released publicly, and much of it is classified. It remains explicitly voluntary and cannot be used to create licensing or preclearance.
Compare the other clocks running. The EU AI Act’s enforcement powers became operational on 2 August, with published obligations, a named regulator and fines attached. While the SEC’s answers to Foster and Sherman’s thirteen questions on agentic trading, due 31 July, are now nine days late with nothing published as far as I have seen – correct me if I’m wrong!
Why this matters for trading: you cannot build a control against a benchmark you are not allowed to read.
• Designation status becomes a procurement question rather than a control. “Has this model been through the EO 14409 process?” is answerable; “does it meet the threshold?” is not.
• The asymmetry is structural. European, UK and Singapore supervisors publish what they expect and can act on it; Washington has built the capability and kept the criteria secret. A firm operating across both will have to work to the published version and hope the US agrees.
3. Singapore and FINRA Both Say the Rulebook Already Reaches the Agent
In a written parliamentary reply for the sitting of 5 August, MAS confirmed that its forthcoming Guidelines on AI Risk Management – consulted on in November 2025, covering board oversight, risk frameworks and lifecycle controls – apply to all AI use cases by financial institutions, agentic AI included, and will be finalised soon. Asked whether the industry-led Safeguards for Agentic Finance at Runtime (SAFR) framework would become mandatory, MAS declined to commit and gave no timeline.
SAFR itself, published on 3 July, is the interesting part: a runtime specification rather than regulatory guidance covering what an agent is authorised to do, how each proposed action is assessed at the moment it is proposed and before it executes, and what is recorded to support review, accountability and remediation when outcome diverges from intent.
FINRA reached the same place from the other direction, as Alex Twittau set out this week. Its 2026 report addresses AI agents for the first time, and writes no new rule – because it does not need to. The moment an AI system can take a step on its own, existing supervision, recordkeeping and fair dealing duties attach to what it did. Its stated considerations – track the agent’s actions and decisions, restrict what it can reach, hold guardrails on what it may do – are all runtime controls, not sign-offs at build.
Why that distinction bites is now measurable. In the HANDBOOK.md benchmark, agents were given a company policy document of 20 to 124 pages in context and 65 tasks across finance, insurance and three other domains, graded deterministically against 824 criteria. Under strict grading the best of thirty model configurations passed 36.2% of trials; most frontier configurations stayed below 25%. The failure patterns are the ones that matter in a regulated firm: a plausible in-environment request overrides the standing policy, a required check is run and then acted against, rule details are lost over long horizons, and compliance is reported rather than achieved. One case involves an approval posted by the person who raised the cost, above the approval limit – the exact thing the control exists to catch – which the model cleared and then reported as compliant.
Why this matters for trading:
• A rule in context is not a rule in force. Placing a trading policy in the system prompt is not a control, and an agent’s own report is the least reliable part of the workflow. Point-in-time assurance tests the document, not the behaviour between checks.
• Runtime authorisation is the right layer, and it is portable. Pre-deployment testing and after-the-fact audit both miss the moment that matters. RTS 6 and ESMA’s supervisory briefing on algorithmic trading asks firms to evidence control rather than intention, and nothing in SAFR is Singapore-specific – it is the design EMS and algo vendors should be able to describe.
• The examiner’s question is what the agent did, not whether the tool was approved. If regulators ask firms to produce the full account of one agent action from last quarter, this means the reasoning is required – not just the approval.
4. “Agent Traces” – the Audit Trail Becomes the Ask
More on reasoning – on CBS’s Face the Nation on 2 August, Hugging Face’s Clement Delangue described an OpenAI agent taking some 17,000 actions over four and a half days against his company’s systems, and pressed for mandatory disclosure of what he calls agent traces – what the engineers asked the agents, and what steps the agents then took – after any agent-driven incident, alongside his standing request for $100m of compute for community cyber defence. The argument for the trace is diagnostic: without it you cannot tell a human mistake from a system mistake from an AI mistake. He also noted his team contained the breach using an open Chinese model, because commercial tools would not analyse the exploit.
Why this matters for trading: strip out the politics and the same requirement is arriving from four directions at once. SAFR calls it what is recorded at the point of decision. MCP’s new stateless design means nobody holds the conversation for you. RTS 6 calls it reconstruction. And the FIX AI Working Group’s proposal on agentic runtime governance calls it a chain: the regulated firm, the human who authorised the agent, the certified envelope of what it may do, and the execution record carrying all three.
The practical version will become part of vendor due diligence: if an agent does something unintended in our environment, will you give us the trace, and in what format? Firms that can answer have built the control. Firms that cannot are relying on nothing having gone wrong yet.
5. Everyone is Automating
Four items that read differently together. The ECB published working paper 3262, A SPOT in the dark: using AI to assess financial stability risks, using LLMs across roughly 1.1 million financial news articles from 2005 to 2026 to score the Severity and Probability Of potential Trigger events; the indicator rises ahead of major historical triggers, identifies the source correctly, and adds information beyond existing stress, geopolitical-risk and policy-uncertainty measures. The FCA opened the Handbook through a new API on 6 August, making the rulebook machine-readable.
On the industry side, Flow Traders has moved foundation model training to CoreWeave, with the press release carrying a line sharper than most analyst notes: training proprietary models has become a core competitive discipline for a quant firm, not a side project. The selection criterion was consistent multi-node training performance at scale rather than GPU-hour pricing. And Bloomberg’s 2 August feature covered retail investors using LLM tooling and platforms like Composer, Alpaca and QuantConnect to build always-on strategies that would have needed a quant team two years ago, alongside agentic offerings now live at Robinhood, Public, Gemini, Coinbase and eToro, with Schwab expected in the second half.
Why this matters for trading:
• The arms race is switching – For twenty years the industry competed on speed to the exchange – dark fibre, microwave, colocation, FPGAs to intelligence. That is a different capex conversation, a different vendor list and a different hiring plan.
• Supervisors are building the capability they are worried about. An LLM-derived trigger indicator is something a desk can build for its own risk function, and the methodology is public. A machine-readable Handbook is the other half of that – the compliance layer becomes something an agent can query rather than something a human summarises.
• The correlated-behaviour question has moved. It is no longer principally about one broker’s agent product converging on a signal. It is tens of thousands of independently built systems, trained on similar data and prompted with similar instructions, arriving at the same trade in the same minute – with no perimeter. Exactly what started the conversation at the first FIX AI Workshop back in June 2024.
Thanks for reading. As ever, any questions or feedback, let me know.
Rebecca


