— AI Trading Newsletter

AI in Trading 2026: Capability does not equal Green to Go

As agents increasingly go mainstream, capability does not equate to authority – meaning can the desk keep up with the necessary governance as well as the infrastructure

All AI eyes were back on Washington this week. On 29 September President Trump signed an executive order directing federal agencies to call AI “Super Intelligence”, and the heads of Google, Anthropic, Meta, OpenAI, xAI and Nvidia signed the White House Accord on Super Intelligence, committing to capability monitoring, internal oversight and independent audits. There is no enforcement, by design: Vice President JD Vance told the companies to own the risk rather than seek regulation, and the President called the accord “morally binding”. Critics say it lets firms define safety for themselves.

London took a different route. In an Insights article, Bank of England Governor Andrew Bailey argued for understanding and testing self-improving AI before and after deployment, ahead of any new regulation, because, as he put it, “models will behave unexpectedly”. Westminster is divided: Parliament’s Joint Committee on Human Rights has called for a dedicated AI Bill and a single regulator so harm cannot fall between sectoral remits, with a government response due by mid-November. Meanwhile in Brussels, ESMA has made digital innovation a Union strategic supervisory priority from 2027, starting with firms’ use of AI and tokenisation.

For Heads of Trading, the direction is clear even if the model is not: whatever the approach, the controls around the model are the firm’s responsibility, and from 2027 EU supervisors will be checking them. Here are the five themes I learnt this week on AI in Trading:

1. Governing agents at the point of action

Nasdaq is rolling out a governed agentic AI environment inside Calypso. Clients can connect Calypso’s agents to their own AI through a Model Context Protocol (MCP) layer, in a sandbox with live oversight and no external data retention, which Nasdaq expects to help scale institutional adoption.

One World AI set out the principle: once an AI system can act outside itself, capability and authority must be separated, with authority checked at the point of action and evidence that does not rely on the model’s own account. The UK’s AI Security Institute agrees, calling defences beyond alignment, such as cybersecurity controls, “essential”. Training a model to behave is not a control.

Why this matters for trading:

  • Capability is not authority. A trader can type any order, but limits and pre-trade checks decide what reaches the market. Agents need the same separation, checked when they act, not when they are prompted.
  • Alignment is not a control. A desk cannot evidence a well-behaved model to a regulator. It can evidence entitlements, sandboxing, logging and human sign-off.
  • Evidence must sit outside the model. If the agent’s own explanation is the only record of why an order was sent, there is no audit trail.
  • Watch the workflows. Every child order can pass its checks while the parent strategy drifts. Test cumulative exposure, not just individual actions.
  • Every agent needs a named owner. A vendor can host the agent but cannot hold the accountability. That requires agent identification and human oversight, which we are working on in the FIX AI Working Group under GLEIF’s verifiable LEI (vLEI).

2. The buy side is learning to trust AI on its own terms

Millennium has rolled out more than 1,600 AI “digital twins” across its 7,000-plus staff, after a pilot in which 97% of users engaged daily. Twins start with zero system access, gain permissions one system at a time, never exceed their user’s entitlements, and carry their own identity and audit trail. Team-level AI Teammates follow, including a digital risk analyst co-developed with Anthropic.

At Balyasny, the ~$38bn multi-strategy fund, Chief AI Officer Charlie Flanagan explained in an interview published by Anthropic that every new model is tested on thousands of the firm’s own checkable tasks, standalone and inside its agent environment. Controls sit outside the model: approved data only, minimal tools, every action logged and human review of important output. A merger-arbitrage deal package now takes under a day rather than three to five.

Vendors are following. LTX’s agentic BondGPT lets credit traders automate repetitive tasks with a simple instruction, while data providers are opening up to clients’ own agents: S&P Global Energy via MCP servers, Rystad Energy inside Microsoft 365 Copilot, Energy Aspects through general AI assistants, and Wood Mackenzie via a shared MCP gateway where entitlements decide who sees what. The protocol matters less than the groundwork, clear definitions, units, timestamps and permissions, so an agent answers correctly and can show its sources.

Why this matters for trading:

  • Test on your own workflows. Before a model touches execution, test it on a private, checkable set built from the desk’s own orders, fills and TCA, the same discipline already applied to brokers and algos.
  • Start agents at zero access. Grant OMS and EMS permissions one system at a time, never beyond the trader’s own, with a separate identity and audit trail to answer the SM&CR question before it is asked.
  • Start with the repetitive work. Monitoring, pre-trade preparation and reporting are where the gains are emerging, freeing traders to focus where it matters, just as DMA let traders handle more tickets in the late 1990s.

3. Implementation in practice: integration, analytics, compliance and compute

Integration comes first. Trading Technologies has acquired TRAFiX, adding equities and equity options to its multi-asset platform. CEO Justin Llewellyn-Jones wants to end the “swivel effect” of traders juggling three or four EMSs, arguing that modern, modular technology is what makes data, and therefore AI, usable.

Analytics is the next gap. In The TRADE, Ingenuity Trading’s Naz Al-Khudairi argues the institutional algo experience is broken: customisations take weeks, and asking what a live algo is doing mid-order gets an answer tomorrow, if at all. AI has made fixing this cheaper and more accessible. BTON, meanwhile, pools execution data across firms to predict slippage and recommend brokers pre-trade, inside the EMS.

Compliance is the quieter bottleneck. Approval needs expertise in both the technology and the regulation, which usually sit in different departments, so the default answer is no. Even a tightly designed system, where a rules engine decides, the AI only explains and every call is logged, gives compliance nothing to measure against. ESMA’s own data shows how little client-facing AI exists as a result.

Compute is becoming the new latency. A study of a quant firm’s research agent found each request carried a median 172,353 tokens of history, only 1,656 of them new. Because each step depends on the last, the work cannot run in parallel: one session took 24.8 minutes on a dedicated 16xB200 cluster, but 93.1 minutes when 64 sessions shared it. As the FT reports, agents are burning through budgets and forcing a rethink of how AI is priced ahead of the frontier lab IPOs. Markets are responding: Architect’s AI Exchange, under regulatory review, plans GPU-hour futures and options.

Why this matters for trading:

  • Fragmented stacks break agents. If order state and fills mean different things in each EMS, an agent will be wrong quickly and consistently – which is why we are working on greater standardisation in the FIX AI Working Group.
  • Make intra-order transparency a broker review criterion.  AI is making real-time algo analytics cheap pushing TCA further away from look back to real-time and switching how analysis is incorporated (see point 4) – brokers will need to keep pace.
  • Treat agent-ready data like any new feed. Before connecting a vendor’s MCP server, check licensing for agent use, entitlements, and whether every answer traces to its source and timestamp.
  • Agree with compliance up front. Map each agent to the frameworks it will be judged against, such as best execution, RTS 6 and SM&CR. Supervisory guidance within those frameworks would unlock more adoption than new law, and ESMA’s 2027 focus on AI may provide it.
  • Budget compute like market data. Shared capacity nearly quadruples run time, so decide which agents get priority when markets move, include compute in vendor due diligence, and keep research agents in a sandbox, not production.

4. When signals shorten, what is the reference point?

Information is going stale faster. RavenPack research by Alan Liu and colleagues found news sentiment signals for the 1,000 largest US stocks work best on only the last two to four hours of news. Since 2020, the open-to-close information ratio rises from 0.14 using all news to 1.06 at two hours, though breadth collapses at one hour.

More of that flow will come from agents. Robinhood is rolling out agentic trading accounts that monitor markets and trade on customer instructions, with trade-by-trade approval on by default but optional.

Why this matters for trading:

  • Benchmarks were built for human time. If a signal decays within hours, VWAP may already contain the news the order was chasing, measuring lateness rather than execution quality.
  • Decision time is the new arrival time. TCA needs to timestamp when the information arrived and when the decision was made, not just when the order reached the desk.
  • Are models trained on multi information streams – not just the good data points: Pre-trade cost models and AI pricing calibrated on benign data will mislead.
  • Crowded windows mean crowded trades. If retail and institutional agents trade on the same few hours of news, they will pile into the same names at the same time, moving the benchmark with them.

5. ESMA’s 2027 supervisory priorities: what they mean for the desk

ESMA’s priorities don’t mention best execution or market structure, but they will reach the desk all the same. From 2027, “innovation with investor safeguards” will have national regulators mapping firms’ use of AI and tokenisation in processes affecting client outcomes, likely including ML-driven broker selection, predictive TCA, smart order routing and LLM trader tools. Supervisors will examine governance, testing, data quality and bias, run initial checks on the most affected firms, and have named over-reliance on a few technology providers as a key risk. DORA supervision is also intensifying: more on-site inspections, smaller firms in scope, and closer scrutiny of third-party registers and incident data.

Why this matters for trading:

  • 2027 is the mapping year. Build the inventory now, including AI embedded in vendor tools: broker algos, EMS features, TCA and pricing models.
  • Best execution is the benchmark. AI will be judged against existing obligations: best execution, RTS 6 where it applies, SM&CR and UCITS/AIFMD governance. For every AI-influenced routing or broker decision, show who owns it, how it was tested and how it serves the client.
  • Concentration is a double risk. One provider for your EMS, analytics and AI is both an innovation-priority concern and a DORA third-party risk.
  • DORA reaches the desk. OMS/EMS, FIX networks, market data and TCA providers belong in the register of information. A trading outage is a reportable ICT incident, an agent with OMS permissions is an attack surface, and the disciplines in theme 1 double as DORA evidence.
  • Use the window. ESMA plans to engage with the market on benefits and share positive use cases. Bring a well-governed, evidenced example, through trade bodies or the FIX AI Working Group, and help shape the approach rather than just receive it.

One to watch: tokenisation, 23/6 and market noise

Tokenisation and extended hours aren’t moving much liquidity yet, but they are adding noise. The SEC’s innovation exemption lets Tokenised Securities Venues trade tokenised NMS stocks outside Reg NMS, but caps them at 75 names and less than 0.25% of monthly volume; Larry Tabb called it “kludgy at best”. Coinbase won CFTC approval to clear fully collateralised derivatives, Goldman Sachs’ $100bn Treasury fund added tokenised settlement, Baillie Gifford took a tokenised bond fund into four markets, and the FCA brought UK crypto firms into full regulation. Bruce Markets, backed by PEAK6 and Robinhood, plans the first 24/7 US equity trading ecosystem, subject to regulatory approval.

When the Bank of England and FCA asked the market what tokenisation is good for, the answer came back as collateral. But finality set by contract rather than statute doesn’t survive a counterparty insolvency, and as Tradeweb’s Liz Kirby warned, clearing, margining and settlement still run on banking hours. For the desk, that means deciding which thin-hour and TSV prints count for valuation, TCA and surveillance. It also means redefining the close, VWAP and arrival price for a longer day, and deciding who owns the Saturday margin call.

The industry has moved from DMA to algos to complex algos, and agents are the next step on the same path. The principles haven’t changed: know what is acting, on whose authority, against what data, and what happens when it goes wrong. What has changed is the infrastructure, and the benchmarks we measure it against, which now have to work around the clock.

Thanks for reading as always.
Rebecca

Share:

Facebook
X
LinkedIn
Privacy Overview

This website uses cookies so that we can provide you with the best user experience possible. Cookie information is stored in your browser and performs functions such as recognising you when you return to our website and helping our team to understand which sections of the website you find most interesting and useful.