AI in Trading 2026: Claude Code and the Rise of Agent-Led Markets
The buzz around Claude Code this week highlights a structural shift in how LLMs could be deployed in secondary markets trading. As discussed in last week’s newsletter (https://www.mindfulmarkets.ai/ai-in-trading-2026-from-shifts-to-signals-what-market-participants-are-watching/), AI-assisted execution is moving from experimentation into core workflows, with machine-generated signals embedded directly into execution, risk, and surveillance frameworks, reshaping market microstructure in the process. Claude Code illustrates how quickly this transition could occur by replacing single-threaded tools with modular, autonomous systems that breakdown workflows into parallel tasks, dynamically reusing skills, and escalate to human oversight – all with built-in observability and controls that closely resemble institutional trading teams – but retaining the crucial trader oversight. With access to broad technical and scientific capabilities, AI is rapidly becoming a scalable research and execution layer, driving not incremental efficiency gains but a reorganisation of trading, risk, and monitoring functions. Recent developments – ranging from ELPs adopting advanced AI models to claims of AI-driven funds outperforming incumbents (https://bit.ly/4b2aKsZ) underscores the extent to which AI has already crossed into trading infrastructure. What matters now is not the technology, but firms’ data quality, governance, management practices, as well as their operating discipline. Here’s what I learnt this week on AI in trading:
1. So what can Claude Code do that is different?
Claude Code appears to represent a shift from AI as an assistant to AI as deployable, institution-grade infrastructure that mirrors how trading organisations actually operate (https://github.com/anthropics/claude-code). Its CLI-based approach offers firms the ability to install agents, commands, and integrations, configuring AI as reusable “teams” with defined roles, workflows, controls, and real-time observability – mapping directly onto execution, risk, compliance, and surveillance functions rather than fragile, one-off prompts. By making AI behaviour repeatable, shareable, and governable, Claude Code lowers the barrier to embedding multi-agent systems across the trade lifecycle, but it does not lower the standard required to make them work. As highlighted in the linked post (https://bit.ly/4jRSBk2), tools alone do not create agents – systems do. Agents that lack memory, fail when context shifts, or cannot reason across steps are still just prompts. Claude Code may be a game-changer, but the real test will be whether firms can design robust, stateful, multi-agent architectures: “it lowers the barrier to entry, but it does not lower the bar”.
2. What will be the wider impact for trading?
If Claude Code accelerates the shift toward agent-led growth, this could have significant implications for liquidity provision as market making and risk capital increasingly rely on predictive analytics rather than bank balance sheets alone. Firms already invested in advanced technology – particularly HFTs and electronic liquidity providers – gain increasing structural advantages over traditional incumbents, a dynamic seemingly illustrated by JPMorgan’s move to build a quant unit to compete directly with ELPs (https://www.bloomberg.com/news/articles/2026-01-15/jpmorgan-forms-new-quant-group-to-fend-off-market-maker-rivals). Emerging research and industry commentary suggest that AI – including the use of transformer-based models trained on high-frequency order-book data – is increasingly applied to predictive analytics and execution strategies, potentially reshaping how liquidity and risk are managed. While these developments may enhance market efficiency under normal conditions, they also raise concerns about feedback loops, model correlation, and systemic risk, highlighting governance and disciplined AI deployment as critical priorities – as well as reflecting the reality that leading firms are increasingly technology companies built around capital (https://www.hedgeco.net/news/01/2026/hedge-funds-2026-reset-governance-shifts-employee-ownership-and-ai-first-trading-infrastructure.html) – not necessarily as the stewards of secondary market activity to support the real economy.
3. Third-party vendors willing customers – or hostages?
Claude Code’s ability to generate and configure software – and to automate large parts of the coding, integration, and testing process – also potentially has significant implications for many simple SaaS platforms, which may increasingly be bypassed as teams build and customise functionality internally rather than purchasing external services. By reducing the cost and friction of developing routine features such as reporting, analytics, document workflows, or lightweight CRM tools, Claude Code weakens reliance on third-party vendors whose value lies primarily in incremental functionality. This dynamic echoes an industry observation that “SaaS companies don’t have customers, they have hostages,” (https://x.com/aakashgupta/status/2012393275685278080?s=12&t=hIZyz92X18xX5rIQ8fkNzQ) reflecting subscription models built on lock-in rather than differentiation. Vendors whose propositions rest on simple, repeatable services may therefore be pushed to move up the value chain into complex, hard-to-automate offerings such as deep connectivity, specialised compliance, or regulated integration – or face growing commoditisation. The likely outcome is greater bifurcation: commodity SaaS features become internalised, while high-complexity, high-integration services gain greater share and importance in trading workflows.
4. Understanding risk: from LLM-as-a-Judge to Agent-as-a-Judge
The shift from LLM- to Agent-as-a-Judge has important implications for multi-agent trading workflows because it demands greater verification. Single agents act as passive judges, producing reasonable-sounding outputs without validating data, assumptions, or actions – largely viewed as insufficient for trading. When multi-agent systems decompose the trade lifecycle into roles such as signal generation, validation, risk, compliance, and surveillance, context management and information integrity become critical. Dedicated “judge” agents that actively verify claims using tools, data, and cross-checks improve control and accountability, surface errors earlier, and shift evaluation from outcomes alone to full decision traces. As agents become more reactive or self-evolving, governance requirements increase, requiring clear constraints, continuous evaluation, and human escalation. Multi-agent trading will be therefore less about speed and more about building auditable, institution-grade systems where evaluation quality, data integrity, and operating discipline matter as much as model capability. Read more here – https://bit.ly/4pJkvQk.
5. A fundamental change in how risk is evaluated
When multi-turn agents operate across multiple states, using tools and memory as they adapt step by step, errors can emerge, propagate, and compound – forcing risk management to evolve beyond monitoring static model outputs. In its January 9, 2026 article Demystifying evals for AI agents (https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents), Anthropic outlines an evaluation framework that distinguishes tasks from trials, traces full multi-step interactions as transcripts rather than focusing solely on outcomes, and runs evaluations end-to-end through an evaluation harness. Central to this is the grader, which may be code-based, model-based, or human-led, each with distinct trade-offs. Applied to multi-agent systems, this implies continuous, transcript-level supervision using clear task definitions, repeated trials to measure variability, and layered grading that combines deterministic controls, AI reviewers, and targeted human oversight. Risk must be managed as a distribution rather than a point estimate, with memory introducing additional hazards such as drift and contamination that require explicit governance. Hallucinations become a systems risk, mitigated through source-gated claims, claim-level verification, and containment to prevent cascading failures.
6. A cultural challenge, not just a technical one
JPMorgan’s decision to form a new quantitative trading and research group to compete with electronic liquidity providers highlights a deeper divide between global banks and technology-native ELPs. In a LinkedIn discussion led by the CEO of XTX Markets (https://bit.ly/49sW8BJ), one comment stood out: “Wrote a simple market-making model during my tenure in that e-trading group mentioned by the Bloomberg report. It took six months for an internal independent model review group to approve.” This illustrates the structural latency often embedded in large institutions. While such governance is essential for managing systemic risk at scale, it contrasts sharply with the speed at which ELPs can design, test, and deploy models, raising questions about whether traditional banks – despite strong earnings, balance sheets, and renewed investment in quant capabilities – can continue to match the pace of innovation in markets increasingly shaped by fast-moving, technology-native firms. One to keep watching.
As always, thank you for reading – and a continued happy January.
Rebecca


