Ai
September 30, 2026
1 views
2 min read

I Used AI Agents to Write Most of a Live-Trading Codebase: Introducing The Gate That Made It Safe

Curated by Patrick
Source: HackerNoon
I Used AI Agents to Write Most of a Live-Trading Codebase: Introducing The Gate That Made It Safe
Tech Daily Byte Analysis

Over the past 21 months the engineer behind TopSet—a self‑funded, end‑to‑end trading system that selects ML‑driven portfolios, rebalances on schedule and executes via broker APIs—let large‑language‑model agents write roughly 1,600 pull requests, a volume he could not have typed himself. To keep the system safe, he shifted his focus from line‑by‑line code review to test‑centric validation: every PR must pass a CI pipeline that runs lint, mypy, about 7,500 unit tests across 331 modules, a database‑emptiness assertion, migration rollbacks and Terraform validation on two AWS accounts. A weekly “heavy” suite adds 100 stochastic mock‑broker rebalancing runs, deterministic model‑snapshot checks and leakage scans, exposing race conditions that surface only under concurrent load. This strategy caught a critical bug where a zero‑amount order slipped past a guard, triggered a database constraint failure, and caused the scheduler to repeatedly retry a stuck rebalancing—a failure the developer only discovered after the live system reproduced it.

The approach reflects a broader shift where developers treat AI‑generated code as a productivity accelerator but compensate for its hallucinations with rigorous automated testing. TopSet’s reliance on mock brokers and probabilistic fills mirrors industry practices at firms like QuantConnect and Alpaca, which also simulate market microstructure to validate algo behavior before deployment. By integrating LLMs such as Claude Code, Cursor and GitHub Copilot into a continuous‑integration pipeline, the engineer demonstrates a practical model for “AI‑first” development that still respects the zero‑tolerance environment of live finance. The emphasis on integration and stress tests, rather than unit tests alone, underscores a growing consensus that AI‑written code can pass superficial checks yet fail in complex stateful interactions.

The experience warns that test suites can still give false confidence: a green build did not catch the zero‑amount bug because the guard prevented the decision but not the ORM mutation. Future safeguards may need to include mutation‑aware assertions or runtime invariants that verify object state before persistence. Watching how TopSet evolves its test hygiene—especially around snapshot management and flaky‑test mitigation—will indicate whether AI‑augmented development can scale safely in high‑stakes domains like automated trading.

Key Takeaways

The TopSet engineer generated most of his live‑trading code with Claude Code, Cursor and Copilot, producing 1,600 PRs in 21 months.

All changes must clear a CI gate that runs 7,500 unit tests, type checks, Terraform validation and a database‑emptiness assertion across two AWS accounts.

Weekly stress runs simulate 100 probabilistic broker fills, exposing concurrency bugs that unit tests miss.

A zero‑amount order bug slipped past the guard because the test only verified the decision, not the ORM state mutation, highlighting limits of test‑only safety nets.

About the Source

This analysis is based on reporting by HackerNoon. Here is a short excerpt for context:

1,600 AI-agent pull requests on a system that trades real money. Tests became the gate. Here are three times a green build lied, and what I changed.
Read the original at HackerNoon

More in Ai