AI + Quant Engineering

Building production-grade AI-orchestrated quantitative research systems. Writing about architecture decisions, engineering trade-offs, and lessons learned.

Cross-Model Review Is Architecture, Not Heuristic

Single-LLM self-reflection has a structural blind spot — the model tends to confirm its own output. QuantGPT enforces a hard rule in factor-mine SKILL Phase 0.5: Claude must consult DeepSeek before designing a new factor family. Not a suggestion, a hard rule. This isn’t redundancy — it’s the antidote to structural bias.

May 8, 2026 · 8 min

Agent-Native Architecture: Designing Systems for Agents, Not Humans

When the system operator changes from a human to an LLM Agent, design principles need fundamental rethinking. Humans need GUIs and documentation. Agents need semantically clear tools and constraints that throw errors.

May 1, 2026 · 5 min

Harness Is Governance: Constraining Agents with Code, Not Prompts

The mainstream approach to Agent governance is writing rules in prompts. But LLMs can ignore prompts. Real governance lives outside the Agent, in the Harness layer — tools define what’s possible, errors define what’s forbidden, code defines where the boundaries are.

May 1, 2026 · 5 min

Why I Build Agent Infrastructure, Not Agents

While everyone is building Agent frameworks, I chose a different path: build infrastructure for Agents, not the Agent itself. Not because I can’t build Agents, but because the Agent layer is a consumable — infrastructure is an asset.

May 1, 2026 · 7 min

Skill Orchestration > Agent Loop Chains: Why Dumb Pipelines + Smart Tools Beat Smart Pipelines + Dumb Tools

Current AI Agent frameworks obsess over building complex loop chains: Planner → Executor → Reflector → Re-planner. I chose the opposite: tools are stateless pure functions, and the LLM decides the call sequence itself. Not because loop chains aren’t cool — but because they put decision authority in the wrong place.

April 30, 2026 · 9 min

The Expression Parser Is a Compiler, Not eval()

QuantGPT’s core is an 870+ line hand-written recursive descent parser supporting 80+ operators, automatic cross-sectional/time-series grouping, and dual-mode compilation. Not because I didn’t know eval() is simpler — but because what eval() can’t do happens to be what matters most.

April 29, 2026 · 4 min

API Guard Pattern: Why Calling Functions Directly Is Forbidden

QuantGPT uses threading.local to enforce a runtime guard: all backtest calls must go through the API boundary. Direct function calls raise an exception. Not because the function is dangerous — but because a system without boundaries can’t be audited.

April 28, 2026 · 3 min

Anti-Overfit Is Architecture, Not a Plugin

Most backtest systems treat anti-overfit as an optional add-on check — run the backtest, then test for overfitting if you feel like it. QuantGPT builds it into the scoring system and evolution engine: anti-overfit results directly affect factor scores, and the evolution engine reads anti-overfit metrics to decide its next strategy. Factors that haven’t proven robustness don’t even qualify for iteration.

April 27, 2026 · 6 min

If Research Isn't Reproducible, It Isn't Research

The most common lie in ML research is ’the results looked great last time.’ What code was used last time? What data version? What parameters? Nobody can say. I used filesystem transactions (temp directory → atomic rename) to create immutable snapshots of every iteration, turning ’last time’s results’ from a memory into a queryable fact.

April 3, 2026 · 8 min

Let AI's Code Run — But Don't Let It Run Away

AI-generated code must be executed — otherwise it’s just text. But execution means risk. I didn’t choose container isolation or RestrictedPython. Instead I designed a three-layer defense: reject dangerous structures at compile time via AST, replace the entire builtins at runtime, and enforce OS-level resource limits as a backstop. Each layer handles a different class of risk. Overlapping but not redundant.

April 3, 2026 · 8 min