Cross-Model Review Is Architecture, Not Heuristic
Single-LLM self-reflection has a structural blind spot — the model tends to confirm its own output. QuantGPT enforces a hard rule in factor-mine SKILL Phase 0.5: Claude must consult DeepSeek before designing a new factor family. Not a suggestion, a hard rule. This isn’t redundancy — it’s the antidote to structural bias.
Agent-Native Architecture: Designing Systems for Agents, Not Humans
When the system operator changes from a human to an LLM Agent, design principles need fundamental rethinking. Humans need GUIs and documentation. Agents need semantically clear tools and constraints that throw errors.
Harness Is Governance: Constraining Agents with Code, Not Prompts
The mainstream approach to Agent governance is writing rules in prompts. But LLMs can ignore prompts. Real governance lives outside the Agent, in the Harness layer — tools define what’s possible, errors define what’s forbidden, code defines where the boundaries are.
Why I Build Agent Infrastructure, Not Agents
While everyone is building Agent frameworks, I chose a different path: build infrastructure for Agents, not the Agent itself. Not because I can’t build Agents, but because the Agent layer is a consumable — infrastructure is an asset.
Skill Orchestration > Agent Loop Chains: Why Dumb Pipelines + Smart Tools Beat Smart Pipelines + Dumb Tools
Current AI Agent frameworks obsess over building complex loop chains: Planner → Executor → Reflector → Re-planner. I chose the opposite: tools are stateless pure functions, and the LLM decides the call sequence itself. Not because loop chains aren’t cool — but because they put decision authority in the wrong place.
The Expression Parser Is a Compiler, Not eval()
QuantGPT’s core is an 870+ line hand-written recursive descent parser supporting 80+ operators, automatic cross-sectional/time-series grouping, and dual-mode compilation. Not because I didn’t know eval() is simpler — but because what eval() can’t do happens to be what matters most.
API Guard Pattern: Why Calling Functions Directly Is Forbidden
QuantGPT uses threading.local to enforce a runtime guard: all backtest calls must go through the API boundary. Direct function calls raise an exception. Not because the function is dangerous — but because a system without boundaries can’t be audited.
Anti-Overfit Is Architecture, Not a Plugin
Most backtest systems treat anti-overfit as an optional add-on check — run the backtest, then test for overfitting if you feel like it. QuantGPT builds it into the scoring system and evolution engine: anti-overfit results directly affect factor scores, and the evolution engine reads anti-overfit metrics to decide its next strategy. Factors that haven’t proven robustness don’t even qualify for iteration.
If Research Isn't Reproducible, It Isn't Research
The most common lie in ML research is ’the results looked great last time.’ What code was used last time? What data version? What parameters? Nobody can say. I used filesystem transactions (temp directory → atomic rename) to create immutable snapshots of every iteration, turning ’last time’s results’ from a memory into a queryable fact.
Let AI's Code Run — But Don't Let It Run Away
AI-generated code must be executed — otherwise it’s just text. But execution means risk. I didn’t choose container isolation or RestrictedPython. Instead I designed a three-layer defense: reject dangerous structures at compile time via AST, replace the entire builtins at runtime, and enforce OS-level resource limits as a backstop. Each layer handles a different class of risk. Overlapping but not redundant.