Author here. Since late July, GPT-5.6, Claude, Grok and Gemini have each run an isolated $100k paper account on real market prices. Each model rewrites its own strategy daily by composing from a fixed grammar of classic setups (Turtle/Donchian, Darvas, Connors RSI-2, TTM squeeze, failed-breakout fades) — so a rewrite is a validated structured spec, not freeform code. A fifth account runs a frozen rulebook as the control. After three weeks the frozen rulebook is +15.6%, the best model +5.7%, S&P +5.1%.
Three things I measured that I didn't expect:
1. Daily self-rewriting adds almost nothing. Correlation between rewrite count and performance across arms: r = 0.078. Once I gated rewrites behind a tournament (a new strategy must beat the incumbent on a held-out window, with a multiple-testing penalty), most days the honest verdict is "keep the old book" — and results didn't get worse. The learning is front-loaded.
2. Paper-to-live slippage was 4x my modeled cost. I mirror one lane into a small real-money account. Across 16 real round trips in one session: mean -0.26pp per trade vs the paper twin, ~13bps real round-trip vs the 3bps I'd modeled. Paper was breakeven that day; the real account lost money. For high-churn strategies that gap IS the strategy.
3. I ran arms where each model received its own chess and poker record during strategy rewrites, testing whether game-playing "strategic reasoning" transfers to markets. The no-games control beat both game-trained arms by 6-10pp. Not detected.
Honest caveats: one 3-week window, an up-tape that flatters an always-long rulebook, paper fills on the four AI accounts, n=4 models. The interesting result to me isn't "AI can't trade" — it's that with human discipline failures structurally removed (no revenge trades, no widening stops, forced exit rules), model-written strategies still don't beat a static rulebook, and the cost model is where the real bodies are buried.
Everything is public — every trade from all five accounts, losses included, no signup to watch. Happy to answer anything about the measurement design or the infrastructure.
At first, I thought that websites designed by an AI were just tacky.
A couple of months later, I was starting to really dislike them.
Now I'm at the point where I just hit the back button.
I’m at step 0—asking “This site is awful, is this AI?”
I should have gleamed it from the visual design language, but it was the writing that clued me in.
Author here. Since late July, GPT-5.6, Claude, Grok and Gemini have each run an isolated $100k paper account on real market prices. Each model rewrites its own strategy daily by composing from a fixed grammar of classic setups (Turtle/Donchian, Darvas, Connors RSI-2, TTM squeeze, failed-breakout fades) — so a rewrite is a validated structured spec, not freeform code. A fifth account runs a frozen rulebook as the control. After three weeks the frozen rulebook is +15.6%, the best model +5.7%, S&P +5.1%.
Three things I measured that I didn't expect:
1. Daily self-rewriting adds almost nothing. Correlation between rewrite count and performance across arms: r = 0.078. Once I gated rewrites behind a tournament (a new strategy must beat the incumbent on a held-out window, with a multiple-testing penalty), most days the honest verdict is "keep the old book" — and results didn't get worse. The learning is front-loaded.
2. Paper-to-live slippage was 4x my modeled cost. I mirror one lane into a small real-money account. Across 16 real round trips in one session: mean -0.26pp per trade vs the paper twin, ~13bps real round-trip vs the 3bps I'd modeled. Paper was breakeven that day; the real account lost money. For high-churn strategies that gap IS the strategy.
3. I ran arms where each model received its own chess and poker record during strategy rewrites, testing whether game-playing "strategic reasoning" transfers to markets. The no-games control beat both game-trained arms by 6-10pp. Not detected.
Honest caveats: one 3-week window, an up-tape that flatters an always-long rulebook, paper fills on the four AI accounts, n=4 models. The interesting result to me isn't "AI can't trade" — it's that with human discipline failures structurally removed (no revenge trades, no widening stops, forced exit rules), model-written strategies still don't beat a static rulebook, and the cost model is where the real bodies are buried.
Everything is public — every trade from all five accounts, losses included, no signup to watch. Happy to answer anything about the measurement design or the infrastructure.
[dead]