Claude Code self-improved on business workflows through conversation, not memories
I instructed it how to pull task descriptions (spreadsheets, CRM records, email, ...) from a simulated business workflow server and told it to get better
held-out success rose from 34% to 44%
Claude Code and Codex can improve agent prompts and perform engineering practices like iteratively evaluating their changes
but what happens without training data?
surprisingly, Claude Code’s improvements performed as well without training data as with it in several settings
I was curious what Codex and Claude Code actually do as agent engineers under light instruction
they perform engineering practices like failure mode analysis and ad-hoc evaluations, even without explicit access to such tools