I build tools for coding agents, then run controlled experiments to check whether they work.
emulo mines your Claude Code and Codex session logs into a
you.md profile an agent reads before a task, so it behaves like something that already knows how
you work. 275 stars.
VibeRaven is a cockpit for AI coding agents: map the repo, control what an agent can touch, see what changed before you ship.
Most claims about agent tooling are untested, including mine. So in August I ran a pre-registered replication against my own product, with the decision rule fixed before a single run executed.
60 generations, three conditions: nothing appended, my mined profile, and a placebo, a length-matched profile of a designer who does not exist and who never mentions a single thing the experiment measures.
None of my three predictions separated the profile from the placebo. One separated in reverse of the direction I predicted, on my profile's own stated rule. The strongest number from the earlier run did not replicate.
I am publishing it anyway, including the reversal and the four reasons the earlier positive result was weaker than it looked. A tool that has never been tested against a control is a claim, not a result.
Recent open source: two fixes to Cabinet, an install failure on Windows that nobody in the tracker had named, and four shipped skill defaults broken by an upstream rename.
I also make cinematic motion pieces built from real data rather than mockups.



