I make my agent prove every fix by recording a before and after.
This is a vibe coding post classified by Jev as Code review & safety (a workflow), kept by the Vibe Coding Radar because it carries real work, not commentary.
I make my agent prove every fix by recording a before and after. Every pull request has both. One page took 850 milliseconds to load. The agent got it down to 61. It wrote the test, measured both numbers, and put them in the pull request. Sometimes the agent watches its own after-video and finds that the feature is not complete. Then it goes back to build. No one tells it to. You pay for this with setup. Write the two skills first: evidence-driven testing, and before and after. Then review the screenshots and merge.
Posted by The Startup Ideas Podcast (SIP) 🧃 (47.9k followers) 4 days ago · 30 likes · 3.5k views · view the original post on X. Kept by the Vibe Coding Radar as Code review & safety.
More vibe coding work like this
- Existing coding benchmarks stop at the first working version. We are releasing Vibe Code… — @ValsAI
- Introducing Scorers: Agents that grade your agents. — @warpdotdev
- Hot tip: you can ask your agent to verify its own work and catch issues. The trick is to… — @MiaAI_lab
- 一个需求下去,Agent 一口气动了十几个文件,改动一行行翻过去,到底牵连了哪几个模块看不出来,合并了才发现碰到别处。 — @GitHub_Daily
- GPT-6 Astra helps @cognition’s Devin back up “it works” with tests before the team ships. — @OpenAIDevs
- You ask Codex to build a page. — @godofprompt
- At Anthropic, Claude now writes 80% of our code. Engineers ship 8x more code per quarter. — @addyosmani
- GPT-6 Astra in Codex is helping test code end to end at @perplexity_ai. — @OpenAIDevs
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 1.8k posts from 1.8k X accounts over the last 14 days, 141 tools, 12 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-19 01:18 UTC. Full method.