Introducing DuetBench-2, our latest benchmark for measuring whether self improving…
This is a vibe coding post classified by Jev as Other (a launch), kept by the Vibe Coding Radar because it carries real work, not commentary.
Introducing DuetBench-2, our latest benchmark for measuring whether self improving agents create improvements that last. After passing 93% of diagnostic tasks and outperforming humans on complex agent building tasks, Duet quickly outgrew the original benchmark. 🧵
Posted by Decagon (6.1k followers) 9 days ago · 35 likes · 3k views · view the original post on X. Kept by the Vibe Coding Radar as Other.
More vibe coding work like this
- AI video creation+editing — @shushant_l
- Please watch. This is a fully custom wasm + wgpu custom engine I'm having a ton of fun… — @Dimillian
- Since I was a child, I have wanted my own Jarvis. — @startupideaspod
- Been quietly building our own open world action game for with loop of agents for a while… — @ziwenxu_
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 1.8k posts from 1.8k X accounts over the last 14 days, 141 tools, 12 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-19 01:18 UTC. Full method.