Vibe Coding Radar
Support
LiveUpdated 2026-09-19 01:18 UTC

Introducing DuetBench-2, our latest benchmark for measuring whether self improving…

Introducing DuetBench-2, our latest benchmark for measuring whether self improving agents create improvements that…

This is a vibe coding post classified by Jev as Other (a launch), kept by the Vibe Coding Radar because it carries real work, not commentary.

Introducing DuetBench-2, our latest benchmark for measuring whether self improving agents create improvements that last. After passing 93% of diagnostic tasks and outperforming humans on complex agent building tasks, Duet quickly outgrew the original benchmark. 🧵

Posted by Decagon (6.1k followers) 9 days ago · 35 likes · 3k views · view the original post on X. Kept by the Vibe Coding Radar as Other.

More vibe coding work like this

Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 1.8k posts from 1.8k X accounts over the last 14 days, 141 tools, 12 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-19 01:18 UTC. Full method.