Introducing Scorers: Agents that grade your agents.
This is a vibe coding post classified by Jev as Code review & safety (a launch), kept by the Vibe Coding Radar because it carries real work, not commentary.
Introducing Scorers: Agents that grade your agents. Use LLM-as-a-judge to grade past coding agent sessions on: • Quality • Efficiency • Compliance • Or any custom dimension Scores feed into performance measurements and automatic self-improvement for software factories
Posted by Warp (58.9k followers) 2 days ago · 165 likes · 109.2k views · view the original post on X. Kept by the Vibe Coding Radar as Code review & safety.
More vibe coding work like this
- Existing coding benchmarks stop at the first working version. We are releasing Vibe Code… — @ValsAI
- Hot tip: you can ask your agent to verify its own work and catch issues. The trick is to… — @MiaAI_lab
- 一个需求下去,Agent 一口气动了十几个文件,改动一行行翻过去,到底牵连了哪几个模块看不出来,合并了才发现碰到别处。 — @GitHub_Daily
- GPT-6 Astra helps @cognition’s Devin back up “it works” with tests before the team ships. — @OpenAIDevs
- I make my agent prove every fix by recording a before and after. — @startupideaspod
- You ask Codex to build a page. — @godofprompt
- At Anthropic, Claude now writes 80% of our code. Engineers ship 8x more code per quarter. — @addyosmani
- GPT-6 Astra in Codex is helping test code end to end at @perplexity_ai. — @OpenAIDevs
Every post is read and classified by Jev (TypeSafe): what it is, which market it belongs to, and whether the link is a real tool. 1.8k posts from 1.8k X accounts over the last 14 days, 141 tools, 12 markets. Collected every 5 minutes, fully re-ranked every hour — last update 2026-09-19 01:18 UTC. Full method.