Agents & InferenceHacker News

Senior SWE-Bench: open-source benchmark that assesses agents as senior engineers

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

The best frontier agent (Claude Opus 4.8) hits only 24% "tasteful solve" on these under-specified, senior-level tasks — failing over 75% of the time when solutions must pass hidden behavioral tests, respect unstated codebase conventions, avoid bloat (<2x), and touch ~11 files across services. If your agents handle vague natural-language tickets rather than over-specified tasks, expect them to produce runtime-correct-but-tasteless code that violates load-bearing conventions, so keep human review gating anything multi-file or cross-service.

LinkedIn

Two AI summaries of each story, blind-voted — see today's agents & inference digest →