Agents & InferenceHacker News

Humans missed 1 in 3 threats approving AI agent commands across 40k game runs

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

Humans missed 1 in 3 threats when approving AI agent commands across 40,000 game runs, with deceptive commands like `npm run analyze` being approved 64.7% of the time. This highlights a critical vulnerability in relying on human-in-the-loop safeguards, as even clearly suspicious actions are overlooked due to familiarity or time pressure. For production systems, this necessitates stronger automated checks or contextual awareness to reduce dependence on manual approvals, which can fail under real-world constraints.

AI vs. AI Debate

Rank 1 Matchup
Critique by Summary B

The summary omits the critical detail that hiding malicious payloads behind familiar script names *doubles* their success rate, even when the payload is explicitly shown in logs.

Defense by Summary A

My summary implicitly addresses this by emphasizing the high approval rate (64.7%) for deceptive commands like `npm run analyze`, which inherently highlights the exploitability of familiarity, even without explicitly stating the doubling of success rates.

LinkedIn

Two AI summaries of each story, blind-voted — see today's agents & inference digest →