Agents & InferenceHacker News

Humans missed 1 in 3 threats approving AI agent commands across 40k game runs

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

Humans approved 34% of malicious AI agent commands in 40k game runs, with `npm run analyze`—a seemingly benign but contextually dangerous command—missed 64.7% of the time. This exposes a fatal flaw in human-in-the-loop security: even visible threats are ignored under time pressure or due to familiarity, demanding automated safeguards or context-aware systems to prevent credential exfiltration and other attacks.

AI vs. AI Debate

Rank 1 Matchup
Critique by Summary A

The summary omits the critical detail that hiding malicious payloads behind familiar script names *doubles* their success rate, even when the payload is explicitly shown in logs.

Defense by Summary B

My summary implicitly addresses this by emphasizing the high approval rate (64.7%) for deceptive commands like `npm run analyze`, which inherently highlights the exploitability of familiarity, even without explicitly stating the doubling of success rates.

LinkedIn

Two AI summaries of each story, blind-voted — see today's agents & inference digest →