Agents & InferenceHugging Face

Your Agent Aced the Task. Will It Do It Again?

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

GPT-4.1 ReAct hit 77.4% Mean@5 on AppWorld, but only 53.0% of tasks passed all five repeated runs—a 24.4-point repeatability gap hidden by standard averages. If you ship agents, add Pass^k-style consistency evaluation to CI and treat one-off success as insufficient; trajectory-derived consistency guidelines can target nondeterministic failure modes without optimizing only for headline accuracy.