Agents & InferencearXiv

Format Sensitivity Index: Token-Controlled Prompt Wrapper Robustness and Schema Compliance in LLM Benchmarking

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

A single formatting change in prompt wrappers can swing LLM accuracy by up to 30x—even when token count is held constant. This means your production agents may silently fail schema compliance or drop performance when switching templates, wrappers, or even minor formatting tweaks, forcing you to either lock in brittle prompt structures or add costly compliance checks and variance reporting to every eval.

LinkedIn

Two AI summaries of each story, blind-voted — see today's agents & inference digest →