Agents & InferencearXiv

Retrieval-augmented method boosts fallacy detection F1 to 0.864 in political debates

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

Steering RAG retrieval by argumentative relations (support/attack) rather than plain semantic similarity lifted macro-F1 to 0.864 for fallacy detection and 0.725 for classification, beating non-retrieval baselines across 42 configs and 14 models. The practical takeaway: for reasoning-quality tasks, what you retrieve should be conditioned on discourse structure between claims, not just surface text match—generic embedding retrieval leaves substantial accuracy on the table when the target is relational judgment.