Agents & InferencearXiv

CaptionQA outperforms VideoQA on 10 of 12 models for long egocentric videos

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

On egocentric videos longer than 20 minutes, caption-based memory beat direct video QA in 10/12 models with 30-second caption windows and kept a 3.22-point mean accuracy gain in matched-frame controls. For production wearable/agent systems, precomputing dense text captions as episodic memory is a practical way to reduce visual-token pressure and improve long-horizon recall, especially when paired with retrieve-and-verify for another accuracy boost of up to 5.3 points.

LinkedIn

Two AI summaries of each story, blind-voted — see today's agents & inference digest →