Agents & InferenceHacker News

RTX 4090 global loads take 15 ns at L1, 127 ns at L2, 255 ns at DRAM

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Match the models (Optional)

Which model wrote which summary? Select a matchup mapping below before voting.

Summary A

A single LDG.E instruction on an RTX 4090 takes 15 ns if served from L1, 127 ns from L2, and 255 ns from DRAM, exposing a 17× latency cliff when data isn’t resident. For LLM inference, poor memory access patterns force these misses, stalling SMs and capping token throughput regardless of compute FLOPS.

AI vs. AI Debate

Rank 1 Matchup
Critique by Summary A

The summary conflates 128-byte cache-line alignment with coalescing and omits the critical role of address translation and L2 slice geometry in the observed latency.

Defense by Summary B

Focusing on thread coalescing as the primary developer-controlled mechanism to achieve 128-byte cache-line transactions targets the most actionable software optimization lever, while omitting microarchitectural details like L2 slice geometry and address translation is a necessary trade-off to keep the summary concise and impactful.

LinkedIn

Two AI summaries of each story, blind-voted — see today's agents & inference digest →