Concept detail
KV cache size
0 tagged videos · 1 connection videos · degree 0
Publication timeline
- 2026-091 video
Source videos
Observed neighbors
GQA/MLA
Attention architecture change reduces the number/format of cached key-value tensors.
1 source videos
Deep dive on LLM Inference at Scale — Harshul Jain, Audible & Tanmay Sah, Independent AI ResearcherRelated tagged insights
No core insights are tagged with this concept.