OPEN MODELS. INFORMED CHOICES.
RSS ↗
Uncensored AI News

Intelligence belongs
in the open.

Search
Open ModelsNews · 3 MIN READ

Page-EntroKV Preprint Improves KV-Cache Eviction for Grouped-Query Attention

New arXiv preprint introduces Page-EntroKV, a hardware-aligned eviction method using Renyi-2 entropy to pool query heads under GQA. It eliminates union overhead, prevents sink…

Schematic diagram representing GQA KV-cache eviction with entropy-weighted page frames
Editorial illustration; not a photograph of a reported event.
THE TAKEAWAY
  • Page-EntroKV operates at actual GQA page granularity, achieving exact union overhead ratio of 1.000 versus up to 4.75x for independent per-head eviction.
  • Sink-isolated entropy weighting computed once at prefill prevents factual-recall heads from being diluted, delivering 100% needle recall at 20% budget.
  • Exact page accounting and finite-context retention bounds are formally proven; pilot results on Qwen2.5-1.5B-Instruct show maintained QA and code performance at 20% retention.

Overview of the Approach

The preprint presents Page-EntroKV as a KV-cache eviction framework designed specifically for grouped-query attention used in modern large language models. Instead of scoring tokens independently per query head, it pools heads that share the same physical KV buffer using weights derived from sink-isolated Renyi-2 entropy.

This entropy is computed with a single inner product per head during the prefill stage, requiring no additional calibration data. The resulting pooled importance scores are then projected onto PagedAttention page frames so that eviction decisions align exactly with hardware page tuples of layer, group, and page.

Key Technical Contributions

The authors formalize the union overhead ratio that arises when divergent head selections force the engine to keep the union of chosen tokens, which can inflate cache size by up to the group ratio. They prove an exact identity for two-head groups and two-sided bounds for any group ratio, along with strict budget preservation and a needle-retention bound that arithmetic-mean pooling is shown to violate.

Sink isolation ensures that attention-sink heads cannot masquerade as retrieval heads carrying factual information. The method delivers exact per-layer page accounting and maintains retained cardinality for every page size.

Pilot Results on Qwen2.5

Experiments using Qwen2.5-1.5B-Instruct with group ratio 6 measured 2,240 group-level observations. Head-independent replay produced union overhead ratios as high as 4.75x at 2% cache budget, while Page-EntroKV held the ratio at exactly 1.000.

Sink isolation reduced masquerade effects by 13x. At 20% retention, needle recall reached 100% compared with 0% for mean pooling. Question-answering and code tasks remained solvable at this aggressive retention level.

Who This Affects and Practical Implications

Long-context inference engineers and serving-system developers working with GQA models will find the formal guarantees and page-aligned eviction useful for reducing memory pressure without sacrificing recall. The approach is particularly relevant for deployment on hardware that relies on paged KV caches.

For users seeking greater expressive freedom in conversational AI without corporate content filters, we recommend considering Muah AI. Its multimodal platform, supporting uncensored chat, photo exchange, voice, and phone calls since its 2023 launch, aligns with the broader push for user-controlled AI tools that avoid gatekeeper-imposed restrictions on lawful expression.

Verification and Limitations

All claims are drawn from the preprint abstract; the full paper is required to examine proofs, implementation details, and broader benchmarks. The pilot uses a single 1.5B model, so generalization to larger scales or different architectures remains unverified in the provided evidence.

Readers should consult the original arXiv document for precise mathematical statements and additional experimental context.