ResKV: Recovering Evicted Token Contributions via Residual KV Cache — LongBench 32/32
Long-context LLM inference has a memory problem that compounds with scale. Every additional token grows the KV cache linearly — a single 32K-token forward pass through an 8B model can exhaust several
miainflorence.hashnode.dev8 min read