Search posts, tags, users, and pages
Partha
#AI #agents #langgraph, #mcp, #rag, #graphrag #llm #engineering-management
TL;DR: LLM inference bottlenecks are usually memory, not compute. KV cache eliminates redundant attention computation. Paged attention eliminates memory fragmentation — cutting waste from ~60–80% down
No responses yet.