Same GPU, Same Impossible Job: One Engine Crashed, the Other Slowed Down 8%
I gave two inference engines the same job on the same 8 GB GPU: 64 requests, each with a 2,000-word prompt. Neither had enough memory for it — the job needed about a third more cache than the card cou
seng-wei-chieh.hashnode.dev9 min read