Hardware
Dense, MoE, and the SSE Bug That Ate My Chat Stream
I had three vllm serve processes running on the H200 cluster, one prompt piped into all three through the same OpenAI-compatible client, and I was watching what should have...
Read Article
Four Sparks and a Closet
Part of my job is keeping an H200 cluster alive for research projects and workloads heavy enough to need it. Eight of those GPUs wired into one NVLink domain can...
Read Article