(Ycombinator) AI Inference Optimization: Multi-GPU Kernels, Efficiency, and Hardware Co-Design | OSMU Blog