Senior Deep Learning Software Engineer, Inference
Design and optimize GPU-accelerated deep learning inference software for large-scale language and generative AI models. Contribute to high-performance frameworks like vLLM and SGLang, focusing on performance tuning across NVIDIA accelerators. Collaborate with research and engineering teams to implement cutting-edge algorithms and optimize model serving pipelines using CUDA, Triton, and other low-level tools.