Member of Technical Staff, ML Systems
Own speed and efficiency across the stack—from low-level GPU kernels to distributed inference engines and multi-node training/serving systems.
What you’ll do
- Optimize system and GPU performance of training and inference for image, video, and world-model workloads
- Profile and remove bottlenecks at kernel, memory, system, and cluster levels (Nsight and related tooling)
- Implement low-level optimizations in CUDA / Triton
- Design and improve distributed inference / training engines for diffusion models
- Own communication performance across GPUs and nodes: NCCL, RDMA / InfiniBand / RoCE, and disaggregated serving
- Build benchmarking and regression harnesses so performance gains stick in production
- Hardware-aware kernel, runtime, and model co-design
Minimum qualifications
- 3+ years in deep learning inference/training systems, distributed systems, or high-performance computing
- Proficiency in CUDA, with hands-on GPU profiling (e.g. Nsight)
- Strong systems instincts across GPU architecture, parallel programming, and compute kernels
- Experience with distributed multi-GPU debugging and optimization
- Familiarity with PyTorch and performance-critical model execution
- Comfort working onsite with the team in the Bay Area
Preferred qualifications
- Experience optimizing diffusion, video, VLM, or other multimodal models for training and inference
- Deep familiarity with modern inference / runtime stacks (vLLM, SGLang, TensorRT-LLM, or equivalent)
- Experience with NCCL, RDMA, InfiniBand/RoCE, or disaggregated serving
- Knowledge of ML compilers and graphs (Triton, torch.compile, XLA, TensorRT)
- Contributions to open-source ML systems
- Background co-designing models and runtimes for emerging accelerators
To apply, email careers@tensorscale.io.