Machine Learning
Machine Learning Systems Engineer
- Location
- Palo Alto, CA
- Compensation
- $200,000 - $300,000
- Employment type
- Permanent
- On-site
Our client builds diffusion-based large language models (dLLMs) that generate entire responses in parallel instead of token-by-token — up to 10x faster and more efficient than traditional autoregressive LLMs, while matching best-in-class quality. They shipped the world's first commercially available diffusion LLM in early 2025 and are now deploying large-scale diffusion models at Fortune 500 companies. This is a 20-person team based in Palo Alto that has raised $56M to date.
Role responsibilities
You'll join a small, elite team of researchers and engineers building the infrastructure that trains and serves frontier diffusion language models.
- Build and scale infrastructure for training and inference of diffusion-based LLMs (dLLMs)
- Work across the stack — from low-level GPU/CUDA kernel optimization to distributed training orchestration
- Partner directly with researchers and founders on core model performance breakthroughs
- Help deploy large-scale diffusion LLMs for enterprise customers
Requirements
- 2-5 years of experience in ML systems engineering, with a focus on infrastructure for training and inference systems
- Comfort working across the stack, from CUDA/kernel-level optimization to distributed training systems
- Experience with high-performance inference frameworks (vLLM, TensorRT, ONNX Runtime, SGLang, or similar) is a strong plus
- Excited to work in a small, high-ownership, in-person team
Tech stack: vLLM, TensorRT, ONNX Runtime, PyTorch, TensorFlow, CUDA, Docker, Kubernetes, Python, AWS, Azure, Kubeflow, SGLang
Compensation: $200K–$300K base plus competitive equity. Visa sponsorship, including H-1B, available.
Work policy: 5 days/week in-office in Palo Alto, CA.
Interview process: intro call, technical coding rounds, then a panel with the founders and reference checks.
