Skip to content
Go Reach HQ
All opportunities

Machine Learning

Machine Learning Systems Engineer

Location
Palo Alto, CA
Compensation
$200,000 - $300,000
Employment type
Permanent
  • On-site

Our client builds diffusion-based large language models (dLLMs) that generate entire responses in parallel instead of token-by-token — up to 10x faster and more efficient than traditional autoregressive LLMs, while matching best-in-class quality. They shipped the world's first commercially available diffusion LLM in early 2025 and are now deploying large-scale diffusion models at Fortune 500 companies. This is a 20-person team based in Palo Alto that has raised $56M to date.

Role responsibilities

You'll join a small, elite team of researchers and engineers building the infrastructure that trains and serves frontier diffusion language models.

  • Build and scale infrastructure for training and inference of diffusion-based LLMs (dLLMs)
  • Work across the stack — from low-level GPU/CUDA kernel optimization to distributed training orchestration
  • Partner directly with researchers and founders on core model performance breakthroughs
  • Help deploy large-scale diffusion LLMs for enterprise customers

Requirements

  • 2-5 years of experience in ML systems engineering, with a focus on infrastructure for training and inference systems
  • Comfort working across the stack, from CUDA/kernel-level optimization to distributed training systems
  • Experience with high-performance inference frameworks (vLLM, TensorRT, ONNX Runtime, SGLang, or similar) is a strong plus
  • Excited to work in a small, high-ownership, in-person team

Tech stack: vLLM, TensorRT, ONNX Runtime, PyTorch, TensorFlow, CUDA, Docker, Kubernetes, Python, AWS, Azure, Kubeflow, SGLang

Compensation: $200K–$300K base plus competitive equity. Visa sponsorship, including H-1B, available.

Work policy: 5 days/week in-office in Palo Alto, CA.

Interview process: intro call, technical coding rounds, then a panel with the founders and reference checks.