AI Research
Member of Technical Staff — Machine Learning
- Location
- San Francisco, CA
- Compensation
- $200,000 - $350,000
- Employment type
- Permanent
- On-site
Our client is developing AI foundation models efficient enough to run entirely on a phone, with the goal of matching the performance of top labs like OpenAI or Google in a dramatically smaller footprint. They're a 5-person team based in San Francisco's Financial District, founded in 2025, with $9M raised to date.
Role responsibilities
We're looking for self-motivated researchers and engineers who want to meaningfully contribute to training powerful, efficient models — whether that means low-level GPU optimization or new optimization theory.
Relevant experience includes any of the following:
- Work on large language models at a research lab (e.g. OpenAI, Google, Mistral, Z.ai, Qwen, DeepSeek, Ai2) or in academia
- Pretraining language models and large-scale AI infrastructure — model parallelism, software/hardware co-design for training throughput, monitoring large training runs
- Post-training language models — reinforcement learning for LLMs (environments, infrastructure, training), instruction data curation, or tool use
- Inference/systems optimization — contributions to vLLM, SGLang, Dynamo, MegaKernels, or similar; systems-level understanding of batching, KV cache pressure, and long context
- Low-level kernel design — CUDA, C++, CuTe, Triton, PTX, TileLang, or similar DSLs
Requirements
- 1+ years of experience in theoretical LLM research or as an ML research engineer at a top-tier tech company
- Academic papers, community contributions (e.g. NanoGPT Speedrun, Marin), or equivalent hands-on experience are all valued
Compensation: $200K–$350K plus competitive equity. Visa sponsorship, including H-1B, available.
Work policy: 5 days/week in-office in San Francisco's Financial District.
Interview process: initial interview, technical interview, take-home or work trial.
