Ashby · 15h ago
Inference Engineer
Together AI·Infrastructure
Published band
$180K–$260K + eq.
Above the median for this category.
- Location
- US remote
- Visa
- Considered
- HQ
- San Francisco · 150–250
- Source
- Ashby
Inference, fine-tuning, and GPU clusters for labs that want open weights at frontier speed. Remote US with research hubs.
Together is where open models become a business. Inference quality is the moat.
You will live in kernels, schedulers, and the last 15% of utilization.
What they ask
- Inference serving (vLLM, TensorRT-LLM, or similar)
- Comfortable reading CUDA or Triton
- Measurement culture
Nice if you have it
- Speculative decoding research
- MoE serving
- vLLM
- Kernels
- CUDA
- Serving