Ashby · 15h ago

Inference Engineer

Together AI·Infrastructure

Published band

$180K–$260K + eq.

Above the median for this category.

Location
US remote
Visa
Considered
HQ
San Francisco · 150–250
Source
Ashby

Inference, fine-tuning, and GPU clusters for labs that want open weights at frontier speed. Remote US with research hubs.

Together is where open models become a business. Inference quality is the moat.

You will live in kernels, schedulers, and the last 15% of utilization.

What they ask

  • Inference serving (vLLM, TensorRT-LLM, or similar)
  • Comfortable reading CUDA or Triton
  • Measurement culture

Nice if you have it

  • Speculative decoding research
  • MoE serving
  • vLLM
  • Kernels
  • CUDA
  • Serving
Apply on source