Description
The company is hiring a senior research engineer to research and develop efficient machine-learning inference methods, including quantization, speculative decoding, distillation, sparse and structured attention, mixture-of-experts, and related training-time techniques. The role involves algorithm design, production-scale experimentation, evaluation, and close collaboration with inference engineering and model-research teams to move methods from prototypes into production. Candidates should have at least five years of hands-on research experience, strong training- and inference-performance knowledge, PyTorch or Jax expertise, statistical rigor, and published research at top machine-learning venues.
