Description
d-Matrix is hiring a Senior Staff ML Researcher for its Algorithms team to develop and evaluate methods that improve large language model inference on DNN accelerators. The role focuses on numerical precision, model compression, sparsity, pruning, distillation, low-rank approximation, transformer workloads, memory efficiency, and efficient execution, requiring research prototypes in Python and collaboration with hardware, compiler, and software teams. Candidates need a master's degree or PhD in a quantitative field, at least five years of relevant hands-on experience, strong machine-learning and applied-mathematics knowledge, and experience with frameworks such as PyTorch, JAX, or TensorFlow.
