Description
G-Research is hiring an ML Performance Engineer in London to optimise large-scale machine learning training and inference workloads across distributed GPU, CPU, and memory-intensive infrastructure. The role involves profiling, benchmarking, tuning, and developing reference implementations, libraries, and tools while collaborating with research, infrastructure, systems, architecture, and platform teams to improve the compute stack and guide long-term platform decisions. Candidates need a computer science degree or equivalent experience, distributed workload optimisation experience, Python, CUDA, HPC schedulers, Kubernetes, deep learning frameworks such as PyTorch, parallel programming, Linux systems expertise, and performance profiling tools. Benefits include competitive compensation, bonus, healthcare and life assurance, pension contributions, annual leave, meals, and other workplace benefits.
