Description
Tower Research Capital is hiring a Machine Learning Training Systems Engineer to build and optimize systems for training machine learning models at scale. The role focuses on benchmarking heterogeneous compute platforms, optimizing distributed training across CPUs, GPUs, and accelerators, improving GPU kernel and framework performance, and designing efficient training infrastructure for on-premises and cloud environments. The position requires at least three years of experience with large-scale machine learning training workloads and strong knowledge of Python, C++, GPU technologies, distributed-training systems, and performance-analysis tools.
