Summary from listing
The role focuses on building and optimizing infrastructure for large-scale machine learning and deep learning projects, including LLMs, GPUs, cloud platforms, containerization, distributed training, and MLOps. Responsibilities include managing and optimizing resources, troubleshooting hardware and memory issues, supporting data scientists and machine learning engineers, and leading technical aspects of projects while ensuring timely and error-free deliverables.
