Description
Inference.net is seeking a highly skilled and experienced Machine Learning Systems Engineer to optimize their AI inference stack. The successful candidate will be responsible for making inference as fast and efficient as possible, working across the full inference stack from CUDA kernels to serving frameworks. This role involves implementing and productionizing optimization techniques, deep diving into inference frameworks, profiling and optimizing GPU utilization, and experimenting with novel inference techniques. The position reports directly to the founding team and offers significant autonomy and resources to push the boundaries of model serving performance.
