Description
Cohere is seeking a highly skilled and experienced engineer to join their fast-growing Model Efficiency team. This team focuses on building reliable ML systems and improving LLM inference efficiency by developing techniques to optimize model execution in production. The role involves working across the inference stack to identify and resolve performance bottlenecks, collaborating with modeling and systems teams, and contributing to the acceleration of inference. The company embraces a remote-friendly environment with offices in major global cities, and the Model Efficiency team is concentrated in the EST and PST time zones. Cohere offers a diverse and inclusive culture with various perks, including health and dental benefits, parental leave, and personal enrichment benefits.

