Skip to main content

Model Efficiency Engineer at Cohere

Setup
Remote, Hybrid
Location
New York, New York
Type
Full-time
Posted

Description

Cohere is seeking a highly skilled and experienced engineer to join their fast-growing Model Efficiency team. This team focuses on building reliable ML systems and improving LLM inference efficiency by developing techniques to optimize model execution in production. The role involves working across the inference stack to identify and resolve performance bottlenecks, collaborating with modeling and systems teams, and contributing to the acceleration of inference. The company embraces a remote-friendly environment with offices in major global cities, and the Model Efficiency team is concentrated in the EST and PST time zones. Cohere offers a diverse and inclusive culture with various perks, including health and dental benefits, parental leave, and personal enrichment benefits.

For job seekers

Ready to find a role that actually fits?

Upload your résumé, start a Job Search Thread, and let Metaintro rank real openings against your experience — then guide you from search to offer.

Match

Compare live roles against your current evidence.

Position

Turn proof projects into role-specific applications.

Improve

Use market feedback to keep the skill plan current.

Return to navigation