Description
Etched is seeking an ML Accelerator Performance Engineer Intern to design and develop performance analysis and profiling tools for its custom ML accelerators. The intern will collect and analyze hardware counters, execution traces, and memory behavior; trace host-side runtime and accelerator activity; correlate performance events across CPUs, accelerators, storage, networking, and distributed workloads; and build visualization and analysis tools. The role requires strong C++ or Rust programming, computer architecture knowledge, and experience or strong interest in low-level performance analysis, profiling, and optimization. The internship is fully in-person in San Jose, California.
