Description
The role leads the architectural definition, modeling, specification, and lifecycle management of high-performance machine-learning compute IP and acceleration blocks for Google Cloud’s TPU silicon. It involves defining compute dataflows, numerical formats, sparsity, KV-cache optimization, and performance, latency, power, and area projections; partnering with AI research and software compiler teams; and applying hardware-software co-design, computer architecture, and ML hardware systems expertise. The position requires a bachelor’s degree and 15 years of relevant experience, with preferred qualifications including a master’s or PhD, extensive AI accelerator leadership, and knowledge of modern deep-learning workloads and memory subsystems. The role is available in Tel Aviv and Haifa, Israel.
