Description
fal is hiring an experienced Software Engineer to build and evolve its core Python/Rust platform for request routing, AI workload orchestration, scheduling, GPU autoscaling, file storage, queueing, and related infrastructure. The role involves designing systems for 100x traffic growth, using AI to automate complex systems, profiling CPU and memory performance, and improving reliability and scalability. Candidates should have at least five years of experience building distributed compute and orchestration platforms, strong distributed systems knowledge, production-scale system experience, observability expertise, and familiarity with AI/ML infrastructure, high-performance systems programming, multi-tenant platforms, networking, and GPU workloads.
