Description
Arcana is hiring an Applied AI Engineer focused on inference optimization and agent systems. The role involves improving time-to-first-token and streaming performance, designing parallel Plan-Execute-Synthesize agent pipelines, building reliable Temporal orchestration, enforcing structured outputs, routing models across providers, and maintaining evaluation and regression testing infrastructure. The engineer will also work on model serving, cold-start optimization, async workers, and observability. The posting seeks candidates who have shipped production inference systems or eval harnesses and lists Go, Python, Temporal, Kafka, PostgreSQL, and Docker as relevant technologies.
