Summary from listing
Dragonfly is sourcing a Senior Inference Optimization Engineer for a portfolio company building privacy-first consumer AI infrastructure. The role focuses on improving LLM inference throughput, latency, and cost per token at scale, including GPU infrastructure, batching, attention, KV cache, speculative decoding, quantization, distributed inference, profiling, and load balancing. The position is remote in the United States, with compensation of $180,000–$250,000 for senior-level candidates.

