Description
The Small Language Model Engineer will fine-tune and train small language models using Hugging Face, TRL, and adapter methods such as LoRA, QLoRA, and PEFT. The role also optimizes models for inference through quantization, pruning, and knowledge distillation; deploys them to edge devices, mobile platforms, and local servers; builds end-to-end MLOps pipelines; and monitors accuracy, latency, and hardware utilization in production. Deployment experience, ONNX knowledge, and MLOps tooling are beneficial.

