Machine Learning Engineer — Inference Optimization
About the Role We’re looking for a Machine Learning Engineer to own and push the limits of model inference performance at scale . You’ll work at the intersection of research and production—turning cutting-edge models into fast, reliable, and cost-efficient systems that serve real users. This role is ideal for someone who enjoys deep technical work, profiling systems down to the kernel/GPU level, and translating research ideas into production-grade performance gains. What You’ll Do Optimize inference latency, throughput, and cost for large-scale ML models in production Profile and bottleneck GPU/CPU inference pipelines (memory, kernels, batching, IO) Implement and tune techniques such as:...
Apply at Featherless AISummary and match are AI-generated; verify details on the original posting.
Listed via Himalayas. All sources