Job Description
Model Inference
Engineer.
Amber Group·Hong Kong · Shenzhen · Singapore
We are looking for a passionate Model Inference Engineer to build and maintain high-performance inference service infrastructure. You will own the full technical chain — from GPU cluster deployment and model optimization through to serving stable, efficient inference externally. The ideal candidate pairs deep technical grounding with commercial thinking, finding the best balance between performance, cost and user experience.
- Experience
- 3 yrs+
AI infra / model deployment - Education
- Bachelor’s or above
CS · AI · Electronics - Location
- Hong Kong / Shenzhen
Singapore — any
02
Requirements · Must-have
- Education: Bachelor’s degree or above in Computer Science, AI, Electronics or a related field
- Experience: 3+ years in AI infrastructure or model deployment
- Core technical skills:
- Mastery of at least one major inference framework (TensorRT, vLLM, TGI, Triton Inference Server, DeepSpeed-Inference)
- Proficient with PyTorch/TensorFlow model export and deployment workflows
- Deep understanding of GPU architecture; CUDA programming or kernel optimization experience preferred
- Familiar with Docker, Kubernetes and containerized deployment
- Capable in Linux system operations and network debugging
Amber Group
Listed on mempool.asia · Jobs