mempool.asia
Amber Group · HIRING
Job Description

Model Inference
Engineer.

Amber Group·Hong Kong · Shenzhen · Singapore
We are looking for a passionate Model Inference Engineer to build and maintain high-performance inference service infrastructure. You will own the full technical chain — from GPU cluster deployment and model optimization through to serving stable, efficient inference externally. The ideal candidate pairs deep technical grounding with commercial thinking, finding the best balance between performance, cost and user experience.
Experience
3 yrs+
AI infra / model deployment
Education
Bachelor’s or above
CS · AI · Electronics
Location
Hong Kong / Shenzhen
Singapore — any
01

Key Responsibilities

1Inference infrastructure build & maintenance
  • Design, deploy and maintain GPU-based inference clusters (multi-GPU, multi-node)
  • Own deployment, version management and lifecycle maintenance of open-source and in-house small models
  • Build highly available, scalable inference service architecture, ensuring stability (SLA ≥ 99.9%)
2Inference performance optimization
  • Accelerate and tune inference for LLMs and multimodal models (e.g. MiniMax H3)
  • Apply quantization (INT8/INT4), operator fusion and graph optimization to cut latency and raise throughput
  • Optimize VRAM management strategy to maximize GPU utilization
  • Deliver scenario-specific optimization, balancing inference speed, accuracy and cost
3Distributed deployment & network optimization
  • Implement parallel inference deployment across multi-GPU, multi-node environments
  • Optimize inter-node communication; familiarity with RDMA, InfiniBand and other high-speed networking
  • Solve load balancing, data parallelism and pipeline parallelism problems in distributed inference
4Cost control & commercial support
  • Monitor and analyze resource consumption; continuously optimize cost per request
  • Build cost accounting models to support product pricing and commercial decisions
  • Evaluate cost-performance of hardware options (GPU models, cloud vs self-hosted)
5Technical service & collaboration
  • Provide stable, reliable inference APIs for internal products and external customers
  • Partner with algorithm teams to turn trained models into production-grade inference services
  • Write technical documentation; establish monitoring and troubleshooting processes
02

Requirements · Must-have

  • Education: Bachelor’s degree or above in Computer Science, AI, Electronics or a related field
  • Experience: 3+ years in AI infrastructure or model deployment
  • Core technical skills:
    • Mastery of at least one major inference framework (TensorRT, vLLM, TGI, Triton Inference Server, DeepSpeed-Inference)
    • Proficient with PyTorch/TensorFlow model export and deployment workflows
    • Deep understanding of GPU architecture; CUDA programming or kernel optimization experience preferred
    • Familiar with Docker, Kubernetes and containerized deployment
    • Capable in Linux system operations and network debugging
03

Nice to have

  • Production deployment experience with LLMs or multimodal models (VLM)
  • Familiarity with training workflows and the pain points across the training-to-inference chain
  • Knowledge of distributed training frameworks (DeepSpeed, Megatron-LM, FSDP)
  • Hands-on experience with model compression — quantization, pruning, distillation
  • Familiar with cloud GPU instance selection and optimization (AWS, Azure, Alibaba Cloud)
  • Experience building an inference serving platform from scratch
04

Soft skills

Commercial thinking

Cost-aware; frames technical choices in business terms

Problem solving

Strong at locating and resolving complex systemic issues

Appetite for learning

Deeply curious about AI; tracks the frontier proactively

Communication

Explains technical plans clearly; collaborates across teams

Results oriented

Focused on business value rather than technical perfection alone

Follow
@MempoolAsia
on X
Amber Group
Listed on mempool.asia · Jobs
Role & apply
Export view2x1080px wide long-form · opens the bare artboard — screenshot it, or use 2x for a 2160px capture