mempool.asia

AI Inference Engineer

Build high-performance inference services and find the best balance between performance, cost and user experience.

Hong Kong · Shenzhen · Singapore3+ years experience
Build efficient inference with us
Hong KongShenzhenSingaporeFull-time
Sign in to see the application email.

Overview

We are looking for a passionate AI Inference Engineer to build and maintain high-performance model-serving infrastructure. You will own the full technical chain — from GPU cluster deployment and model optimisation to delivering stable, efficient inference services externally. The ideal candidate combines deep technical skills with a commercial mindset, and can find the right balance between performance, cost and user experience.

Responsibilities

Inference infrastructure

  • Design, deploy and maintain GPU-based model inference clusters (multi-GPU, multi-node)
  • Own deployment, versioning and lifecycle management of open-source models and in-house small models
  • Build a highly available, scalable serving architecture and keep the service stable (SLA ≥ 99.9%)

Inference performance optimisation

  • Accelerate and tune inference for LLMs and multimodal models (e.g. MiniMax H3)
  • Apply quantisation (INT8/INT4), operator fusion, graph optimisation and similar techniques to cut latency and raise throughput
  • Optimise GPU memory management to maximise utilisation
  • Tailor optimisations to specific business scenarios, balancing speed, accuracy and cost

Distributed deployment & networking

  • Deploy models for parallel inference across multi-GPU, multi-node environments
  • Optimise inter-node communication; familiarity with RDMA, InfiniBand and other high-speed interconnects
  • Solve load balancing, data-parallel and pipeline-parallel problems in distributed inference

Cost control & commercial support

  • Monitor and analyse resource consumption and keep driving down cost per request
  • Build cost models that inform product pricing and commercial decisions
  • Evaluate the price/performance of hardware options (GPU SKUs, cloud vs self-hosted)

Technical service & collaboration

  • Provide a stable, reliable inference API for internal products and external customers
  • Work with the algorithm team to turn trained models into production-grade services efficiently
  • Write technical docs and set up monitoring and incident-response processes

Requirements

  • Bachelor's degree or above in computer science, AI, electronics or a related field
  • 3+ years in AI infrastructure or model deployment
  • Expert in at least one mainstream inference framework: TensorRT, vLLM, TGI, Triton Inference Server, DeepSpeed-Inference, etc.
  • Fluent in PyTorch / TensorFlow model export and deployment workflows
  • Deep understanding of GPU architecture; CUDA programming or kernel-optimisation experience preferred
  • Familiar with containerised deployment (Docker, Kubernetes)
  • Comfortable with Linux operations and network debugging

Nice to have

  • Production deployment experience with large language models (LLM) or multimodal models (VLM)
  • Familiar with model training and the pain points across the full train-to-serve pipeline
  • Knowledge of distributed training frameworks: DeepSpeed, Megatron-LM, FSDP, etc.
  • Hands-on experience with model compression: quantisation, pruning, distillation
  • Familiar with selecting and tuning cloud GPU instances (AWS, Azure, Alibaba Cloud, etc.)
  • Experience building an inference platform from scratch

What we value

  • Commercial mindset: cost-aware, able to weigh technical choices from a business angle
  • Problem solving: good at isolating and fixing complex systemic issues
  • Curiosity: strong interest in AI and proactive about following the frontier
  • Communication: explains technical proposals clearly and collaborates well across teams
  • Results-driven: focused on business value rather than technical perfection for its own sake

Apply

Hong Kong, Shenzhen or Singapore. Sign in to see the application email and send your CV straight to the hiring contact.

Sign in to see the application email.

← Back to rolesShare poster →mempool.asia / pool / roles
Amber GroupHong Kong · Shenzhen · Singapore