Ai Infrastructure Engineer - GPU & Colocation

Hồ Chí Minh
Thỏa thuận
Toàn thời gian
Hạn nộp hồ sơ: 18/10/2026 (Còn 27 ngày)
Ứng tuyển sớm để được ưu tiên
Kết nối với Nhà tuyển dụng để tìm hiểu thông tin và gia tăng cơ hội trúng tuyển
OMess
Nhà tuyển dụng đang online

Mô tả công việc

We are building our own GPU infrastructure to run AI models, support AI-driven software development, and power AI capabilities in our software products. We are looking for a hands-on engineer to build this infrastructure in a colocation data centre and take responsibility for its ongoing operation.
You will turn workload requirements into a practical infrastructure setup-from selecting servers and coordinating installation to configuring Linux, deploying model-serving environments, and keeping systems secure and reliable.
Your Responsibilities
Plan and Build the Infrastructure
• Translate AI workload requirements into GPU, compute, storage, and networking configurations.
• Evaluate hardware compatibility, supplier proposals, and costs.
• Coordinate procurement, delivery, installation, and commissioning.
• Confirm rack space, power, cooling, connectivity, and access requirements with the colocation provider.
• Install and configure servers, networking equipment, cabling, and remote management.
• Maintain infrastructure documentation and asset inventories.
Configure GPU and AI Systems
• Configure and maintain Linux servers, GPU drivers, CUDA environments, and container runtimes.
• Deploy and maintain model-serving infrastructure in collaboration with the AI and engineering teams.
• Diagnose GPU, memory, storage, and network performance issues.
• Benchmark workloads and improve resource utilisation.
• Automate deployment and configuration to make environments reproducible.
Manage Networking and Security
• Configure network segmentation, firewalls, VPNs, and secure administrative access.
• Implement access controls, credential management, system hardening, and patching.
• Separate workloads and environments according to product and security requirements.
• Troubleshoot connectivity issues with infrastructure providers.
Own Reliability and Operations
• Establish monitoring and alerting for hardware health, GPU utilisation, capacity, and service availability.
• Implement backup and recovery procedures and test restoration.
• Maintain operational runbooks and troubleshoot infrastructure incidents.
• Coordinate hardware replacements, maintenance windows, and remote-hands support.
• Identify single points of failure and propose practical improvements.
• Establish clear support coverage and escalation procedures with the team.
Manage Capacity and Costs
• Track utilisation, operating costs, and capacity constraints.
• Recommend upgrades based on measured demand and performance.
• Assess trade-offs between owned hardware, rented GPU capacity, and cloud services.
• Plan expansion while avoiding unnecessary complexity and overprovisioning.

Yêu cầu

What You Bring
• Hands-on experience deploying and operating physical servers in a data centre or colocation environment.
• Strong Linux administration and troubleshooting skills.
• Practical experience with GPU servers, NVIDIA drivers, CUDA, and AI workloads.
• Solid networking knowledge, including switching, routing, VLANs, firewalls, and VPNs.
• Experience with containers, monitoring, backups, and infrastructure automation.
• An understanding of rack power, cooling, hardware compatibility, and remote server management.
• A disciplined approach to security, documentation, maintenance, and recovery.
• The ability to work independently and coordinate with vendors and data centre providers.
• Professional English for technical documentation and collaboration.
• Willingness to perform on-site installation and maintenance when required.
Useful Additional Experience
• Operating AI inference and model-serving systems.
• Managing multi-GPU workloads and high-speed networking.
• Using infrastructure-as-code and configuration-management tools.
• Supporting environments with workload isolation and availability requirements.
How You Will Work
• You will own the implementation and reliable operation of our AI infrastructure, working closely with the AI and software engineering teams to align infrastructure with application and model requirements.
• This is a hands-on role covering both the initial build and ongoing operations, with direct involvement in model deployment, performance optimisation, capacity planning, and infrastructure decisions.
What Success Looks Like
• Infrastructure meets agreed workload, security, and operational requirements.
• AI workloads run reliably with measurable performance and utilisation.
• Monitoring, tested recovery procedures, and operational documentation are in place.
• Infrastructure costs are transparent and expansion decisions are based on evidence.
• Incidents are resolved effectively, with clear coordination across suppliers and service providers.

Quyền lợi

Thưởng
Theo quy định công ty

Thông tin khác

NGÀY ĐĂNG
18/09/2026
CẤP BẬC
Nhân viên
NGÀNH NGHỀ
Công Nghệ Thông Tin/Viễn Thông > System/Cloud/DevOps Engineer
KỸ NĂNG
Gpu Servers, Nvidia Drivers, Cuda, Linux System Administration, Switching, Routing, Vlans, Firewalls, Vpns, Containers, Monitoring, Backups, Automation, Work Independently
LĨNH VỰC
Phần Mềm CNTT/Dịch vụ Phần mềm
NGÔN NGỮ TRÌNH BÀY HỒ SƠ
Tiếng Anh
SỐ NĂM KINH NGHIỆM TỐI THIỂU
Không yêu cầu
QUỐC TỊCH
Không hiển thị
Xem thêm

Thông tin chung

  • Thu nhập: Thương lượng

Nơi làm việc

  • 152 Võ Văn Kiệt, Bến Thành, Hồ Chí Minh

Công ty TNHH Sapawoo

Xem trang công ty
Địa chỉ công ty: 160 Lê Thánh Tôn
Quy mô: Từ 10 - 25 nhân viên
Thông tin công việc
Vị trí:
Nhân viên
Hình thức làm việc:
Toàn thời gian
Việc làm tương tự
WORKIT tuyển dụng Chuyên viên lập trình AI
CÔNG TY CỔ PHẦN WORKIT
Hồ Chí Minh
CÔNG TY CỔ PHẦN WORKIT
thỏa thuận
AI Engineer Intern
Innotech Vietnam Corporation
Hồ Chí Minh, Phú Yên
Innotech Vietnam Corporation
You'll love it
QA Engineer Python, Automation, AI, ML
Motorola Solutions
Hồ Chí Minh
Motorola Solutions
You'll love it
Middle/Senior Developer (NodeJS, ReactJS, AI, Leadership)
Công Ty TNHH LOTUS APP
Hồ Chí Minh
Công Ty TNHH LOTUS APP
Thỏa thuận
[HCM] Công Ty UP4DISTRIBUTION Tuyển Dụng Nhân Viên AI Agent Developer, Project Manager Full-time 2026
CÔNG TY TNHH UP4DISTRIBUTION
Hồ Chí Minh
CÔNG TY TNHH UP4DISTRIBUTION
Thỏa thuận
Cảnh báo dấu hiệu lừa đảo tuyển dụng
Thu phí và cung cấp thông tin
  • Phí hồ sơ, đồng phục, đặt cọc.
  • Yêu cầu nộp bản gốc giấy tờ.
  • Cung cấp mã OTP.
Hứa hẹn trúng tuyển 100%
  • Không yêu cầu trình độ.
  • Không cần thử việc.
Phỏng vấn bất thường
  • Địa điểm xa văn phòng công ty.
  • Phỏng vấn qua Telegram.
Yêu cầu làm nhiệm vụ
  • Tải app, nạp tiền.
  • Làm nhiệm vụ nhận thưởng.
Tin tuyển dụng sơ sài
  • Mô tả công việc chung chung
  • Nhiệm vụ đơn giản, thu nhập khủng
  • Lỗi chính tả, đánh máy.
Đội ngũ hỗ trợ của JobOKO sẵn sàng đồng hành, tư vấn và giới thiệu những cơ hội việc làm phù hợp, giúp Ứng viên tự tin phát triển sự nghiệp và chinh phục mục tiêu nghề nghiệp bền vững.
Đội ngũ hỗ trợ của JobOKO luôn chủ động tư vấn các giải pháp tuyển dụng tối ưu, cam kết đồng hành và hỗ trợ Quý Nhà tuyển dụng đạt được hiệu quả tuyển dụng bền vững.