Senior Devops/mlops Engineer (5+ Years of Experience)
Hạn nộp hồ sơ: 23/10/2026 (Còn 29 ngày)
Ứng tuyển sớm để được ưu tiên
Kết nối với Nhà tuyển dụng để tìm hiểu thông tin và gia tăng cơ hội trúng tuyển
Nhà tuyển dụng đang online
Mô tả công việc
We are looking for Senior MLOps & Infrastructure Engineers to build and operate the hybrid AI computing platform behind VinFast's ADAS and autonomous driving programme - on-premise GPU clusters running perception model training and large-scale inference, a multi-petabyte sensor data archive, and the platform services used daily by our engineering and annotation teams.
This is a fullstack infrastructure role. You will work across compute, storage, networking, automation and security. We do not divide the team into narrow specialists: everyone owns a part of the platform end to end - provisioning, deployment, monitoring and incident response - and is expected to grow across the whole stack over time. Work is assigned according to what the programme needs at that moment, not by a fixed job boundary.
Much of the platform runs on-premise, including air-gapped segments. Experience operating without a public-cloud safety net is valued here
Key Responsibilities
1. AI Compute & Serving
- Operate and scale GPU clusters across training, inference and annotation workloads; manage scheduling, utilisation and capacity
- Deploy and scale high-throughput model serving; tune GPU memory and runtime performance
- Manage multi-GPU distributed training and reproducible environment sandboxing
- Scale pipeline orchestration on Kubernetes for large data processing jobs
2. Platform Automation & Delivery
- Drive GitOps-based deployment and maintain infrastructure-as-code across the platform
- Build and maintain CI pipelines, secure container builds and release automation
- Build monitoring, logging and alerting so that failures are detected and actionable, never silent
- Lead incident response and drive follow-up actions to closure
3. Data Infrastructure & Security
- Operate storage at multi-petabyte scale: object storage, local storage, tiering, backup and recovery
- Operate the ingestion path for large sensor data deliveries, with integrity verification and metadata extraction
- Implement access control and single sign-on across platform services, with audit logging
- Apply data protection measures to sensitive content before it reaches external users
This is a fullstack infrastructure role. You will work across compute, storage, networking, automation and security. We do not divide the team into narrow specialists: everyone owns a part of the platform end to end - provisioning, deployment, monitoring and incident response - and is expected to grow across the whole stack over time. Work is assigned according to what the programme needs at that moment, not by a fixed job boundary.
Much of the platform runs on-premise, including air-gapped segments. Experience operating without a public-cloud safety net is valued here
Key Responsibilities
1. AI Compute & Serving
- Operate and scale GPU clusters across training, inference and annotation workloads; manage scheduling, utilisation and capacity
- Deploy and scale high-throughput model serving; tune GPU memory and runtime performance
- Manage multi-GPU distributed training and reproducible environment sandboxing
- Scale pipeline orchestration on Kubernetes for large data processing jobs
2. Platform Automation & Delivery
- Drive GitOps-based deployment and maintain infrastructure-as-code across the platform
- Build and maintain CI pipelines, secure container builds and release automation
- Build monitoring, logging and alerting so that failures are detected and actionable, never silent
- Lead incident response and drive follow-up actions to closure
3. Data Infrastructure & Security
- Operate storage at multi-petabyte scale: object storage, local storage, tiering, backup and recovery
- Operate the ingestion path for large sensor data deliveries, with integrity verification and metadata extraction
- Implement access control and single sign-on across platform services, with audit logging
- Apply data protection measures to sensitive content before it reaches external users
Yêu cầu
- 4+ years in MLOps, DevOps, SRE or HPC platform engineering, with production ownership of AI/ML infrastructure
- Kubernetes administration at production scale: Helm, ingress (Traefik or Envoy), CNI, and GitOps (Flux or ArgoCD)
- GPU & HPC: SLURM, NVIDIA Container Toolkit, CUDA runtime tuning, multi-GPU memory debugging
- Strong Linux systems skills and infrastructure-as-code (SaltStack or Ansible)
- Python and Bash for automation, including Airflow DAGs and custom operators
- Object storage, and a monitoring and logging stack (Prometheus, Grafana or equivalent)
- Willingness to work across the full stack - compute, storage, networking, automation and security - rather than within a single specialty
- Good communication in English - technical documentation and working with international partners
- Kubernetes administration at production scale: Helm, ingress (Traefik or Envoy), CNI, and GitOps (Flux or ArgoCD)
- GPU & HPC: SLURM, NVIDIA Container Toolkit, CUDA runtime tuning, multi-GPU memory debugging
- Strong Linux systems skills and infrastructure-as-code (SaltStack or Ansible)
- Python and Bash for automation, including Airflow DAGs and custom operators
- Object storage, and a monitoring and logging stack (Prometheus, Grafana or equivalent)
- Willingness to work across the full stack - compute, storage, networking, automation and security - rather than within a single specialty
- Good communication in English - technical documentation and working with international partners
Quyền lợi
Thưởng
• Competitive income (including 13th-month salary, performance bonuses, and other rewards as regulated by Vingroup).
Chăm sóc sức khoẻ
• High-quality health insurance
Khác
• Working in a safe, modern, civilized, and professional environment with numerous opportunities for personal development, and even take the lead.
• Competitive income (including 13th-month salary, performance bonuses, and other rewards as regulated by Vingroup).
Chăm sóc sức khoẻ
• High-quality health insurance
Khác
• Working in a safe, modern, civilized, and professional environment with numerous opportunities for personal development, and even take the lead.
Thông tin khác
NGÀY ĐĂNG
23/09/2026
CẤP BẬC
Nhân viên
NGÀNH NGHỀ
Công Nghệ Thông Tin/Viễn Thông > System/Cloud/DevOps Engineer
KỸ NĂNG
Infrastructure-As-Code, DevOps, Kubernetes, SRE, MLOps
LĨNH VỰC
Ô tô
NGÔN NGỮ TRÌNH BÀY HỒ SƠ
Bất kỳ
SỐ NĂM KINH NGHIỆM TỐI THIỂU
4
QUỐC TỊCH
Không hiển thị
Xem thêm
23/09/2026
CẤP BẬC
Nhân viên
NGÀNH NGHỀ
Công Nghệ Thông Tin/Viễn Thông > System/Cloud/DevOps Engineer
KỸ NĂNG
Infrastructure-As-Code, DevOps, Kubernetes, SRE, MLOps
LĨNH VỰC
Ô tô
NGÔN NGỮ TRÌNH BÀY HỒ SƠ
Bất kỳ
SỐ NĂM KINH NGHIỆM TỐI THIỂU
4
QUỐC TỊCH
Không hiển thị
Xem thêm
Thông tin chung
- Thu nhập: Thương lượng
Nơi làm việc
- Hà Nội, Vietnam
Việc làm tương tự khác
CÔNG TY TNHH CÔNG NGHỆ HUAWEI VIỆT NAM
Hà Nội, Hà Nam
Thương lượng
Tổng Công ty Cổ phần Bưu chính Viettel - Viettel Post
Hà Nội
1000 - 3500
Công Ty Cổ Phần Sản Xuất Và Kinh Doanh VinFast - Thành Viên Của Vingroup
Xem trang công ty- Địa chỉ công ty: Khu Kinh tế Đình Vũ - Cát Hải, đảo Cát Hải - Thị trấn Cát Hải - Huyện Cát Hải - Hải Phòng.
- Quy mô: Từ 5000 - 10000 nhân viên
- Lĩnh vực: Sản xuất / Vận hành sản xuất, Kinh doanh, Ô tô - Xe máy, Cơ khí chế tạo / Điện / Điện tử / Tự động hóa
Thông tin công việc
Vị trí:
Nhân viên
Hình thức làm việc:
Toàn thời gian
Việc làm tương tự
Cảnh báo dấu hiệu lừa đảo tuyển dụng
Đội ngũ hỗ trợ của JobOKO sẵn sàng đồng hành, tư vấn và giới thiệu những cơ hội việc làm phù hợp, giúp Ứng viên tự tin phát triển sự nghiệp và chinh phục mục tiêu nghề nghiệp bền vững.
Hotline CSKH
1900.63.63.84
Công ty Cổ phần JobOKO Toàn cầu
Đội ngũ hỗ trợ của JobOKO luôn chủ động tư vấn các giải pháp tuyển dụng tối ưu, cam kết đồng hành và hỗ trợ Quý Nhà tuyển dụng đạt được hiệu quả tuyển dụng bền vững.
Hotline CSKH
0962.107.888
Công ty Cổ phần JobOKO Toàn cầu