Senior Site Reliability Engineer SRE, GCP, Kubernetes

Hồ Chí Minh
Thỏa thuận
4 - 10 năm kinh nghiệm
Hạn nộp hồ sơ: 08/10/2026 (Còn 28 ngày)
Ứng tuyển sớm để được ưu tiên
Kết nối với Nhà tuyển dụng để tìm hiểu thông tin và gia tăng cơ hội trúng tuyển
OMess
Nhà tuyển dụng đang online

Mô tả công việc

Top 3 Reasons To Join Us
Build impactful products from the ground up
Tackle high-scale, real-world tech challenges
Grow with a multidisciplinary tech team
The Job

DatVietVAC Group Holdings is looking for a Senior SRE Engineer to lead the infrastructure architecture, platform reliability, and production operations of our Fan Commerce Platform.

The platform operates across Google Cloud for computing and scaling and CMC Cloud for local data residency. You will be responsible for building a secure, scalable, and highly available infrastructure capable of supporting high-traffic on-sale periods and major entertainment events while maintaining operational efficiency and infrastructure costs within the approved budget.

1. Cloud Infrastructure Architecture

  • Design and implement multi-environment infrastructure, including Production, Staging, and Development, using Google Cloud services such as GKE, Cloud SQL, Memorystore, and Pub/Sub.
  • Design the hybrid-cloud architecture between Google Cloud and CMC Cloud, including private connectivity, VPN configuration, network segmentation, and data-flow separation.
  • Build and manage infrastructure as code using Terraform.
  • Develop and maintain CI/CD pipelines and automation tools that support engineering teams throughout the software development and release lifecycle.
  • Establish a comprehensive observability platform covering metrics, logs, traces, dashboards, and alerting.
  • Ensure that the infrastructure architecture supports scalability, maintainability, security, and local data-residency requirements.

2. Site Reliability and Incident Response

  • Define and manage Service Level Indicators, Service Level Objectives, and error budgets for core platform services.
  • Take ownership of platform availability, reliability, scalability, and operational readiness.
  • Lead capacity planning and load testing for high-traffic product launches, ticket or merchandise on-sales, and live events.
  • Design autoscaling strategies, traffic-management mechanisms, and overload-protection measures.
  • Participate in the 24/7 production on-call rotation and lead the response to high-severity incidents.
  • Lead blameless postmortems, identify root causes, and ensure that corrective and preventive actions are completed.
  • Develop incident-response procedures, operational runbooks, and disaster-recovery plans.
  • Coach and mentor engineers in production operations, incident management, and reliability practices.

3. Security and Compliance

  • Implement infrastructure security controls, including WAF, DDoS protection, and bot-management solutions using tools such as Cloudflare and Google Cloud Armor.
  • Manage secrets, encryption, identity, and access controls across cloud environments.
  • Establish backup, recovery, and disaster-recovery processes and conduct periodic recovery drills.
  • Collaborate with relevant teams on penetration testing, vulnerability remediation, security audits, and compliance requirements.
  • Ensure appropriate monitoring and audit logging for sensitive infrastructure and operational activities.

4. Performance and Cost Optimization

  • Monitor, analyze, and optimize cloud infrastructure costs through rightsizing, committed-use discounts, storage lifecycle management, and egress optimization.
  • Investigate performance bottlenecks and implement solutions to improve platform efficiency, scalability, and resilience.
  • Provide infrastructure and reliability recommendations to the Project Lead and engineering teams.
  • Evaluate and propose technologies that support the platform's long-term architecture and business requirements.

Your Skills and Experience

Education

  • Bachelor's degree or higher in Information Technology, Software Engineering, Computer Science, or a related field.

Experience

  • At least five years of experience as a DevOps Engineer, Site Reliability Engineer, System Engineer, Cloud Engineer, or in a similar infrastructure role.
  • At least one year of experience at an equivalent senior level or in a technical leadership role.
  • Proven experience designing or operating production systems with high traffic or significant peak-load events, such as flash sales, on-sales, live events, entertainment platforms, or e-commerce platforms.
  • Hands-on experience operating Kubernetes and cloud infrastructure in a production environment.
  • Production experience with Google Cloud Platform is mandatory.

Technical Knowledge and Skills

  • Advanced knowledge of networking and security concepts, including TCP/IP, HTTP/1.1, HTTP/2, HTTP/3, DNS, gRPC, VPC peering, Cloud VPN, and on-premises-to-cloud connectivity.
  • Strong understanding of distributed systems, microservices, clustering, replication, failover, load balancing, and autoscaling.
  • Strong hands-on experience with Kubernetes, particularly Google Kubernetes Engine, and container technologies in production environments.
  • Proficiency in infrastructure as code and automation tools, particularly Terraform and Ansible.
  • Strong experience building and maintaining CI/CD pipelines using GitHub Actions, Jenkins, or equivalent tools.
  • Strong Linux administration skills, including experience with Ubuntu or CentOS.
  • Hands-on experience with Google Cloud services; additional AWS or Azure experience is an advantage.
  • Experience with performance testing, load testing, bottleneck analysis, capacity planning, and system optimization.
  • Knowledge of API gateways, reverse proxies, ingress controllers, and service mesh architecture.
  • Experience building observability systems using metrics, logs, traces, dashboards, and alerts.
  • Knowledge of SLI/SLO management, error budgets, incident response, postmortems, backup, and disaster recovery.
  • Experience with infrastructure security solutions such as WAF, DDoS protection, bot management, secrets management, and access control.

Soft Skills

  • Strong technical leadership and cross-functional collaboration skills.
  • Ability to mentor engineers and guide teams through incident response and production operations.
  • Calm, structured, and decisive when handling critical production incidents.
  • Strong analytical, troubleshooting, and risk-management skills.
  • Data-driven approach to reliability, performance, capacity, and cost optimization.
  • Strong attention to detail and a high sense of ownership.
  • Ability to explain technical decisions and trade-offs clearly to both technical and non-technical stakeholders.

Language Skills

  • Basic English communication skills.
  • Good ability to read and understand technical documentation in English.

Other Requirements

  • Willingness to participate in the production on-call rotation and support major on-sale periods or live events when required.
  • Ability to work under pressure during critical releases, peak-traffic events, and production incidents.

Why You'll Love Working Here

• Full statutory insurance, including Social Insurance, Health Insurance and Unemployment Insurance, based on 100% of the official salary and in compliance with Vietnamese labor regulations.
• Working hours: Monday to Friday, from 8:30 AM to 5:30 PM, with a one-hour lunch break.
• 14 days of annual leave.
• Company-provided working equipment, including a laptop or desktop computer.
• Employee parking area.
• Annual PMP performance bonus, subject to individual KPI achievement and the Company's business performance.
• Periodic health check-ups.
• Employee engagement programs and internal activities throughout the year.

Yêu cầu

Kubernetes, Terraform, CI/CD, System Architecture, DevOps, GCP

Quyền lợi

• Full statutory insurance, including Social Insurance, Health Insurance and Unemployment Insurance, based on 100% of the official salary and in compliance with Vietnamese labor regulations.
• Working hours: Monday to Friday, from 8:30 AM to 5:30 PM, with a one-hour lunch break.
• 14 days of annual leave.
• Company-provided working equipment, including a laptop or desktop computer.
• Employee parking area.
• Annual PMP performance bonus, subject to individual KPI achievement and the Company's business performance.
• Periodic health check-ups.
• Employee engagement programs and internal activities throughout the year.

Thông tin chung

  • Thu nhập: You'll love it

Nơi làm việc

  • 222 Pasteur, Phường Xuân Hòa, Ho Chi Minh
Việc làm tương tự khác
CÔNG TY CỔ PHẦN DỊCH VỤ VÀ KỸ THUẬT SMC
Hồ Chí Minh, Đồng Nai, Tây Ninh
CẠNH TRANH
Cennext Co., Ltd
Hà Nội, Hồ Chí Minh, Đà Nẵng, Hải Phòng, Bà Rịa - Vũng Tàu
Từ 15 - 25 triệu VND

CÔNG TY CỔ PHẦN DATVIET VAC GROUP HOLDINGS

Xem trang công ty
Địa chỉ công ty: 222 Pasteur - Phường Võ Thị Sáu - Quận 3 - TP. Hồ Chí Minh
Quy mô: Từ 101 - 500 nhân viên
Lĩnh vực: Truyền thông/Internet/Online Media, Biên tập/ Báo chí/ Truyền hình
Thông tin công việc
Vị trí:
Hình thức làm việc:
Toàn thời gian
Việc làm tương tự
Kỹ Sư Phụ Trách Hồ Sơ Thanh Quyết Toán Công Trình - Đi Làm Ngay Tại TP. Hồ Chí Minh
Công ty TNHH đầu tư và xây lắp FSC
Hồ Chí Minh
Công ty TNHH đầu tư và xây lắp FSC
15 - 25 triệu VND + Phụ Cấp
[HCM] Kỹ Sư Triển Khai Bản Vẽ Shopdrawing (Fresher) - Không yêu cầu kinh nghiệm, Ưu tiên sinh viên mới tốt nghiệp
CÔNG TY TNHH SANKEN SCUBE
Hồ Chí Minh
CÔNG TY TNHH SANKEN SCUBE
Thỏa Thuận
QC Ngành May Mặc | Thu Nhập 20-25 Triệu | 3+ Năm Kinh Nghiệm
CÔNG TY TNHH J-LONG VIỆT NAM
Hồ Chí Minh
CÔNG TY TNHH J-LONG VIỆT NAM
20 - 25 triệu VND
Kế toán trưởng Ưu tiên Nam có kinh nghiệm trong lĩnh vực xây dựng
CÔNG TY CỔ PHẦN TẬP ĐOÀN FECON
Hồ Chí Minh
CÔNG TY CỔ PHẦN TẬP ĐOÀN FECON
Thoả thuận
KỸ SƯ THIẾT KẾ ĐIỆN - DỰ ÁN SÂN BAY LONG THÀNH
CÔNG TY CP JESCO ASIA
Hồ Chí Minh
CÔNG TY CP JESCO ASIA
12 triệu - 18 triệu VND
Cảnh báo dấu hiệu lừa đảo tuyển dụng
Thu phí và cung cấp thông tin
  • Phí hồ sơ, đồng phục, đặt cọc.
  • Yêu cầu nộp bản gốc giấy tờ.
  • Cung cấp mã OTP.
Hứa hẹn trúng tuyển 100%
  • Không yêu cầu trình độ.
  • Không cần thử việc.
Phỏng vấn bất thường
  • Địa điểm xa văn phòng công ty.
  • Phỏng vấn qua Telegram.
Yêu cầu làm nhiệm vụ
  • Tải app, nạp tiền.
  • Làm nhiệm vụ nhận thưởng.
Tin tuyển dụng sơ sài
  • Mô tả công việc chung chung
  • Nhiệm vụ đơn giản, thu nhập khủng
  • Lỗi chính tả, đánh máy.
Đội ngũ hỗ trợ của JobOKO sẵn sàng đồng hành, tư vấn và giới thiệu những cơ hội việc làm phù hợp, giúp Ứng viên tự tin phát triển sự nghiệp và chinh phục mục tiêu nghề nghiệp bền vững.
Đội ngũ hỗ trợ của JobOKO luôn chủ động tư vấn các giải pháp tuyển dụng tối ưu, cam kết đồng hành và hỗ trợ Quý Nhà tuyển dụng đạt được hiệu quả tuyển dụng bền vững.