Site Reliability Engineer

Hồ Chí Minh
Thỏa thuận
2 năm kinh nghiệm
Hạn nộp hồ sơ: 17/09/2026 (Còn 33 ngày)
Ứng tuyển sớm để được ưu tiên
Kết nối với Nhà tuyển dụng để tìm hiểu thông tin và gia tăng cơ hội trúng tuyển
OMess
Nhà tuyển dụng đang online

Mô tả công việc

Top 3 Reasons To Join Us
Flexible Friday afternoon
18 Annual Leave + 5 Recharge Days/ Year
Hybrid working model
The Job

Roles & Responsibilities

Incident Management and Reliability

  • Lead or coordinate incident response for assigned markets and services, with independent ownership of routine incidents and supported ownership of complex multi-system incidents
  • Facilitate communication during incidents and drive timely resolution, including clean shift handoffs across global time zones
  • Contribute to post-incident reviews and root cause analysis activities
  • Ensure corrective and preventive actions are identified, tracked, and completed for owned services

Monitoring, Alerting, and Observability

  • Implement and continuously optimize monitoring, logging, alerting, and tracing solutions
  • Develop meaningful alerts based on service behavior, customer impact, and business priorities
  • Build and maintain dashboards that provide actionable insights into system performance and reliability

Platform and Market Ownership

  • Own day-to-day SRE responsibilities for one or more assigned markets, platforms, or services
  • Ensure monitoring coverage, dashboards, runbooks, and alerting configurations remain accurate, effective, and up to date
  • Assess platform health, identify reliability risks, and raise improvements before incidents occur
  • Partner with engineering teams to ensure new features and services meet reliability requirements before production release
  • Track reliability metrics for owned services, including Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets

Platform Engineering, Automation, and AI

  • Develop and maintain tools, scripts, and automation that reduce manual effort and improve operational efficiency
  • Identify and eliminate repetitive tasks through automation, with a bias toward self-service capabilities over ticket queues
  • Contribute to auto-healing and auto-remediation capabilities: detection rules, remediation runbooks as code, and automated response workflows
  • Build and extend internal platform tooling using infrastructure as code and CI/CD pipelines rather than manual configuration
  • Apply team best practices for the responsible use of AI within SRE workflows, including AI-assisted diagnosis, triage, and remediation

Team Contribution

  • Share knowledge through documentation, training sessions, and post-incident learning activities
  • Review runbooks and operational documentation to maintain quality standards
  • Contribute to the continuous improvement of team processes, standards, and ways of working

Your Skills and Experience
  • 2+ years of experience in SRE, DevOps, production support, or infrastructure engineering roles
  • Hands-on experience with monitoring and observability tooling (e.g., Datadog, Prometheus, Grafana, CloudWatch, or similar)
  • Working knowledge of at least one major cloud provider (AWS preferred)
  • Proficiency in at least one scripting or programming language (e.g., Python, Bash, Go) for automation, with demonstrated examples of automating away manual operational work
  • Experience participating in incident response and on-call or shift-based operations
  • Understanding of SLI/SLO concepts and reliability engineering fundamentals
  • Ability to work follow-the-sun shift rotations, including structured handoffs with teams in other regions
  • Strong written and verbal English communication skills for cross-region collaboration
  • Experience with Kubernetes, Docker, and container orchestration in production
  • Infrastructure as code experience (e.g., Terraform, CloudFormation)
  • CI/CD and GitOps pipeline experience (e.g., GitLab CI, ArgoCD)
  • Exposure to platform engineering concepts: internal developer platforms, self-service tooling, developer portals (e.g., Port, Backstage)
  • Experience buildiExperience supporting distributed systems across multiple markets or regions
  • Familiarity with AI-assisted operations tooling
  • Relevant certifications (AWS, CKA, or similar)ng or contributing to auto-remediation or event-driven automation workflows.

Why You'll Love Working Here

Attractive Benefits:

  • 100% salary during probation period
  • Annual Leave: 18 days/ year
  • Five "Recharge Days" - Extra days, in addition to company holidays.
  • Flexible Friday afternoon
  • Full salary insurance
  • 13th-month bonus
  • Gift + 1 day off for birthday
  • Advanced health insurance (Generali)
  • Regular engagement activities: sport clubs, internal event...
  • Support Macbook and Monitor

Yêu cầu

DevOps, English, Kubernetes, Docker, Python, AWS

Quyền lợi

Attractive Benefits:

  • 100% salary during probation period
  • Annual Leave: 18 days/ year
  • Five "Recharge Days" - Extra days, in addition to company holidays.
  • Flexible Friday afternoon
  • Full salary insurance
  • 13th-month bonus
  • Gift + 1 day off for birthday
  • Advanced health insurance (Generali)
  • Regular engagement activities: sport clubs, internal event...
  • Support Macbook and Monitor

Thông tin chung

  • Thu nhập: You'll love it

Pizza Hut Digital & Technology

Xem trang công ty
Địa chỉ công ty: Tan Binh, Ho Chi Minh
Quy mô: Từ 101 - 500 nhân viên
Thông tin công việc
Vị trí:
Nhân viên
Hình thức làm việc:
Toàn thời gian
Việc làm tương tự
Automation & Process Excellence Specialist
CÔNG TY TNHH BUYMED LOGISTICS
Hồ Chí Minh
CÔNG TY TNHH BUYMED LOGISTICS
Thỏa thuận
Chuyên viên - Cung ứng hàng hóa (Demand Planning)
Công ty Cổ phần Vàng bạc Đá quý Phú Nhuận - PNJ
Hồ Chí Minh
Công ty Cổ phần Vàng bạc Đá quý Phú Nhuận - PNJ
15,000,000 - 20,000,000 VNĐ
AI Engineer
Công ty Cổ phần CITICS
Hồ Chí Minh
Công ty Cổ phần CITICS
You'll love it
[HCM] Product Manager Embeeded Finance
Công ty Cổ phần Bán lẻ Kỹ thuật số FPT
Hồ Chí Minh
Công ty Cổ phần Bán lẻ Kỹ thuật số FPT
Thương lượng
Junior Developer (Cloud/Odoo) - Lương Upto 15 Triệu
Công ty Cổ phần Phân Phối Quốc tế
Hà Nội, Hồ Chí Minh, Hòa Bình, Phú Thọ
Công ty Cổ phần Phân Phối Quốc tế
10 - 15 triệu VNĐ
Cảnh báo dấu hiệu lừa đảo tuyển dụng
Thu phí và cung cấp thông tin
  • Phí hồ sơ, đồng phục, đặt cọc.
  • Yêu cầu nộp bản gốc giấy tờ.
  • Cung cấp mã OTP.
Hứa hẹn trúng tuyển 100%
  • Không yêu cầu trình độ.
  • Không cần thử việc.
Phỏng vấn bất thường
  • Địa điểm xa văn phòng công ty.
  • Phỏng vấn qua Telegram.
Yêu cầu làm nhiệm vụ
  • Tải app, nạp tiền.
  • Làm nhiệm vụ nhận thưởng.
Tin tuyển dụng sơ sài
  • Mô tả công việc chung chung
  • Nhiệm vụ đơn giản, thu nhập khủng
  • Lỗi chính tả, đánh máy.
Đội ngũ hỗ trợ của JobOKO sẵn sàng đồng hành, tư vấn và giới thiệu những cơ hội việc làm phù hợp, giúp Ứng viên tự tin phát triển sự nghiệp và chinh phục mục tiêu nghề nghiệp bền vững.
Đội ngũ hỗ trợ của JobOKO luôn chủ động tư vấn các giải pháp tuyển dụng tối ưu, cam kết đồng hành và hỗ trợ Quý Nhà tuyển dụng đạt được hiệu quả tuyển dụng bền vững.