Senior DevOps & Site Reliability Engineer (SRE)
Hạn nộp hồ sơ: 07/10/2026 (Còn 27 ngày)
Ứng tuyển sớm để được ưu tiên
Kết nối với Nhà tuyển dụng để tìm hiểu thông tin và gia tăng cơ hội trúng tuyển
Nhà tuyển dụng đang online
Mô tả công việc
We are orchestrating a hyperscale, "Best-of-Breed" digital transformation for a premier multimedia conglomerate. Our operational hub is Odoo CE 18 (heavily extended via OCA), integrated synchronously with Camunda 8, Bravo 10 ERP, and ResourceSpace DAM.
In an environment handling massive volumes of digital syndication, high-speed Procure-to-Pay (P2P) micro-transactions, and heavy media production workflows, infrastructure is not just a support function-it is the bedrock of our operational survival. We do not tolerate downtime, database bloat, or silent failures.
As our Senior DevOps & SRE, you will own the PRODUCTION environments. You will architect the deployment lifecycles, engineer proactive observability matrices, and govern a hybrid-cloud topology that spans global providers and localized Vietnamese infrastructure. You will work alongside the Lead Solution Architect to ensure our enterprise architecture runs with absolute resilience, sub-second latency, and uncompromising data integrity.
Core Responsibilities
High-Velocity CI/CD & Python Automation: Architect, script, and maintain uncompromising CI/CD pipelines capable of handling high-frequency, multi-daily deployments. You will leverage advanced Python scripting to automate the entire lifecycle (Build, Test, Deploy), ensuring zero-downtime releases and mathematical certainty before code hits Production.
Advanced Observability & APM: You will not wait for users to report errors. You will deploy and manage a comprehensive, real-time observability stack. This includes centralized logging, end-user experience monitoring (APM/RUM), and intelligent, threshold-based alerting to detect and neutralize PostgreSQL bottlenecks, API polling failures, or queue_job queue blockages before they impact the business.
Hybrid-Cloud Architecture: Provision, secure, and scale our infrastructure across a multi-cloud matrix. You will manage workloads spanning global hyperscalers (AWS, GCP, BytePlus) and localized sovereign clouds (CMC Cloud, FPT Cloud) to ensure optimal latency and regulatory compliance for our financial and operational data.
Media Storage Tiering & Data Lifecycle: Because our ecosystem relies on a strict "Pointer Architecture" to keep Odoo lean, you will oversee the infrastructure connecting our cloud environments with the enterprise network storage clusters used by our media production units. You will execute and optimize unified data lifecycle policies and storage tiering designs, ensuring that petabytes of heavy binary assets (managed via ResourceSpace) are stored, archived, or purged efficiently without degrading ERP performance.
Database & Worker Tuning: Execute deep PostgreSQL tuning (shared buffers, work memory, vacuuming strategies) and manage Odoo's multi-processing worker threads to sustain high-volume asynchronous webhook traffic without locking the database.
In an environment handling massive volumes of digital syndication, high-speed Procure-to-Pay (P2P) micro-transactions, and heavy media production workflows, infrastructure is not just a support function-it is the bedrock of our operational survival. We do not tolerate downtime, database bloat, or silent failures.
As our Senior DevOps & SRE, you will own the PRODUCTION environments. You will architect the deployment lifecycles, engineer proactive observability matrices, and govern a hybrid-cloud topology that spans global providers and localized Vietnamese infrastructure. You will work alongside the Lead Solution Architect to ensure our enterprise architecture runs with absolute resilience, sub-second latency, and uncompromising data integrity.
Core Responsibilities
High-Velocity CI/CD & Python Automation: Architect, script, and maintain uncompromising CI/CD pipelines capable of handling high-frequency, multi-daily deployments. You will leverage advanced Python scripting to automate the entire lifecycle (Build, Test, Deploy), ensuring zero-downtime releases and mathematical certainty before code hits Production.
Advanced Observability & APM: You will not wait for users to report errors. You will deploy and manage a comprehensive, real-time observability stack. This includes centralized logging, end-user experience monitoring (APM/RUM), and intelligent, threshold-based alerting to detect and neutralize PostgreSQL bottlenecks, API polling failures, or queue_job queue blockages before they impact the business.
Hybrid-Cloud Architecture: Provision, secure, and scale our infrastructure across a multi-cloud matrix. You will manage workloads spanning global hyperscalers (AWS, GCP, BytePlus) and localized sovereign clouds (CMC Cloud, FPT Cloud) to ensure optimal latency and regulatory compliance for our financial and operational data.
Media Storage Tiering & Data Lifecycle: Because our ecosystem relies on a strict "Pointer Architecture" to keep Odoo lean, you will oversee the infrastructure connecting our cloud environments with the enterprise network storage clusters used by our media production units. You will execute and optimize unified data lifecycle policies and storage tiering designs, ensuring that petabytes of heavy binary assets (managed via ResourceSpace) are stored, archived, or purged efficiently without degrading ERP performance.
Database & Worker Tuning: Execute deep PostgreSQL tuning (shared buffers, work memory, vacuuming strategies) and manage Odoo's multi-processing worker threads to sustain high-volume asynchronous webhook traffic without locking the database.
Yêu cầu
Required Expertise (The Non-Negotiables)
Python Automation Mastery: You are a highly proficient Python programmer. You write clean, scalable automation scripts to replace manual infrastructure tasks, deployment steps, and system recovery protocols.
Hybrid-Cloud Command: Demonstrable, hands-on experience provisioning and managing infrastructure on AWS, GCP, and BytePlus, paired with a strong working knowledge of domestic Vietnamese cloud providers (CMC Cloud, FPT Cloud).
Production Odoo & PostgreSQL: Proven experience deploying, scaling, and tuning Odoo (preferably v16+) and PostgreSQL in high-availability, clustered PRODUCTION environments.
Observability & Telemetry Stack: Deep expertise in deploying centralized logging (e.g., ELK/EFK stack, Loki), metrics collection (Prometheus, Grafana), and Application Performance Monitoring (APM) tools (e.g., Datadog, New Relic, Sentry).
Infrastructure as Code (IaC): Mastery of Terraform, Ansible, or similar IaC tools to ensure our infrastructure is version-controlled, reproducible, and immutable.
Nice-To-Haves
Experience managing high-throughput network-attached storage (NAS) and media asset lifecycles.
Familiarity with the orchestration deployment of Camunda 8 (Zeebe, Operate, Tasklist) on Kubernetes.
Experience with disaster recovery (DR) protocols and Point-in-Time Recovery (PITR) for mission-critical financial databases.
Python Automation Mastery: You are a highly proficient Python programmer. You write clean, scalable automation scripts to replace manual infrastructure tasks, deployment steps, and system recovery protocols.
Hybrid-Cloud Command: Demonstrable, hands-on experience provisioning and managing infrastructure on AWS, GCP, and BytePlus, paired with a strong working knowledge of domestic Vietnamese cloud providers (CMC Cloud, FPT Cloud).
Production Odoo & PostgreSQL: Proven experience deploying, scaling, and tuning Odoo (preferably v16+) and PostgreSQL in high-availability, clustered PRODUCTION environments.
Observability & Telemetry Stack: Deep expertise in deploying centralized logging (e.g., ELK/EFK stack, Loki), metrics collection (Prometheus, Grafana), and Application Performance Monitoring (APM) tools (e.g., Datadog, New Relic, Sentry).
Infrastructure as Code (IaC): Mastery of Terraform, Ansible, or similar IaC tools to ensure our infrastructure is version-controlled, reproducible, and immutable.
Nice-To-Haves
Experience managing high-throughput network-attached storage (NAS) and media asset lifecycles.
Familiarity with the orchestration deployment of Camunda 8 (Zeebe, Operate, Tasklist) on Kubernetes.
Experience with disaster recovery (DR) protocols and Point-in-Time Recovery (PITR) for mission-critical financial databases.
Quyền lợi
Competitive income commensurate with qualifications and experience;
At least 12 days of annual leave per year;
Full-time employees are entitled to 01 paid day off during their birthday month;
Full participation in compulsory insurance schemes in accordance with applicable labor laws;
Premium health insurance coverage sponsored by the Company upon signing the Labor Contract;
Annual health check-up.
At least 12 days of annual leave per year;
Full-time employees are entitled to 01 paid day off during their birthday month;
Full participation in compulsory insurance schemes in accordance with applicable labor laws;
Premium health insurance coverage sponsored by the Company upon signing the Labor Contract;
Annual health check-up.
Thông tin khác
Thời gian làm việc
Thứ 2 - Thứ 6 (từ 09:00 đến 18:00)
Thứ 2 - Thứ 6 (từ 09:00 đến 18:00)
Thông tin chung
- Thu nhập: Thoả thuận
Nơi làm việc
- Hồ Chí Minh: Phường Tân Định (Quận 1 cũ)
Việc làm tương tự khác
CÔNG TY CỔ PHẦN TẬP ĐOÀN YEAH1
Xem trang công ty- Địa chỉ công ty: 191 Nam Kỳ Khởi Nghĩa, Phường 7, Quận 3, Thành phố Hồ Chí Minh
- Quy mô: Từ 101 - 500 nhân viên
- Lĩnh vực: Truyền thông/Internet/Online Media, Marketing / Truyền thông / Quảng cáo
Thông tin công việc
Vị trí:
Nhân viên
Hình thức làm việc:
Toàn thời gian
Việc làm tương tự
DevOps Engineer Cloud Native / Kubernetes
Endava Limited Liability Company
Hà Nội, Hồ Chí Minh, Phú Thọ, Phú Yên
You'll love it
Infrastructure (Linux/Azure/Automation)
CÔNG TY TNHH GIẢI PHÁP PHÂN TÍCH DỮ LIỆU INSIGHT DATA
Hồ Chí Minh
Thoả thuận
Cảnh báo dấu hiệu lừa đảo tuyển dụng
Đội ngũ hỗ trợ của JobOKO sẵn sàng đồng hành, tư vấn và giới thiệu những cơ hội việc làm phù hợp, giúp Ứng viên tự tin phát triển sự nghiệp và chinh phục mục tiêu nghề nghiệp bền vững.
Hotline CSKH
1900.63.63.84
Công ty Cổ phần JobOKO Toàn cầu
Đội ngũ hỗ trợ của JobOKO luôn chủ động tư vấn các giải pháp tuyển dụng tối ưu, cam kết đồng hành và hỗ trợ Quý Nhà tuyển dụng đạt được hiệu quả tuyển dụng bền vững.
Hotline CSKH
0962.107.888
Công ty Cổ phần JobOKO Toàn cầu