Mô tả công việc
Job Purpose
We are hiring a Senior Observability Engineer / Platform Owner to drive enterprise observability enablement and ongoing operational visibility across critical digital manufacturing systems. This role will lead incoming dashboard, log, trace, alerting, and telemetry requests from ETS, MES, UFE, APS, AIP, SAP, Teamcenter, and related teams, turning them into practical, scalable, and maintainable observability solutions. The role is expected to own requirement clarification, solution design, prioritization, quality review, and cross-team execution, while leading 2 interns on day-to-day delivery and maintenance.
Key Responsibilities
• Serve as the single intake point for observability requests across ETS, MES, UFE, APS, AIP, SAP, Teamcenter, and related teams.
• Design and break down solutions covering dashboards, alerts, logs, traces, matrix configuration, and telemetry collection.
• Evaluate priority, delivery sequencing, dependencies, and effort, and maintain the backlog.
• Review deliverables from 2 interns to ensure quality, consistency, reuse, and maintainability.
• Handle complex topics such as cross-system tracing, log ingestion boundaries, metric definitions, and alert threshold design.
• Drive key initiatives such as HVR sync alerting, MES logs to Loki, Alloy standardization, and Teamcenter / SAP observability visibility.
• Build and maintain runbooks, handover documentation, templates, and configuration standards.
• Work closely with K8s / SRE, DBA, Windows host admins, and application owners to land observability solutions.
• Lead known issue management, top error log analysis, and rule design for alert-to-known-issue closure.
• Define use cases for distributed trace support dashboards, trace detail workflows, and AI-assisted analysis support.
• Govern the boundary between Grafana-managed alerts and data source-managed alerts, including notification paths, Alertmanager configuration ownership, and template strategy.
• Design solutions for complex dashboard data models such as QMS multi-table queries, panel reuse, hidden variables, and GUI embedding constraints.
• Define onboarding and integration approaches for external metrics such as Confluent Kafka metrics, AdminTools-based metrics, HVR sync metrics, and IIoT / Ignition health metrics.
• As a longer-term evolution target, help transition the observability platform toward AIOps capability by preparing usable data foundations for anomaly detection, AI-assisted RCA, and intelligent alerting.
• Drive the capture and organization of key AIOps inputs such as metrics, logs, alerts, application dependency trees / graphs, application change events, historical issues, and RCA records.
• Work with SRE, ETS, and application teams to establish feedback loops for auto-RCA outputs, operational knowledge capture, and future optimization.
• Identify which alerting, log analysis, and known issue scenarios are suitable for future anomaly detection, intelligent analysis, or agent-assisted troubleshooting, without disrupting current observability foundation priorities.
Yêu cầu
Qualifications
• Strong hands-on experience with Grafana, Prometheus, Loki, Tempo, or similar observability platforms.
• Solid experience in dashboard design, PromQL or equivalent query languages, metric modeling, and alert design.
• Good understanding of how logs, traces, and metrics work together in end-to-end troubleshooting.
• Experience in requirement clarification, solution design, and technical coordination across teams.
• Ability to read and adjust telemetry and collection configuration such as Alloy, exporters, scrape jobs, and remote_write.
• Strong documentation, review, and mentoring skills.
• Understanding of OpenTelemetry, OpenObserve, collectors / gateways, distributed trace querying, and log correlation analysis.
• Familiarity with Grafana contact points, webhooks, AlertmanagerConfig, Teams / Power Automate style notification workflows.
• Ability to handle complex dashboard data patterns, including SQL joins, panel reuse, hidden variables, iframe constraints, and multi-environment datasource design.
• Familiarity with telegraf, custom SQL metrics, ScrapeConfig, and third-party metrics onboarding approaches.
• Understanding of AIOps prerequisites such as data quality, historical RCA capture, event correlation, change event ingestion, and usable operational knowledge bases.
• Basic understanding of anomaly detection, intelligent alerting, AI-assisted RCA, and agentic workflow concepts, with the ability to judge when they are appropriate and when they are premature.
Nice to have:
• Experience in manufacturing or enterprise platforms such as MES, APS, SAP, Teamcenter, AIP, or UFE.
• Experience with Loki migrations, OpenTelemetry, distributed tracing, or alert governance.
• Experience leading interns or junior engineers.
• Experience with runbooks, KT materials, known issue repositories, or RCA knowledge capture.
• Experience with Grafana webhooks, GenAI webhooks, known issue automation, or support dashboards.
• Experience with QMS / reporting-style dashboards, SQL optimization, or GUI integration.
• Experience onboarding AdminTools, infra-monitor-tools, HVR sync, or Confluent Kafka metrics.
• Experience with AIOps, anomaly detection, auto-RCA, LLM / GenAI-assisted operations, event correlation, or operational knowledge management.
• Experience organizing application dependency graphs, change events, incident datasets, or RCA datasets for intelligent operations use cases.
Quyền lợi
Thưởng
13th , Retention, Perfomance
Thông tin khác
Xem thêm
Thông tin chung
Nơi làm việc