Middle Data Engineer (Apache Spark, Trino)
- Thỏa thuận
- 3 năm kinh nghiệm
Hạn nộp hồ sơ: 12/11/2026 (Còn 58 ngày)
Ứng tuyển sớm để được ưu tiên
Kết nối với Nhà tuyển dụng để tìm hiểu thông tin và gia tăng cơ hội trúng tuyển
Nhà tuyển dụng đang online
Top 3 reasons to join us
Attractive salary
Good working environment
Flexible management
Job description
Role Overview
We are seeking a Middle Data Engineer to build and scale our Production Datahouse on-premise environment. This is a purely technical, hands-on role focused on implementing a high-performance, on-premise data architecture.
You will own the end-to-end technical execution, managing the entire lifecycle of data as it moves from landing ingestion to the final serving layer. This role is not about high-level theory; it is about the hard engineering required to optimize on-premise hardware, tune distributed processing engines, and ensure a seamless data flow across our internal ecosystem.
Key Responsibilities
Data Pipeline Engineering: Design, build, and maintain automated data ingestion pipelines from several source systems using Airbyte and Airflow, ensuring reliability and scalability.
Lakehouse Architecture Implementation: Develop and manage the Medallion Architecture (Bronze/Silver/Gold) using dbt-spark to enable structured, high-quality analytical datasets.
Platform Performance Optimization: Optimize Spark and Trino workloads on Kubernetes and manage the full lifecycle of Apache Iceberg tables (compaction, snapshotting, and maintenance) to improve query performance, storage and reduce latency.
Infrastructure & Reliability Management: Proactively monitor and maintain Data Platform resources using Grafana and Prometheus; analyze query patterns to enhance system stability and performance for production users.
Data Security & Governance: Enforce row- and column-level security using Open Policy Agent (OPA) and develop foundational frameworks for Data Governance and Data Quality, embedding lineage, automated testing, and compliance directly into pipelines.
Enterprise Data Integration: Architect and implement high-performance data bridges to deliver on-premise Lakehouse data into Microsoft Fabric, enabling seamless and secure hybrid-cloud connectivity.
Your skills and experience
3+ years of hands-on experience in Data Engineering or Platform Engineering, building and operating production-grade data platforms.
Strong expertise in distributed processing with Apache Spark and Trino, including performance tuning and optimization.
Proven experience designing and maintaining reliable data pipelines using Airbyte and Apache Airflow.
Solid understanding of modern Lakehouse architectures, dbt-based transformations, and Iceberg table management.
Hands-on experience with Kubernetes, containerized deployments, and CI/CD for data workloads.
Knowledge of data security, governance, and access control practices (row/column-level security, policy enforcement, IAM).
Experience supporting BI/analytics use cases and integrating with enterprise reporting tools (e.g., Microsoft Fabric, Power BI).
Strong troubleshooting, system reliability, and problem-solving skills in complex distributed environments.
Ability to work independently in a highly technical, ownership-driven role.
Why you'll love working here
**easy going and friendly environment
**Remuneration Package:
Working hours: 5 days per week (in a professional and yet young, dynamic environment);
Salary: Competitive remuneration package (based on skills and experience);
12 leave days per year
Attractive salary
Good working environment
Flexible management
Job description
Role Overview
We are seeking a Middle Data Engineer to build and scale our Production Datahouse on-premise environment. This is a purely technical, hands-on role focused on implementing a high-performance, on-premise data architecture.
You will own the end-to-end technical execution, managing the entire lifecycle of data as it moves from landing ingestion to the final serving layer. This role is not about high-level theory; it is about the hard engineering required to optimize on-premise hardware, tune distributed processing engines, and ensure a seamless data flow across our internal ecosystem.
Key Responsibilities
Data Pipeline Engineering: Design, build, and maintain automated data ingestion pipelines from several source systems using Airbyte and Airflow, ensuring reliability and scalability.
Lakehouse Architecture Implementation: Develop and manage the Medallion Architecture (Bronze/Silver/Gold) using dbt-spark to enable structured, high-quality analytical datasets.
Platform Performance Optimization: Optimize Spark and Trino workloads on Kubernetes and manage the full lifecycle of Apache Iceberg tables (compaction, snapshotting, and maintenance) to improve query performance, storage and reduce latency.
Infrastructure & Reliability Management: Proactively monitor and maintain Data Platform resources using Grafana and Prometheus; analyze query patterns to enhance system stability and performance for production users.
Data Security & Governance: Enforce row- and column-level security using Open Policy Agent (OPA) and develop foundational frameworks for Data Governance and Data Quality, embedding lineage, automated testing, and compliance directly into pipelines.
Enterprise Data Integration: Architect and implement high-performance data bridges to deliver on-premise Lakehouse data into Microsoft Fabric, enabling seamless and secure hybrid-cloud connectivity.
Your skills and experience
3+ years of hands-on experience in Data Engineering or Platform Engineering, building and operating production-grade data platforms.
Strong expertise in distributed processing with Apache Spark and Trino, including performance tuning and optimization.
Proven experience designing and maintaining reliable data pipelines using Airbyte and Apache Airflow.
Solid understanding of modern Lakehouse architectures, dbt-based transformations, and Iceberg table management.
Hands-on experience with Kubernetes, containerized deployments, and CI/CD for data workloads.
Knowledge of data security, governance, and access control practices (row/column-level security, policy enforcement, IAM).
Experience supporting BI/analytics use cases and integrating with enterprise reporting tools (e.g., Microsoft Fabric, Power BI).
Strong troubleshooting, system reliability, and problem-solving skills in complex distributed environments.
Ability to work independently in a highly technical, ownership-driven role.
Why you'll love working here
**easy going and friendly environment
**Remuneration Package:
Working hours: 5 days per week (in a professional and yet young, dynamic environment);
Salary: Competitive remuneration package (based on skills and experience);
12 leave days per year
Thông tin chung
- Thu nhập: Thỏa thuận
Việc làm tương tự khác
CÔNG TY CỔ PHẦN GIẢI PHÁP THANH TOÁN VIỆT NAM (VNPAY)
Hồ Chí Minh
Thỏa thuận
Công ty Cổ phần SmartOSC
Hà Nội, Hồ Chí Minh, Đà Nẵng
30 - 56 triệu
IMIP Technology And Solution Consultancy JSC
Xem trang công ty- Địa chỉ công ty: Tầng 13 Tòa nhà CMC - Duy Tân - Cầu Giấy - Hà Nội
- Quy mô: Từ 26 - 100 nhân viên
- Lĩnh vực: Tư vấn/ Chăm sóc khách hàng, IT phần mềm
Thông tin công việc
Vị trí:
Nhân viên
Hình thức làm việc:
Toàn thời gian
Việc làm tương tự
Cảnh báo dấu hiệu lừa đảo tuyển dụng
Đội ngũ hỗ trợ của JobOKO sẵn sàng đồng hành, tư vấn và giới thiệu những cơ hội việc làm phù hợp, giúp Ứng viên tự tin phát triển sự nghiệp và chinh phục mục tiêu nghề nghiệp bền vững.
Hotline CSKH
1900.63.63.84
Công ty Cổ phần JobOKO Toàn cầu
Đội ngũ hỗ trợ của JobOKO luôn chủ động tư vấn các giải pháp tuyển dụng tối ưu, cam kết đồng hành và hỗ trợ Quý Nhà tuyển dụng đạt được hiệu quả tuyển dụng bền vững.
Hotline CSKH
0962.107.888
Công ty Cổ phần JobOKO Toàn cầu