Top 3 reasons to join us
Attractive salary
Good working environment
Flexible management
Job description
Role Overview
We are seeking a Middle Data Engineer to build and scale our Production Datahouse on-premise environment. This is a purely technical, hands-on role focused on implementing a high-performance, on-premise data architecture.
You will own the end-to-end technical execution, managing the entire lifecycle of data as it moves from landing ingestion to the final serving layer. This role is not about high-level theory; it is about the hard engineering required to optimize on-premise hardware, tune distributed processing engines, and ensure a seamless data flow across our internal ecosystem.
Key Responsibilities
Data Pipeline Engineering: Design, build, and maintain automated data ingestion pipelines from several source systems using Airbyte and Airflow, ensuring reliability and scalability.
Lakehouse Architecture Implementation: Develop and manage the Medallion Architecture (Bronze/Silver/Gold) using dbt-spark to enable structured, high-quality analytical datasets.
Platform Performance Optimization: Optimize Spark and Trino workloads on Kubernetes and manage the full lifecycle of Apache Iceberg tables (compaction, snapshotting, and maintenance) to improve query performance, storage and reduce latency.
Infrastructure & Reliability Management: Proactively monitor and maintain Data Platform resources using Grafana and Prometheus; analyze query patterns to enhance system stability and performance for production users.
Data Security & Governance: Enforce row- and column-level security using Open Policy Agent (OPA) and develop foundational frameworks for Data Governance and Data Quality, embedding lineage, automated testing, and compliance directly into pipelines.
Enterprise Data Integration: Architect and implement high-performance data bridges to deliver on-premise Lakehouse data into Microsoft Fabric, enabling seamless and secure hybrid-cloud connectivity.
Your skills and experience
3+ years of hands-on experience in Data Engineering or Platform Engineering, building and operating production-grade data platforms.
Strong expertise in distributed processing with Apache Spark and Trino, including performance tuning and optimization.
Proven experience designing and maintaining reliable data pipelines using Airbyte and Apache Airflow.
Solid understanding of modern Lakehouse architectures, dbt-based transformations, and Iceberg table management.
Hands-on experience with Kubernetes, containerized deployments, and CI/CD for data workloads.
Knowledge of data security, governance, and access control practices (row/column-level security, policy enforcement, IAM).
Experience supporting BI/analytics use cases and integrating with enterprise reporting tools (e.g., Microsoft Fabric, Power BI).
Strong troubleshooting, system reliability, and problem-solving skills in complex distributed environments.
Ability to work independently in a highly technical, ownership-driven role.
Why you'll love working here
**easy going and friendly environment
**Remuneration Package:
Working hours: 5 days per week (in a professional and yet young, dynamic environment);
Salary: Competitive remuneration package (based on skills and experience);
12 leave days per year
Thông tin chung