Top 3 Reasons To Join Us
Disruptive innovations
People-oriented philosophy of doing business
Your footprint on regional on-demand market
The Job
- Build and operate batch and near-real-time data pipelines on our lakehouse platform,
powering reporting, analytics, and data products across the company. - Own assigned data domains end to end - from source ingestion to the trusted,
documented tables that business and technical teams rely on. - Design and maintain data models that turn raw operational data into reusable,
well-structured datasets instead of one-off queries. - Improve pipeline performance, reliability, and infrastructure cost as data volume and
business complexity grow. - Establish and maintain data quality standards, including validation, freshness monitoring,
and alerting on critical datasets. - Keep metadata, lineage, and documentation up to date so data consumers can find and
trust the data they use. - Support BI, Analytics, and Data Science teams by preparing serving datasets and
resolving data issues they raise. - Partner with Backend and Product teams to define event and CDC data contracts for
new and existing services. - Participate in the team's on-call rotation; investigate incidents, restore data SLAs, and
drive follow-up improvements. - Contribute to code review, technical documentation, and the team's deployment and
engineering practices.
Your Skills and Experience
- 2+ years of hands-on experience in Data Engineering, or an equivalent role with
ownership of production data pipelines. - Strong SQL, including window functions, complex CTEs, and query optimization on
tables with hundreds of millions of rows. - Production-level Python: structured, tested, and packaged code - not only notebooks or
one-off scripts. - Hands-on Spark experience (PySpark or Scala): partitioning, shuffle, join strategies, data
skew, and the ability to read and interpret execution plans. - Practical experience with a workflow orchestrator; Airflow strongly preferred.
- Solid understanding of data warehouse and lakehouse fundamentals: dimensional
modeling, incremental loads, CDC handling, snapshot vs delta processing. - Comfortable with Git and CI/CD workflows as part of daily development.
- Working knowledge of Docker and Kubernetes: able to containerize a job, read pod logs
and events, understand resource requests and limits, and debug a failing or evicted pod. - Strong debugging discipline: reads logs and metrics, isolates variables, and verifies
hypotheses with data instead of guessing. - Able to read English technical documentation and open-source code independently.
Nice to have
- Production experience with Apache Iceberg, Delta Lake, or Hudi, including snapshot,
manifest, and metadata layer concepts. - Experience with Trino/Presto, or OLAP engines such as Apache Doris, ClickHouse, or
StarRocks. - Deeper Kubernetes experience: Helm, ArgoCD/GitOps, and resource tuning for Spark
executors. - Streaming experience with Spark Structured Streaming or Flink, including state,
watermark, and exactly-once semantics. - Familiarity with data catalog and quality tooling such as DataHub, OpenMetadata, dbt, or
Great Expectations. - Experience with Vault, External Secrets Operator, Terraform, or Terragrunt.
- Exposure to logistics, marketplace, mobility, e-commerce, or on-demand platform data.
- Experience with billing, reconciliation, or finance data pipelines.
Why You'll Love Working Here
- Physical Wellbeing Benefit: General Insurance, Medical check-up, Accident Insurance, Healthcare Insurance.
- Emotional Wellbeing Benefit: Company Trip, Year End Party, Aha Hour Activities, Special Day Gifts, Aha Club (Badminton, Soccer).
- Financial Wellbeing Benefit: Grab/Be For Work (Tech/Lead Level), Workplace Relocation, 13th Month Salary, PP Appreciate, Annual Leave Remain.