Job Description:
As an Experienced Data Engineer, you will take end-to-end ownership of assigned data pipelines - from data ingestion through to application consumption - with responsibility for testing and basic data quality control.
1. Build and Scale Data Pipelines
- Develop ingestion jobs to bring data from new sources, including REST APIs, XML/CSV/JSON files, and Kafka topics, into the lakehouse; handle incremental loading, checkpoints, and schema changes over time.
- Standardize, clean, and deduplicate data into a common schema; develop matching and data consolidation logic across sources, resolving conflicts when different sources describe the same entity in different ways.
- Build the serving layer for applications and reporting, including upserts to PostgreSQL/MongoDB, ensuring data is delivered on time and remains consistent with the source snapshots.
- Develop and maintain orchestration DAGs, including dependencies, retries, backfills, and scheduling.
2. Data Quality and Operations
- Define data expectations for assigned tables, including schema, constraints, volume, and freshness, and implement automated tests at layer boundaries; investigate and resolve data quality alerts.
- Monitor jobs, debug failures, and troubleshoot common performance issues such as data skew, small files, and OOM (Out of Memory); optimize queries and table structures.
- Write unit tests for transformation logic and follow the team's branching/commit conventions, code review practices, and CI processes.
3. Collaboration
- Work closely with Product, BI, and AI teams to understand data consumption requirements; maintain data documentation including metadata, lineage, and data dictionaries so that downstream systems and AI Agents can use data with the correct context and access permissions.
-------
As a Senior Data Engineer, you will design the data platform so that the team can onboard new data sources quickly without compromising quality. You will establish technical standards for pipelines, data models, and the semantic layer supporting analytics and AI Agents, as well as build data quality measurement frameworks to support release decisions.
You will be responsible for the reliability, performance, and security of the platform handling sensitive cybersecurity data, while providing technical leadership and mentoring to the team.
1. Architecture and Technical Standards
- Design an end-to-end lakehouse architecture for multiple tenants, including data layering, partitioning and table organization strategies, and data contracts between layers.
- Standardize the pipeline development framework, including job templates, resource profiles, flow registration mechanisms, and checkpoints, so that onboarding a new data source becomes a repeatable process rather than requiring a new design from scratch.
- Design data models and semantic layers, including consistent definitions of metrics, entities, and lineage, to support reporting, analytics, and AI Agent data queries.
2. Data Quality, Reliability, and Performance
- Build a data quality measurement framework that runs across the full dataset on every pipeline run, rather than relying only on sample-based validation; define metrics, thresholds, quality gates, alerts, and operational runbooks.
- Size resources based on actual measurements such as data volume, skew, and cardinality, rather than trial and error; address large-scale performance challenges including data skew, small files, compaction, snapshot expiration, and storage costs.
- Establish data governance, including data catalogs, lineage, sensitive data classification, and role-based access control.
- Integrate CI/CD, automated testing, and code review standards into the pipeline development process.
3. Technical Leadership and Collaboration
- Mentor Data Engineers on the team, conduct code and design reviews, and establish clear ownership across data domains.
- Evaluate and select technologies for the data platform architecture; collaborate with Product, AI, BI, and Infrastructure teams to bring data into the product.