Data Engineer
Data Engineer
- Owns the data pipelines that move data from operational source systems onto the data platform — extraction, ingestion, transformation, orchestration, and the day-to-day operational health of those pipelines.
- Also plays a role as a hands-on builder of Foundational Data Products (raw → record-of-truth) and potentially Derivative Data Products (composed from upstream products).
Core skills
- SQL & Python (table stakes); Scala/Java for high-throughput streaming
- Pipeline build & operations: design, develop, deploy, monitor, and remediate batch and streaming pipelines that land source data on the platform; manage backfills, replays, late-arriving data, schema drift, and SLA breaches
- Ingestion patterns: full-load, incremental, change-data-capture (Debezium, Fivetran, Qlik Replicate, GoldenGate), event-driven ingest, API and file-based intake
- Pipeline frameworks: dbt, Apache Spark, Apache Beam, Airflow, Dagster, Prefect
- Streaming: Kafka, Kinesis, Flink, Spark Structured Streaming
- Data product packaging: schema contracts (Avro/Protobuf/JSON Schema), versioning, SLAs/SLOs, data contracts, output-port design (SQL, file, API, event)
- Storage formats: Iceberg, Delta Lake, Hudi (open table formats are now the mesh default)
- Quality & observability: Great Expectations, Soda, Monte Carlo, dbt tests, data lineage (OpenLineage)
- CI/CD for data: GitOps pipelines, unit + integration tests, environment promotion
- AI-adjacent: vector store ingestion (pgvector, Pinecone), feature stores (Feast, Tecton), RAG-ready chunking and embedding pipelines
- Mesh-specific: knowing when to build a derivative product vs. extending an existing one; consuming upstream products through governed input ports rather than reaching into source systems
Data Engineer
- Owns the data pipelines that move data from operational source systems onto the data platform — extraction, ingestion, transformation, orchestration, and the day-to-day operational health of those pipelines.
- Also plays a role as a hands-on builder of Foundational Data Products (raw → record-of-truth) and potentially Derivative Data Products (composed from upstream products).
Core skills
- SQL & Python (table stakes); Scala/Java for high-throughput streaming
- Pipeline build & operations: design, develop, deploy, monitor, and remediate batch and streaming pipelines that land source data on the platform; manage backfills, replays, late-arriving data, schema drift, and SLA breaches
- Ingestion patterns: full-load, incremental, change-data-capture (Debezium, Fivetran, Qlik Replicate, GoldenGate), event-driven ingest, API and file-based intake
- Pipeline frameworks: dbt, Apache Spark, Apache Beam, Airflow, Dagster, Prefect
- Streaming: Kafka, Kinesis, Flink, Spark Structured Streaming
- Data product packaging: schema contracts (Avro/Protobuf/JSON Schema), versioning, SLAs/SLOs, data contracts, output-port design (SQL, file, API, event)
- Storage formats: Iceberg, Delta Lake, Hudi (open table formats are now the mesh default)
- Quality & observability: Great Expectations, Soda, Monte Carlo, dbt tests, data lineage (OpenLineage)
- CI/CD for data: GitOps pipelines, unit + integration tests, environment promotion
- AI-adjacent: vector store ingestion (pgvector, Pinecone), feature stores (Feast, Tecton), RAG-ready chunking and embedding pipelines
- Mesh-specific: knowing when to build a derivative product vs. extending an existing one; consuming upstream products through governed input ports rather than reaching into source systems