JobHabor

Sr. Data Engineer - Spark

Techsa
Location
Cairo, Cairo, Egypt
Workplace
Employment
Full Time
Salary
Apply on the employer’s site

Posted 1mo ago

Building high-scale batch and near-real-time data pipelines deployed on infrastructure we run ourselves (on-prem), not managed cloud services. You will design and operate high-volume analytical data systems end to end, with Apache Spark as the core processing engine for both batch and streaming workloads.

Requirements

  • 7+ years of experience in data engineering and software development
  • Ability to write high-quality code in Java/Scala, Python, or equivalent languages
  • Deep, hands-on production experience with Apache Spark — batch and Spark Structured Streaming (core requirement)

Demonstrated Spark performance tuning

partitioning, caching and persistence, broadcast joins, shuffle reduction, data-skew handling, and Adaptive Query Execution

  • Experience operating Spark on self-managed clusters (YARN, Kubernetes, or standalone) — executor sizing, resource allocation, and multi-tenant workloads
  • Practical experience with Kafka (or equivalent messaging systems) as a Spark source and sink for high-volume workloads, including offset and checkpoint management
  • Practical experience with distributed query engines (e.g., Trino/Presto or similar)
  • Practical experience with ETL / data integration tools, commercial or open-source (e.g., Datastage, Informatica, Apache NiFi, or similar)
  • Practical experience with SQL-based transformation frameworks (e.g., dbt or others)
  • Strong SQL skills and understanding of data modeling and data warehousing for analytical workloads
  • Hands-on experience with real-time / low-latency analytical stores (columnar or OLAP engines, e.g., Apache Pinot/ClickHouse or similar)
  • Practical experience with big-data platforms and distributions (e.g., Cloudera, Hadoop ecosystem, Databricks, or similar)
  • Practical experience containerizing and operating data workloads (Docker; Kubernetes a plus)
  • Experience with workflow orchestration tools (e.g., Airflow or similar)
  • Familiarity with data lake table formats (e.g., Apache Iceberg, Delta Lake, or similar), including schema evolution and compaction
  • Familiarity with data governance / cataloging tools (e.g., DataHub or similar)
  • Familiarity with lakehouse management systems (e.g., Apache Amoro or similar)
  • Familiarity using AI tools for development and debugging (Claude, Cursor, Codex)

Skills

  • Spark
  • Java
  • Scala
  • Python
  • Kubernetes
  • Kafka
  • Trino
  • Presto
  • ETL
  • Informatica
  • NiFi
  • SQL
  • dbt
  • Apache
  • ClickHouse
  • Hadoop
  • Databricks
  • Docker
  • Airflow
  • Apache Iceberg
  • Delta Lake
  • Anthropic Claude
  • Cursor
  • Codex

More jobs at Techsa

All 11

Similar roles