JobHabor

Associate Process Manager

eClerx
Location
Pune, Maharashtra, India
Workplace
Employment
Salary
Apply on the employer’s site

Posted 11d ago

Job Description – GCP Data Engineer

Role Overview

We are looking for an experienced GCP Data Engineer to design, develop, and maintain scalable data pipelines and ETL/ELT frameworks on the Google Cloud Platform. The candidate will work extensively with BigQuery, Python/PySpark, Airflow/Cloud Composer, and Google Cloud Storage to process and transform large volumes of campaign, customer, and clickstream data.

The ideal candidate should have strong hands-on experience in data engineering, excellent SQL and Python skills, and a good understanding of cloud-based data pipeline architecture. The role will involve building reliable production pipelines, optimizing BigQuery performance and costs, implementing automation, and supporting critical data workflows.

Key Responsibilities

Data Pipeline Development

  • Design, develop, and maintain scalable and reliable data pipelines on GCP.
  • Build batch and, where required, near-real-time data ingestion and transformation pipelines.
  • Develop robust ETL/ELT frameworks for large volumes of structured and semi-structured data.
  • Ingest data from multiple sources into Google Cloud Storage and BigQuery.
  • Develop reusable and modular data engineering components.
  • Implement appropriate error handling, retry mechanisms, logging, and data validation.

Required Technical Skills

Must Have

  • 4–8 years of experience in Data Engineering.
  • Strong hands-on experience with GCP.
  • Strong proficiency in SQL, preferably extensive experience with BigQuery SQL.
  • Strong hands-on experience in Python and/or PySpark.
  • Experience developing ETL/ELT data pipelines.
  • Hands-on experience with BigQuery.
  • Hands-on experience with Google Cloud Storage (GCS).
  • Experience with Apache Airflow / Cloud Composer.
  • Good understanding of data pipeline architecture and data engineering best practices.
  • Experience with Git and CI/CD practices.
  • Exposure to Jenkins or similar CI/CD tools.
  • Experience in production support, monitoring, troubleshooting, and performance optimization.

Good to Have

  • Experience working with campaign, marketing, customer, or clickstream data.
  • Experience with large-scale data processing.
  • Knowledge of Dataflow / Apache Beam.
  • Knowledge of Pub/Sub and event-driven architectures.
  • Experience with real-time or streaming data pipelines.
  • Knowledge of data warehousing and dimensional data modeling.
  • Experience with data quality frameworks and validation.
  • Knowledge of GCP IAM and security concepts.
  • Experience with Cloud Monitoring / Logging.
  • Experience working in Agile/Scrum environments.

Candidate Profile

The ideal candidate should

  • Have strong hands-on technical expertise rather than only theoretical knowledge.
  • Be comfortable writing complex SQL and Python/PySpark code.
  • Have experience independently designing and developing data pipelines.
  • Understand how to build scalable and cost-efficient solutions on GCP.
  • Be capable of troubleshooting production issues and taking ownership until resolution.
  • Have good analytical and problem-solving skills.
  • Be comfortable working with multiple stakeholders and managing delivery timelines.
  • Demonstrate good communication and documentation skills.

Skills

  • GCP
  • ETL
  • ELT
  • BigQuery
  • Python
  • PySpark
  • Airflow
  • Cloud Composer
  • Google Cloud Storage
  • SQL
  • Git
  • Jenkins
  • Dataflow
  • Beam
  • Pub/Sub
  • IAM
  • Cloud Monitoring

More jobs at eClerx

All 98

Similar roles