JobHabor

Process Manager

eClerx
Location
Pune, Maharashtra, India
Workplace
Employment
Salary
Apply on the employer’s site

Posted 8d ago

Job Description – GCP Data Engineer

Role Overview

We are looking for an experienced GCP Data Engineer to design, develop, and maintain scalable data pipelines and ETL/ELT frameworks on the Google Cloud Platform. The candidate will work extensively with BigQuery, Python/PySpark, Airflow/Cloud Composer, and Google Cloud Storage to process and transform large volumes of campaign, customer, and clickstream data.

The ideal candidate should have strong hands-on experience in data engineering, excellent SQL and Python skills, and a good understanding of cloud-based data pipeline architecture. The role will involve building reliable production pipelines, optimizing BigQuery performance and costs, implementing automation, and supporting critical data workflows.

Key Responsibilities

Data Pipeline Development

  • Design, develop, and maintain scalable and reliable data pipelines on GCP.
  • Build batch and, where required, near-real-time data ingestion and transformation pipelines.
  • Develop robust ETL/ELT frameworks for large volumes of structured and semi-structured data.
  • Ingest data from multiple sources into Google Cloud Storage and BigQuery.
  • Develop reusable and modular data engineering components.
  • Implement appropriate error handling, retry mechanisms, logging, and data validation.

BigQuery & Data Engineering

  • Develop complex and optimized SQL queries in BigQuery for data transformation and analysis.
  • Optimize BigQuery performance through:
  • Partitioning
  • Clustering
  • Query optimization
  • Efficient table design
  • Appropriate data types and storage strategies
  • Monitor and optimize BigQuery processing costs.
  • Design scalable data models suitable for large-scale campaign, customer, and clickstream datasets.
  • Troubleshoot data quality, performance, and pipeline-related issues.

Airflow / Cloud Composer

  • Develop and maintain DAGs using Apache Airflow / Cloud Composer.
  • Implement scheduling, dependency management, retries, failure handling, and alerting.
  • Monitor production workflows and proactively resolve failed or delayed jobs.
  • Build reusable operators and workflow components where required.
  • Ensure critical data pipelines meet agreed SLAs.

Python / PySpark

  • Develop data processing and transformation logic using Python and/or PySpark.
  • Write clean, reusable, scalable, and production-ready code.
  • Optimize PySpark jobs for performance and efficient resource utilization.
  • Implement appropriate testing and validation for data transformation processes.

GCP Services

Work extensively with GCP services, particularly

  • Google BigQuery
  • Google Cloud Storage (GCS)
  • Cloud Composer / Airflow

Exposure to additional GCP services such as Cloud Functions, Pub/Sub, Dataflow, Cloud Run, Secret Manager, IAM, or Cloud Monitoring would be an advantage.

Automation, CI/CD & DevOps

  • Automate deployment and execution of data engineering workflows.
  • Work with Git-based development and version-control workflows.
  • Implement and maintain CI/CD pipelines using tools such as Jenkins or equivalent.
  • Follow code review, branching, deployment, and release-management processes.
  • Implement monitoring, logging, and alerting for production data pipelines.

Production Support & Troubleshooting

  • Monitor daily data pipeline execution and resolve production issues within defined SLAs.
  • Investigate pipeline failures, data discrepancies, performance issues, and processing delays.
  • Perform root-cause analysis and implement permanent fixes.
  • Coordinate with application, analytics, infrastructure, and business teams to resolve data-related issues.
  • Participate in production deployments and provide post-deployment support.

Stakeholder Collaboration

  • Work closely with Data Analysts, Data Scientists, Product Teams, Business Stakeholders, and Technology Teams to understand data requirements.
  • Translate business requirements into scalable technical solutions.
  • Communicate technical issues, risks, dependencies, and delivery status effectively.
  • Participate in technical discussions, design reviews, and solution development.

Required Technical Skills

Must Have

  • 4–8 years of experience in Data Engineering.
  • Strong hands-on experience with GCP.
  • Strong proficiency in SQL, preferably extensive experience with BigQuery SQL.
  • Strong hands-on experience in Python and/or PySpark.
  • Experience developing ETL/ELT data pipelines.
  • Hands-on experience with BigQuery.
  • Hands-on experience with Google Cloud Storage (GCS).
  • Experience with Apache Airflow / Cloud Composer.
  • Good understanding of data pipeline architecture and data engineering best practices.
  • Experience with Git and CI/CD practices.
  • Exposure to Jenkins or similar CI/CD tools.
  • Experience in production support, monitoring, troubleshooting, and performance optimization.

Good to Have

  • Experience working with campaign, marketing, customer, or clickstream data.
  • Experience with large-scale data processing.
  • Knowledge of Dataflow / Apache Beam.
  • Knowledge of Pub/Sub and event-driven architectures.
  • Experience with real-time or streaming data pipelines.
  • Knowledge of data warehousing and dimensional data modeling.
  • Experience with data quality frameworks and validation.
  • Knowledge of GCP IAM and security concepts.
  • Experience with Cloud Monitoring / Logging.
  • Experience working in Agile/Scrum environments.

Candidate Profile

The ideal candidate should

  • Have strong hands-on technical expertise rather than only theoretical knowledge.
  • Be comfortable writing complex SQL and Python/PySpark code.
  • Have experience independently designing and developing data pipelines.
  • Understand how to build scalable and cost-efficient solutions on GCP.
  • Be capable of troubleshooting production issues and taking ownership until resolution.
  • Have good analytical and problem-solving skills.
  • Be comfortable working with multiple stakeholders and managing delivery timelines.
  • Demonstrate good communication and documentation skills.

Skills

  • GCP
  • ETL
  • ELT
  • BigQuery
  • Python
  • PySpark
  • Airflow
  • Cloud Composer
  • Google Cloud Storage
  • SQL
  • Cloud Functions
  • Pub/Sub
  • Dataflow
  • Cloud Run
  • IAM
  • Cloud Monitoring
  • Git
  • Jenkins
  • Beam

More jobs at eClerx

All 98