JobHabor

Senior Data Engineer

Amgen
Location
India - Hyderabad
Workplace
Employment
Salary
Apply on the employer’s site

Posted 11d ago

Career Category

Engineering

Job Description

Senior Data Engineer

Roles & Responsibilities

  • Design, develop, and maintain scalable databricks pipelines to support structured, semi-structured, and unstructured data processing across the Enterprise Data Fabric.
  • Implement real-time and batch data processing solutions, integrating data from multiple sources into a unified, governed data fabric architecture.
  • Optimize big data processing frameworks using Apache Spark to ensure high availability and cost efficiency.
  • Work with metadata management and data lineage tracking tools to enable enterprise-wide data discovery and governance.
  • Ensure data security, compliance, and role-based access control (RBAC) across data environments.
  • Optimize query performance, indexing strategies, partitioning, and caching for large-scale data sets.
  • Develop CI/CD pipelines for automated data pipeline deployments, version control, and monitoring.
  • Implement data virtualization techniques to provide seamless access to data across multiple storage systems.
  • Collaborate with cross-functional teams, including data architects, business analysts, and DevOps teams, to align data engineering strategies with enterprise goals.
  • Stay up to date with emerging data technologies and best practices, ensuring continuous improvement of Enterprise Data Fabric architectures.

Must-Have Skills

  • Develop pipelines using Databricks (Delta Lake, Spark, notebooks)
  • Hands-on experience in data engineering technologies such as Databricks PySpark, SQL, and Scaled Agile methodologies.
  • Proficiency in workflow orchestration, performance tuning on big data processing.
  • Strong Programming skills and lead teams with technical acumen
  • Experience with Data Fabric, Data Mesh, or similar enterprise-wide data architectures.
  • Ability to quickly learn, adapt and apply new technologies
  • Strong problem-solving and analytical skills
  • Excellent communication and teamwork skills
  • Experience with Scaled Agile Framework (SAFe), Agile delivery practices, and DevOps practices.

Preferred / Strategic Skills (Aligned to Future Data Strategy)

  • Certification:
  • Relevant certifications in Databricks, cloud platforms (AWS/Azure/GCP), or modern data engineering technologies are a plus
  • Experience with:
  • Delta Lake / Lakehouse architectures
  • Data Fabric / Data Mesh concepts
  • Familiarity with:
  • Streaming data (Kafka, event-driven pipelines)
  • Data orchestration tools (Airflow, Databricks Workflows)
  • Exposure to:
  • AI/ML data pipelines and feature engineering
  • Unstructured data processing
  • Understanding of:
  • Data governance frameworks and cataloging tools
  • Security and privacy controls for sensitive data

Good-to-Have Skills

  • Good to have deep expertise in Biotech & Pharma industries
  • Experience in writing APIs to make the data available to the consumers
  • Experienced with SQL/NOSQL database, vector database for large language models
  • Experienced with data modeling and performance tuning for both OLAP and OLTP databases
  • Experienced with software engineering best-practices, including but not limited to version control (Git, Subversion, etc.), CI/CD (Jenkins, Maven etc.), automated unit testing, and Dev Ops

Education and Professional Certifications

  • 8 to 12 years of Computer Science, IT or related field experience
  • Databricks Certificate preferred

Soft Skills

  • Excellent analytical and troubleshooting skills.
  • Strong verbal and written communication skills
  • Ability to work effectively with global, virtual teams
  • High degree of initiative and self-motivation.
  • Ability to manage multiple priorities successfully.
  • Team-oriented, with a focus on achieving team goals.
  • Ability to learn quickly, be organized and detail oriented.
  • Strong presentation and public speaking skills.

.

Skills

  • Databricks
  • Spark
  • RBAC
  • Delta Lake
  • PySpark
  • SQL
  • AWS
  • Azure
  • GCP
  • Kafka
  • Airflow
  • Vector Databases
  • LLM
  • Git
  • Jenkins
  • Maven

More jobs at Amgen

All 125

Similar roles