Big Data Lead
Hexaware- Location
- United States
- Workplace
- —
- Employment
- —
- Salary
- —
Posted 2mo ago
Responsibilities
Development and Maintain Data Pipelines
Design, implement, and optimize end-to-end ETL/ELT pipelines for ingesting, processing, and transforming large volumes of structured and unstructured data.
Utilize Python and Pyspark
Write efficient, scalable and maintainable code in Python and leverage Pyspark for large-scale data processing in distributed computing environments. Also be able to review existing code and identify areas of improvement.
Ensure Data Quality and Integrity
Implement data validation, cleansing, transformation and reconciliation processes to ensure data accuracy and consistency throughout the data lifecycle.
Collaborate with Stakeholders
Work closely with IT teams and business stakeholders to gather data requirements and translate them to technical solutions.
Troubleshoot and Optimize
Monitor job performance, troubleshoot complex data issues and fine-tune for performance and scalability.
Adhere to Best Practices
Participate in code reviews, establish coding standards, and implement CI/CD pipelines for automated testing and deployment.
Skills
- Strong hands-on coding proficiency in Python, Pyspark and SQL (Microsoft SQL Server preferred)
- Experience with big data frameworks (Hadoop, Spark).
- Experience with cloud platforms ( AWS, Azure or GCP)
- Experience with Code versioning tools ( Bitbucket, Github )
- Experience with CI/CD and setting up pipelines.
- Solid understanding of database design principles, data modelling, schemas and data warehousing solutions.
- Excellent problem-solving and analytical skills to troubleshoot complex data issues independently.
Skills
- Pyspark Lead