JobHabor

Remote | Data Engineer — $140,000–$180,000/year

24-MAG
Location
New York · United States
Workplace
Remote
Employment
Full Time
Salary
USD 140,000–180,000/yr
Apply on the employer’s site

Posted today

We are sharing a full-time opportunity for an experienced Data Engineer with strong expertise in Python, SQL, ETL pipelines, data modelling, exploratory data analysis, and scalable data-processing workflows to support research, analytics, and AI/ML initiatives.

The role focuses on transforming structured and unstructured information into reliable, production-quality datasets. The successful candidate will build and maintain scalable data pipelines, optimise SQL and data-processing workflows, improve data quality and validation systems, and collaborate with researchers, data scientists, and engineers supporting AI and machine-learning development.

Key Responsibilities

ETL Pipeline Development

  • Design, develop, and maintain scalable ETL pipelines
  • Build reliable workflows for data extraction, transformation, and loading
  • Improve pipeline efficiency, maintainability, and scalability
  • Automate recurring data-processing tasks
  • Monitor pipelines and troubleshoot operational failures

Data Collection & Transformation

  • Collect data from structured and unstructured sources
  • Clean, normalise, transform, and organise raw datasets
  • Develop repeatable transformation workflows
  • Identify inconsistencies and incomplete records
  • Prepare reliable datasets for downstream research and analytics

Exploratory Data Analysis

  • Conduct exploratory analysis across complex datasets
  • Identify patterns, trends, anomalies, and quality issues
  • Investigate unexpected behaviour in source data
  • Produce analytical summaries supporting technical decisions
  • Communicate findings clearly to interdisciplinary stakeholders

SQL & Database Engineering

  • Write and optimise SQL queries for data extraction and transformation
  • Work with relational database systems such as PostgreSQL and MySQL
  • Design and maintain database schemas
  • Improve query performance and data-access patterns
  • Support reliable and maintainable database workflows

Data Modelling & Architecture

  • Develop logical and physical data models
  • Define database structures appropriate to analytical and operational requirements
  • Evaluate schema design and transformation strategies
  • Support scalable data-storage solutions
  • Maintain consistency across related datasets and systems

Data Quality & Validation

  • Develop automated data-validation workflows
  • Monitor accuracy, integrity, consistency, and completeness
  • Identify and resolve data-quality issues
  • Establish repeatable quality-control procedures
  • Ensure datasets meet downstream analytical and modelling requirements

Python Data Processing

  • Build data-processing workflows in Python
  • Use Pandas and NumPy for transformation and analysis
  • Develop reusable data-processing components
  • Improve performance and reliability of analytical workflows
  • Support automation of repetitive data-engineering processes

AI & Machine-Learning Data Support

  • Prepare datasets for AI and machine-learning initiatives
  • Collaborate with data scientists and researchers on data requirements
  • Support model-development workflows through reliable data preparation
  • Evaluate data suitability for training and evaluation use cases
  • Help structure data pipelines supporting AI/ML experimentation

Automation & Reporting

  • Automate recurring reporting and data-processing activities
  • Build repeatable validation and monitoring processes
  • Reduce manual intervention across routine data workflows
  • Improve operational visibility into pipeline health
  • Support timely delivery of high-quality datasets

Technical Documentation & Troubleshooting

  • Document pipelines, schemas, workflows, and technical decisions
  • Maintain clear operational and development documentation
  • Investigate and resolve pipeline failures
  • Diagnose data-related technical issues
  • Communicate root causes and remediation steps clearly

Ideal Profile

  • Strong proficiency in Python
  • Strong proficiency in SQL
  • Hands-on experience designing and maintaining ETL pipelines
  • Experience conducting exploratory data analysis
  • Proficiency with Pandas and NumPy
  • Experience with PostgreSQL and MySQL
  • Strong understanding of data modelling and database schemas
  • Experience working with structured and unstructured datasets
  • Demonstrated ability to maintain data quality, integrity, and reliability
  • Familiarity with Jupyter Notebook, VS Code, PyCharm, or comparable development environments
  • Strong analytical and problem-solving skills
  • Strong written and verbal communication skills
  • Experience collaborating with researchers, data scientists, or engineering teams is valuable
  • Exposure to AI and machine-learning workflows is advantageous
  • Familiarity with scikit-learn is beneficial
  • Experience with Hugging Face Transformers is a plus
  • Familiarity with AI APIs or comparable AI platforms is advantageous
  • Experience preparing datasets for AI/ML model development is highly valuable

Engagement Details

  • Full-time engagement
  • Fully remote
  • Compensation: $140,000–$180,000/year
  • Work will involve Python, SQL, ETL, exploratory data analysis, database design, data modelling, data-quality assurance, automation, and AI/ML dataset preparation
  • Strong data-engineering fundamentals and analytical judgement are central to this role
  • Responsibilities may involve both structured and unstructured data
  • Regular collaboration with researchers, data scientists, and engineering teams is expected
  • Project scope, data sources, technical requirements, and priorities may evolve based on business and research needs
  • Work must be completed without using confidential, proprietary, regulated, client-identifiable, restricted-access, or otherwise protected information belonging to any employer, client, research institution, data provider, or other third party

About the Platform

This opportunity is available through 24-MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams.

By submitting this application, you acknowledge that your information may be processed by 24-MAG LLC for recruitment and opportunity matching in accordance with our Privacy Policy: https://www.24-mag.com/privacy-policy

Skills

  • Python
  • SQL
  • ETL
  • PostgreSQL
  • MySQL
  • Pandas
  • NumPy
  • Visual Studio Code
  • scikit-learn
  • Hugging Face Transformers

More jobs at 24-MAG

All 96

Similar roles