Remote | Data Engineer — $140,000–$180,000/year
24-MAG- Location
- New York · United States
- Workplace
- Remote
- Employment
- Full Time
- Salary
- USD 140,000–180,000/yr
Posted today
We are sharing a full-time opportunity for an experienced Data Engineer with strong expertise in Python, SQL, ETL pipelines, data modelling, exploratory data analysis, and scalable data-processing workflows to support research, analytics, and AI/ML initiatives.
The role focuses on transforming structured and unstructured information into reliable, production-quality datasets. The successful candidate will build and maintain scalable data pipelines, optimise SQL and data-processing workflows, improve data quality and validation systems, and collaborate with researchers, data scientists, and engineers supporting AI and machine-learning development.
Key Responsibilities
ETL Pipeline Development
- Design, develop, and maintain scalable ETL pipelines
- Build reliable workflows for data extraction, transformation, and loading
- Improve pipeline efficiency, maintainability, and scalability
- Automate recurring data-processing tasks
- Monitor pipelines and troubleshoot operational failures
Data Collection & Transformation
- Collect data from structured and unstructured sources
- Clean, normalise, transform, and organise raw datasets
- Develop repeatable transformation workflows
- Identify inconsistencies and incomplete records
- Prepare reliable datasets for downstream research and analytics
Exploratory Data Analysis
- Conduct exploratory analysis across complex datasets
- Identify patterns, trends, anomalies, and quality issues
- Investigate unexpected behaviour in source data
- Produce analytical summaries supporting technical decisions
- Communicate findings clearly to interdisciplinary stakeholders
SQL & Database Engineering
- Write and optimise SQL queries for data extraction and transformation
- Work with relational database systems such as PostgreSQL and MySQL
- Design and maintain database schemas
- Improve query performance and data-access patterns
- Support reliable and maintainable database workflows
Data Modelling & Architecture
- Develop logical and physical data models
- Define database structures appropriate to analytical and operational requirements
- Evaluate schema design and transformation strategies
- Support scalable data-storage solutions
- Maintain consistency across related datasets and systems
Data Quality & Validation
- Develop automated data-validation workflows
- Monitor accuracy, integrity, consistency, and completeness
- Identify and resolve data-quality issues
- Establish repeatable quality-control procedures
- Ensure datasets meet downstream analytical and modelling requirements
Python Data Processing
- Build data-processing workflows in Python
- Use Pandas and NumPy for transformation and analysis
- Develop reusable data-processing components
- Improve performance and reliability of analytical workflows
- Support automation of repetitive data-engineering processes
AI & Machine-Learning Data Support
- Prepare datasets for AI and machine-learning initiatives
- Collaborate with data scientists and researchers on data requirements
- Support model-development workflows through reliable data preparation
- Evaluate data suitability for training and evaluation use cases
- Help structure data pipelines supporting AI/ML experimentation
Automation & Reporting
- Automate recurring reporting and data-processing activities
- Build repeatable validation and monitoring processes
- Reduce manual intervention across routine data workflows
- Improve operational visibility into pipeline health
- Support timely delivery of high-quality datasets
Technical Documentation & Troubleshooting
- Document pipelines, schemas, workflows, and technical decisions
- Maintain clear operational and development documentation
- Investigate and resolve pipeline failures
- Diagnose data-related technical issues
- Communicate root causes and remediation steps clearly
Ideal Profile
- Strong proficiency in Python
- Strong proficiency in SQL
- Hands-on experience designing and maintaining ETL pipelines
- Experience conducting exploratory data analysis
- Proficiency with Pandas and NumPy
- Experience with PostgreSQL and MySQL
- Strong understanding of data modelling and database schemas
- Experience working with structured and unstructured datasets
- Demonstrated ability to maintain data quality, integrity, and reliability
- Familiarity with Jupyter Notebook, VS Code, PyCharm, or comparable development environments
- Strong analytical and problem-solving skills
- Strong written and verbal communication skills
- Experience collaborating with researchers, data scientists, or engineering teams is valuable
- Exposure to AI and machine-learning workflows is advantageous
- Familiarity with scikit-learn is beneficial
- Experience with Hugging Face Transformers is a plus
- Familiarity with AI APIs or comparable AI platforms is advantageous
- Experience preparing datasets for AI/ML model development is highly valuable
Engagement Details
- Full-time engagement
- Fully remote
- Compensation: $140,000–$180,000/year
- Work will involve Python, SQL, ETL, exploratory data analysis, database design, data modelling, data-quality assurance, automation, and AI/ML dataset preparation
- Strong data-engineering fundamentals and analytical judgement are central to this role
- Responsibilities may involve both structured and unstructured data
- Regular collaboration with researchers, data scientists, and engineering teams is expected
- Project scope, data sources, technical requirements, and priorities may evolve based on business and research needs
- Work must be completed without using confidential, proprietary, regulated, client-identifiable, restricted-access, or otherwise protected information belonging to any employer, client, research institution, data provider, or other third party
About the Platform
This opportunity is available through 24-MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams.
By submitting this application, you acknowledge that your information may be processed by 24-MAG LLC for recruitment and opportunity matching in accordance with our Privacy Policy: https://www.24-mag.com/privacy-policy
Skills
- Python
- SQL
- ETL
- PostgreSQL
- MySQL
- Pandas
- NumPy
- Visual Studio Code
- scikit-learn
- Hugging Face Transformers
More jobs at 24-MAG
All 96Remote | Software Engineer - Open Source Contributions — $50–$100/hour
24-MAG · New York · United States · USD 50–100/hr · today
Remote | Sr. Full-Stack Software Engineer — $50–$100/hour
24-MAG · New York · United States · USD 50–100/hr · today
Remote | Open Source Contributor (GitHub) — $100–$150/hour
24-MAG · New York · United States · USD 100–150/hr · today
Remote | Senior Platform Engineer — $60–$130/hour
24-MAG · New York · United States · USD 60–130/hr · today
Remote | Human Data Manager — $40–$60/hour
24-MAG · New York · United States · USD 40–60/hr · today
Similar roles
Risk and Integrity Data Engineer
Evolution · Tbilisi, Tbilisi, Georgia · today
Staff Data Scientist, Imaging
Biohub · Redwood City, CA (Hybrid) · USD 214,000–294,800/yr · today
Senior/Staff Data Scientist
Render · SF · US/Canada · today
Senior Data Analyst, Avail
Realtor.com Careers · Austin, Texas, United States · today
Senior Analytics Engineer
Redwood Materials · San Francisco, California, United States · today
Data Engineering Manager
Redwood Materials · McCarran, NV · today