JobHabor

Data Engineer, AI & Distributed Systems

Zignal Labs
Location
United States
Workplace
Remote
Employment
Full Time
Salary
USD 120,000–140,000/yr
Apply on the employer’s site

Posted 1mo ago

The employer’s full description could not be read from their board. This is a summary of the posting — follow the apply link for the original.

Responsibilities

  • Develop and operate batch and streaming pipelines for high-volume unstructured data
  • Own pipeline components end to end from implementation through production monitoring
  • Build and extend data pathways for NLP, LLM, and retrieval services
  • Integrate with and tune search and vector stores for semantic search and real-time retrieval
  • Implement and improve microservices and APIs for enterprise analytics
  • Write clean, tested, and maintainable code
  • Participate in code review, CI/CD, and infrastructure-as-code practices
  • Debug production issues and improve system reliability
  • Collaborate with Data Science, ML, Product, and Security teams

Requirements

  • 3+ years building and operating data pipelines in production
  • Strong programming skills in Scala, Java, or Kotlin
  • Working proficiency in Python
  • Hands-on experience with Apache Spark
  • Hands-on experience with Kafka including consumer groups, offsets, and partitioning
  • Practical AWS experience
  • Comfort with Docker
  • Ability to work in a Kubernetes environment
  • Experience with workflow orchestrators like Airflow, Prefect, or Dagster
  • Solid SQL skills
  • Experience with at least one NoSQL or caching layer like Redis, MongoDB, or DynamoDB
  • Sound CS fundamentals in data structures and algorithms
  • Strong written communication skills for asynchronous work

Preferred

  • Experience with Databricks or Delta Lake
  • Experience with Flink or other stream-processing frameworks
  • Experience with vector databases like Pinecone, Qdrant, Milvus, or pgvector
  • Hands-on RAG or embedding pipeline work
  • Experience with Elasticsearch or OpenSearch
  • Experience parsing messy, unstructured, or multilingual text at scale
  • Deep database performance tuning
  • Experience with distributed consensus systems
  • Bachelor's degree in Computer Science, Engineering, or a related field

Skills

  • Scala
  • Java
  • Kotlin
  • Python
  • Apache Spark
  • Kafka
  • AWS
  • Docker
  • Kubernetes
  • Airflow
  • Prefect
  • Dagster
  • SQL
  • Redis
  • MongoDB
  • DynamoDB
  • Databricks
  • Delta Lake
  • Flink
  • Pinecone
  • Qdrant
  • Milvus
  • pgvector
  • Elasticsearch
  • OpenSearch

More jobs at Zignal Labs

Similar roles