JobHabor

Principal Software Engineer- AI Frameworks

Microsoft

Location
United States
Workplace
Employment
Full Time
Salary
Apply on the employer’s site

Posted 2d ago

Define technical vision, architecture, and multi-release strategy for critical AI framework, performance, benchmarking, or developer-productivity capabilities. Lead ambiguous, cross-stack investigations and investments spanning models, frameworks, compilers, runtimes, systems, services, and silicon. Establish common measurement, automation, observability, and engineering mechanisms that turn one-off analyses into scalable platform capabilities. Drive measurable improvements in model onboarding velocity, runtime performance, reliability, hardware utilization, and Azure capacity efficiency. Influence architecture and priorities across teams; align researchers, product groups, infrastructure owners, and hardware partners around clear decisions and execution plans. Provide hands-on technical leadership through prototypes, critical-path implementation, design and code reviews, complex debugging, and operational readiness. Raise the engineering bar by mentoring junior engineers, developing technical leaders, and advancing standards for quality, maintainability, and inclusive collaboration. Bachelor's Degree in Computer Science or a related technical field and 6+ years of technical engineering experience coding in languages such as C++, or Python, or equivalent experience. Deep expertise in GPU or equivalent accelerator programming, compilation, and low-level execution, including intermediate representations, lowering, code generation, instruction-level behavior, memory hierarchy, and synchronization. Expertise in parallelism strategies used in LLM training and inference, including tensor, pipeline, data, and expert parallelism, with the ability to evaluate their suitability for different models and hardware configurations. Expertise in distributed inference acceleration, including prefill/decode disaggregation, KV-cache transfer, collective communication, and compute/communication overlap. Strong understanding of LLM serving architectures such as vLLM, SGLang, or equivalent systems, with a track record of translating architectural improvements into measurable production gains. Demonstrated leadership of cross-team technical initiatives from strategy and design through implementation, deployment, and sustained production impact. A track record of creating reusable platforms, influencing stakeholders, and mentoring engineers. Ability to use AI-assisted development tools effectively and establish practices that improve engineering productivity without compromising correctness, performance, or maintainability.

Skills

  • Azure
  • C++
  • Python
  • LLM
  • vLLM

Similar roles