Principal Software Engineer, AI Infra Management and Ops
Microsoft
- Location
- Location not stated
- Workplace
- —
- Employment
- Full Time
- Salary
- —
Posted 2d ago
Lead the design, development, and evolution of cloud-native services, distributed systems, and platform capabilities that support large-scale production environments. Collaborate across engineering organizations to help define technical direction and architecture for platform infrastructure, cloud services, networking, storage, compute, and operational excellence investments. Guide the development of platform engineering capabilities that improve deployment automation, service lifecycle management, observability, reliability, scalability, and developer productivity. Advance software engineering practices including Continuous Integration and Continuous Delivery (CI/CD), Infrastructure as Code (IaC), testing, security, monitoring, and operational readiness throughout the engineering lifecycle. Contribute to the design and operation of Artificial Intelligence (AI) infrastructure, High Performance Computing (HPC) platforms, bare-metal infrastructure, graphics processing unit (GPU) environments, and large-scale distributed computing systems. Collaborate on datacenter architecture and networking solutions, including rack-scale systems, network topology design, Ethernet fabrics, InfiniBand fabrics, Remote Direct Memory Access (RDMA), Smart Network Interface Cards (SmartNICs), Data Processing Units (DPUs), and accelerated computing platforms. Foster a culture of technical excellence through architecture collaboration, mentorship, knowledge sharing, design review participation, and support for engineering excellence across the broader organization. Bachelor's Degree in Computer Science or related technical field AND 6+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR equivalent experience. Recognized for delivering technical solutions that span multiple engineering teams, organizations, or business areas. Demonstrated ability to simplify complex technical challenges and translate them into practical, scalable, and maintainable solutions. Proven ability to influence technical direction through collaboration, architectural leadership, technical mentoring, and cross-organizational partnerships. Knowledge of large-scale distributed systems, cloud-native platforms, or infrastructure services supporting business-critical production environments. Familiarity with Artificial Intelligence (AI) infrastructure, High Performance Computing (HPC) environments, accelerated computing platforms, or large-scale graphics processing unit (GPU) deployments. Knowledge of datacenter architecture, rack-scale infrastructure, network topology design, and modern datacenter networking technologies. Familiarity with InfiniBand fabrics, NVLink, NVSwitch, Ethernet networking, Remote Direct Memory Access (RDMA), Smart Network Interface Cards (SmartNICs), or Data Processing Units (DPUs) supporting large-scale compute environments. Understanding of bare-metal infrastructure, hardware lifecycle management, fleet operations, infrastructure telemetry, or platform operations at scale. Proficiency in one or more modern programming languages such as Go, Rust, C#, Java, or Python.
Skills
- RDMA
- C
- C++
- C#
- Java
- JavaScript
- Python
- Go
- Rust
Similar roles
Senior Software Engineer
Redwood Materials · remote · Nevada · USD 180,000–237,500/yr · today
Embedded Software Engineer – Power Electronics, Energy Storage
Redwood Materials · San Francisco, California, United States · USD 180,000–237,500/yr · today
Senior Full-Stack Engineer (Go Strength)
Sharesource Australia · Hanoi, Vietnam · today
Forward Deployed Engineer
DevRev · India · today
Software Engineer - ML/Computer Vision (Battery Sorting)
Redwood Materials · McCarran, NV · San Francisco, California, United States · USD 152,500–200,000/yr · today
Full-Stack Developer (Mental Health Care)
Coherent Solutions · Bulgaria · today