Principal Platform Engineer - AI Ops, Backend, Reliability & Security
- Location
- Remote
- Workplace
- Remote
- Employment
- Full Time
- Salary
- —
Posted 27d ago
My IR - Administration - Career - View Job - Principal Platform Engineer - AI Ops, Backend, Reliability & Security Skip to main content
Page loading completed.
Menu
- Browse Jobs
- Application Disclosure
- Privacy
Sign in or Create an account
- Browse Jobs
- Application Disclosure
- Privacy
Sign in or Create an account
Principal Platform Engineer - AI Ops, Backend, Reliability & Security
Posted
02/08/2026
Closing Date
02/11/2026
Job Type
Permanent - Full Time
Location
Remote
Job Category
Information Technology
IR Labs is the innovation lab inside Integrated Research where small, cross‑functional squads chase outsized, industry‑defining opportunities. We operate like a funded startup—rapid sprints, bold experimentation, zero bureaucracy—backed by the global footprint and resources of a public company.
Our charter is simple
turn cutting‑edge AI research into products that customers can’t imagine working without. We target the hardest problems in software and then move fast to ship solutions that create 10x impact.
Our flagship is Agentic SQA - a software quality system that turns noisy signals (static analysis, fuzzers, sanitizers, CI output) into context-aware triage, validated actions, and developer-ready outputs. We’re in beta now, starting narrow on a focused C/C++ risk-analysis workflow targeting a specific bug class, and expanding from there into broader coverage, richer evidence, and stronger control layers. The stack combines LLVM/clang static analysis, a code knowledge graph, and agentic LLM workflows.
The direction is bigger than code review
a systems verification platform that closes the gap where human-speed review is breaking down.
If you thrive on autonomy, crave world‑class technical challenges, and want to see your ideas hit production quickly, IR Labs is your launch pad. Join us and help build the future, one breakthrough at a time.
Before you apply
Agentic SQA is live and free to use. Install and run it. Form a view. We’ll discuss that during the interviews.
Job Description
Who We’re Looking For
The founding Principal Platform Engineer will join a lean team and work closely with senior technical leaders to accelerate execution and deliver high-impact infrastructure to allow AI teams to focus on innovation. If you want to own the core platform that turns AI systems into a reliable, secure, and scalable product and build the backend services that power onboarding and operations, create streamlined “paved roads” for engineering, and productionize AI workloads with strong MLOps foundations, this role was written for you.
What You’ll Do
- Own the product platform, not just the infrastructure: build the core platform layer that turns our AI systems into a reliable, supportable product customers can onboard to and use daily.
- Ship critical backend systems: design and implement core services/APIs that power signup/sign-in, tenancy, permissions, billing/entitlements, provisioning, and internal workflows so customer onboarding is seamless and repeatable.
- Build “paved roads” that prevent engineering drag: create a standardized, self-serve path from feature → deploy → operate so teams can ship without repeated bespoke work, manual steps, or fragile runbooks.
- Keep the AI critical path clean: absorb platform/ops/MLOps work that would otherwise pull our AI Scientist/MLE into toil and customer support, protecting our highest-leverage innovation hours.
- Productionize AI workloads: build pragmatic MLOps and GPU operations foundations (serving/training workflows, model artifact distribution, utilization-aware scheduling, caching/cold-start strategies) so AI cost and latency remain controlled as we scale.
- Create reliability as a system property: define SLIs/SLOs and error budgets, then enforce them through release standards, testing discipline, rollout patterns, and incident practices.
- Embed security into the platform primitives: implement safe-by-default patterns for identity, access boundaries, secrets, and dependency hygiene so security is “how the platform works,” not a separate checklist.
- Design for resilience with cost discipline: architect multi-region strategies (backup/restore, DR, failover) that are as simple as possible while meeting availability goals.
- Make observability actionable: build high-signal metrics/logs/traces and operational dashboards that reduce noise and accelerate diagnosis; run blameless incident reviews that permanently reduce repeat failures.
- Be a force multiplier across a senior team: partner closely with data infra, compiler, and AI leadership - review designs, unblock execution, and ship high-impact work directly when needed.
Desired Skills and Experience
What You Bring to the Table
- Principal-level “builder/operator” experience: you’ve built and run production systems end-to-end - design, implementation, deployment, on-call reality, and iteration under customer pressure.
- Backend engineering strength: you can ship significant backend systems (APIs, workflows, auth boundaries, tenancy) and are comfortable owning core product services, not just infrastructure automation.
- Strong systems coding: deep proficiency in Go and/or Rust (Python acceptable as supporting), plus solid shell skills; you can deliver large, correct changes quickly with tests and quality gates.
- Kubernetes + cloud fluency: strong Kubernetes/EKS experience; you understand service networking, zero-downtime rollouts, autoscaling, and safe multi-tenant patterns. Comfortable with AWS primitives (VPC, IAM, EC2, S3, RDS) and multi-region architecture.
- AI ops / GPU awareness: experience operating GPU-backed workloads or adjacent high-performance compute. You understand utilization, scheduling tradeoffs, and practical cost/performance management in production.
- Reliability engineering mindset: you’ve used SLIs/SLOs/error budgets and know how to make reliability measurable and enforceable through engineering practice (not meetings).
- Security-by-design fundamentals: strong instincts and experience with identity/access patterns, secrets, dependency/supply-chain hygiene, and audit-friendly operational practices.
- Automation-first execution: track record of eliminating operational toil through automation across CI/CD, provisioning, testing, and operations.
- Customer-facing technical judgment: you can take customer feedback and translate it into platform capabilities that reduce friction, reduce support burden, and accelerate adoption.
- Clear communication in ambiguity: you write crisp design docs, make tradeoffs explicit, and collaborate effectively across ML/data/backend in a fast-moving environment.
Our job descriptions often reflect our ideal candidate. If you have a strong foundation of relevant skills and a passion for this field, we encourage you to apply, even if you don't check every box.
What We Offer
- High Impact: ship real features in weeks, not quarters
- Cutting-Edge Tech: work to solve problems no one has cracked before.
- Remote & Flexible: work from anywhere with a culture built on trust, autonomy, and balance.
- Growth & Ownership: own features end-to-end, learn rapidly, and grow with the company as we scale.
- Top-Tier Compensation: competitive salary, performance bonuses, equity upside, and strong benefits.
- Team & Culture: small, senior team that values collaboration, creativity, and building something meaningful together.
- Medical, Dental, Vision Insurance.
- 401k with Employer Contributions.
- Paid Time Off & Birthday Leave.
- Health Savings Account (HSA) Employer Contributions with High-Deductible Health Plan.
- Employer-paid Short-Term/Long-Term Disability Insurance.
- And more!
Compensation Range
- $240,000 - $350,000 base salary
- $50,000 - $80,000 variable compensation
Actual compensation offer to candidate may vary from posted hiring range based upon geographic location, work experience, education, and/or skill level. The pay ratio between base pay and target incentive (if applicable) will be finalized at the offer stage.
At IR, we celebrate, support, and thrive on difference for the benefit of our employees, our products, and our community. We are proud to be an Equal Employment Opportunity employer and encourage applications from all suitable candidates; we never discriminate based on race, religion, national origin, gender identity or expression, sexual orientation, age, or marital, veteran, or disability status.
Apply
Remember Job
Sign In
Email Address Password Log in
Share
Link
Skills
- Go
- Rust
- Python
- Kubernetes
- EKS
- AWS
- VPC
- IAM
- EC2
- S3
- RDS
- MLOps
- GPU
- LLVM
- clang
- CI/CD
Similar roles
SOFTWARE ENGINEERING MANAGER
NMDC Career site · Abu Dhabi, United Arab Emirates · United Arab Emirates · today
Forward Deployed Engineer, Hebrew speaker
Cloudflare · Hybrid · today
Software Engineer
Cloudflare · Location not stated · today
Senior Software Engineer
Appian · Chennai, India · today
Senior Software Engineer -Backend/Full Stack
Crunchyroll · Hyderabad, Telangana, India · today
Software Engineer - Egress (Go/Rust)
Cloudflare · Location not stated · EUR 54,000–75,000/yr · today