SIGN IN
Senior AI Infrastructure & Platform Operations Engineer (remote in the US) jobs in United States
cer-icon
Apply on Employer Site
company-logo

Mirantis · 2 weeks ago

Senior AI Infrastructure & Platform Operations Engineer (remote in the US)

Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. The Senior AI Infrastructure & Platform Operations Engineer will lead the management of expansive AI ecosystems, ensuring the reliability and efficiency of AI service platforms while driving the development of automated operational capabilities.
Artificial Intelligence (AI)Cloud ComputingEnterprise SoftwareDevOpsIaaSIT InfrastructureOpen Source
check
H1B Sponsor Likelynote

Responsibilities

Lead the investigation and resolution of complex infrastructure, networking, and platform-related incidents
Act as a senior escalation point for operational teams during critical service-impacting events
Support large-scale NVIDIA GPU infrastructure and high-performance networking environments
Troubleshoot complex Linux, Kubernetes, networking, storage, and hardware-related issues
Analyze platform performance, capacity, stability, and reliability trends to proactively identify risks
Lead root cause analysis activities and drive long-term corrective actions
Collaborate with engineering teams, hardware vendors, and datacenter personnel to resolve complex technical challenges
Participate in major incident management and service restoration activities
Provide technical leadership for Kubernetes platform operations and supporting infrastructure services
Drive improvements in platform reliability, observability, monitoring, and operational processes
Identify opportunities to automate repetitive operational activities and improve operational efficiency
Contribute to operational readiness reviews, infrastructure changes, upgrades, and service introductions
Support the adoption and operation of AI-powered infrastructure services and operational capabilities through k0rdent AI
Evaluate emerging technologies and operational practices to improve service delivery and platform resilience
Mentor and support AI Infrastructure & Platform Operations Engineers
Share technical knowledge through documentation, training sessions, and operational reviews
Develop and maintain operational standards, runbooks, troubleshooting guides, and best practices
Help define operational processes, escalation paths, and service reliability standards
Act as a trusted technical advisor during operational planning and service improvement initiatives

Qualification

Linux administrationKubernetesNVIDIA GPU infrastructureInfiniBand networkingNVIDIA UFMAI infrastructureHPC environmentsInfrastructure automationInfrastructure-as-CodeObservability platformsGrafanaPrometheusELKOpenTelemetryRoot cause analysisDistributed systemsSite Reliability Engineering

Required

7+ years of experience in infrastructure operations, platform operations, site reliability engineering, network operations, cloud operations, datacenter operations, or related technical roles
Expert-level Linux administration and troubleshooting skills
Strong networking expertise, including experience diagnosing complex performance, connectivity, and reliability issues
Strong experience operating Kubernetes in production environments
Experience supporting large-scale production infrastructure and distributed systems
Proven experience leading technical investigations and managing complex incidents
Experience performing root cause analysis and driving long-term operational improvements
Strong understanding of observability, monitoring, and service reliability practices
Excellent troubleshooting and analytical skills across multiple infrastructure domains
Strong communication, collaboration, and stakeholder management skills

Preferred

NVIDIA GPU infrastructure and accelerated computing platforms
InfiniBand networking and NVIDIA UFM
AI infrastructure environments
HPC environments
Platform Engineering or Site Reliability Engineering (SRE)
Large-scale Kubernetes operations
Infrastructure automation technologies and Infrastructure-as-Code practices
Observability platforms such as Grafana, Prometheus, ELK, or OpenTelemetry
Performance analysis and optimisation of distributed infrastructure platforms
Technical leadership, mentoring, or team lead responsibilities

Benefits

Professional development and training
Attend conferences and working groups
Company outings, happy hours, hackathons, and tech talks

Company

Mirantis

company-logo
Mirantis develops cloud infrastructure and container management software for organizations to build, operate, and scale applications. It is a sub-organization of IREN.

H1B Sponsorship

Mirantis has a track record of offering H1B sponsorships. Please note that this does not guarantee sponsorship for this specific role. Below presents additional info for your reference. (Data Powered by US Department of Labor)
Distribution of Different Job Fields Receiving Sponsorship
62%
Represents job field similar to this job
Engineering and Development
Customer Service and Support
Creatives and Design
Product Management
Trends of Total Sponsorships
*2026 (1)
2025 (3)
2024 (4)
2023 (8)
2022 (6)
2021 (7)
2020 (8)

Funding

Current Stage
Late Stage
Total Funding
$240M
Key Investors
Horizon Technology FinanceIntel CapitalInsight Partners
2026-05-05Acquired
2023-10-10Debt Financing· $20M
2017-11-15Secondary Market

Leadership Team

leader-logo
Shaun O'Meara
Chief Technology Officer
linkedin
leader-logo
Oleg Goldman
SVP of Operations
linkedin
Company data provided by crunchbase