Mirantis · 2 weeks ago
Senior AI Infrastructure & Platform Operations Engineer (remote in the US)
Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. The Senior AI Infrastructure & Platform Operations Engineer will lead the management of expansive AI ecosystems, ensuring the reliability and efficiency of AI service platforms while driving the development of automated operational capabilities.
Artificial Intelligence (AI)Cloud ComputingEnterprise SoftwareDevOpsIaaSIT InfrastructureOpen Source
Responsibilities
Lead the investigation and resolution of complex infrastructure, networking, and platform-related incidents
Act as a senior escalation point for operational teams during critical service-impacting events
Support large-scale NVIDIA GPU infrastructure and high-performance networking environments
Troubleshoot complex Linux, Kubernetes, networking, storage, and hardware-related issues
Analyze platform performance, capacity, stability, and reliability trends to proactively identify risks
Lead root cause analysis activities and drive long-term corrective actions
Collaborate with engineering teams, hardware vendors, and datacenter personnel to resolve complex technical challenges
Participate in major incident management and service restoration activities
Provide technical leadership for Kubernetes platform operations and supporting infrastructure services
Drive improvements in platform reliability, observability, monitoring, and operational processes
Identify opportunities to automate repetitive operational activities and improve operational efficiency
Contribute to operational readiness reviews, infrastructure changes, upgrades, and service introductions
Support the adoption and operation of AI-powered infrastructure services and operational capabilities through k0rdent AI
Evaluate emerging technologies and operational practices to improve service delivery and platform resilience
Mentor and support AI Infrastructure & Platform Operations Engineers
Share technical knowledge through documentation, training sessions, and operational reviews
Develop and maintain operational standards, runbooks, troubleshooting guides, and best practices
Help define operational processes, escalation paths, and service reliability standards
Act as a trusted technical advisor during operational planning and service improvement initiatives
Qualification
Required
7+ years of experience in infrastructure operations, platform operations, site reliability engineering, network operations, cloud operations, datacenter operations, or related technical roles
Expert-level Linux administration and troubleshooting skills
Strong networking expertise, including experience diagnosing complex performance, connectivity, and reliability issues
Strong experience operating Kubernetes in production environments
Experience supporting large-scale production infrastructure and distributed systems
Proven experience leading technical investigations and managing complex incidents
Experience performing root cause analysis and driving long-term operational improvements
Strong understanding of observability, monitoring, and service reliability practices
Excellent troubleshooting and analytical skills across multiple infrastructure domains
Strong communication, collaboration, and stakeholder management skills
Preferred
NVIDIA GPU infrastructure and accelerated computing platforms
InfiniBand networking and NVIDIA UFM
AI infrastructure environments
HPC environments
Platform Engineering or Site Reliability Engineering (SRE)
Large-scale Kubernetes operations
Infrastructure automation technologies and Infrastructure-as-Code practices
Observability platforms such as Grafana, Prometheus, ELK, or OpenTelemetry
Performance analysis and optimisation of distributed infrastructure platforms
Technical leadership, mentoring, or team lead responsibilities
Benefits
Professional development and training
Attend conferences and working groups
Company outings, happy hours, hackathons, and tech talks
Company
Mirantis
Mirantis develops cloud infrastructure and container management software for organizations to build, operate, and scale applications. It is a sub-organization of IREN.
H1B Sponsorship
Mirantis has a track record of offering H1B sponsorships. Please note that this does not
guarantee sponsorship for this specific role. Below presents additional info for your
reference. (Data Powered by US Department of Labor)
Distribution of Different Job Fields Receiving Sponsorship
62%
Represents job field similar to this job
Engineering and Development
Customer Service and Support
Creatives and Design
Product Management
Trends of Total Sponsorships
*2026 (1)
2025 (3)
2024 (4)
2023 (8)
2022 (6)
2021 (7)
2020 (8)
Funding
Current Stage
Late StageTotal Funding
$240MKey Investors
Horizon Technology FinanceIntel CapitalInsight Partners
2026-05-05Acquired
2023-10-10Debt Financing· $20M
2017-11-15Secondary Market
Recent News
The Motley Fool
2026-08-05
GlobeNewswire
2026-08-04
2026-07-22
Company data provided by crunchbase