RxSense · 1 week ago
Principal AI Ops Engineer
RxSense is a healthcare technology company that provides platforms and solutions to improve the management and access of cost-effective pharmacy benefits. The Principal AIOps Engineer will build the platform that makes AI cheap, fast, safe, and observable, owning the infrastructure that every AI-powered product at RxSense depends on.
Cloud ComputingHealthcareInformation TechnologySoftwareCloud ManagementHealth CareInformation Services
Responsibilities
Build and maintain end-to-end deployment pipelines for AI-powered applications, including artifact builds, environment promotion, rollback, and observability hooks. Drive new greenfield deployment platforms from initial build to the default that AI teams ship on
Stand up and operate the runtime and lifecycle infrastructure for production agents, including deployment, versioning, monitoring, rate-limiting, and retirement. Define the deployment contract (config, secrets, tools, memory, evals) and the operational SLOs
Own how the organization provisions, rotates, scopes, and meters access to model provider APIs (Anthropic, OpenAI, and others). Build a key management layer that enforces per-team and per-app quotas, prevents leakage, and gives finance and engineering a clear view of spend
Build evals into the CI/CD pipeline so no agent or LLM-powered service ships without passing a defined eval bar. Design the framework so product teams can author their own evals against a shared harness, and so eval results gate promotion across environments
Stand up self-hosted inference for workloads where managed APIs aren't the right fit, including latency-sensitive paths, regulated data, cost optimization, and vendor redundancy. Own the serving stack, the autoscaling and GPU economics behind it, and the playbook for when a workload belongs to a managed provider versus internal infrastructure
Design and build the shared developer harness that every AI-powered service uses: prompt management, model routing, retries, tracing, eval hooks, and policy enforcement. Set the abstractions that determine how fast every other AI engineer can ship for the next three years
Partner with finance on cost visibility, including token accounting, per-feature cost attribution, and real-time spend observability
Write documentation, runbooks, and clear interfaces so the platform is adoptable by other engineering teams without hand-holding
Participate in code review and promote collaboration and best practices including simplicity, automation, sound design patterns, test coverage, and reusability
Qualification
Required
BS (or higher, e.g., MS or Ph.D.) in Computer Science or related technical field involving coding, or equivalent technical experience
6+ years of platform, infrastructure, or DevOps engineering, with at least 2 years building production infrastructure for AI/ML or LLM-powered systems. We care more about depth and drive than years on a resume
Deep hands-on experience designing and operating CI/CD pipelines for high-velocity engineering organizations, including artifact management, environment promotion, and progressive rollout
Strong AWS background, comfortable down to the IAM, networking, and container orchestration layers
Proven track record building developer platforms or internal tools that other engineering teams adopted by choice, not by mandate
Production experience with LLM-powered applications, including prompt management, model routing, retries, tracing, and the operational realities of running agents or chains in production
Hands-on coding fluency in Python or TypeScript, ideally both. This is a keyboard role, not an architecture-only role
Comfortable operating in a polyglot environment. The RxSense AI engineering stack spans Python, .NET, and TypeScript, and you will deploy and support services across all three
Comfortable owning the cost and reliability conversation with both engineering leadership and finance partners
Strong written communication and a bias toward documentation, runbooks, and clear interfaces
Proven analytical thinking and problem-solving skills
Excellent communication skills, both verbal and written
Preferred
Direct experience integrating with Anthropic, OpenAI, or other frontier model provider APIs at scale, including key management, quota enforcement, and capacity planning
Hands-on experience self-hosting models with vLLM, TGI, SGLang, or similar inference servers, including GPU autoscaling and cost optimization
Built or contributed to an eval framework that gated production deployments
Familiarity with agent runtimes and frameworks such as the Claude Agent SDK, LangGraph, or in-house equivalents
Working familiarity with .NET, enough to read code, debug a deploy, and pair with service owners
Background in healthcare, PBM, pharmacy, or another regulated data environment
FinOps experience, particularly attributing AI spend to features or business units
Kubernetes operator experience or comfort with custom controllers
Experience with Agile development methodologies, preferably both Scrum and Kanban
Company
RxSense
RxSense is a health technology company provides IT solutions for flexible, efficient, and transparent pharmacy benefits administration.
H1B Sponsorship
RxSense has a track record of offering H1B sponsorships. Please note that this does not
guarantee sponsorship for this specific role. Below presents additional info for your
reference. (Data Powered by US Department of Labor)
Distribution of Different Job Fields Receiving Sponsorship
69%
Represents job field similar to this job
Engineering and Development
Marketing
Management and Executive
Product Management
Trends of Total Sponsorships
*2026 (1)
2025 (2)
2024 (3)
2023 (4)
2022 (1)
2021 (5)
2020 (10)
Funding
Current Stage
Growth StageTotal Funding
unknownKey Investors
Parthenon Capital Partners
2020-05-12Private Equity
Leadership Team
Recent News
2026-05-20
2026-05-12
2025-12-20
Company data provided by crunchbase