Apply on Employer Site

Stuut · 23 hours ago

Lead Site Reliability Engineer

United States

Full-time

Remote

Senior Level, Lead/Staff

7+ years exp

Stuut-ai is transforming accounts receivable for B2B companies, making collections smarter and faster. The Lead Site Reliability Engineer will drive the strategy, architecture, and execution of reliability and operational excellence across the platform, ensuring systems are highly available and resilient as the company grows.

Artificial Intelligence (AI)Financial ServicesFinTechSoftware

Responsibilities

Set the Reliability Strategy: define the long-term vision for site reliability, including SLOs/SLIs, error budgets, availability targets, and operational standards

Build & Scale Reliable Infrastructure: architect and maintain resilient, scalable cloud infrastructure across AWS and Kubernetes, ensuring systems are secure, fault-tolerant, and cost-effective

Own Observability & Monitoring: design and evolve monitoring, alerting, and logging systems that provide clear, actionable signals across services and environments

Lead Incident Response & Postmortems: own incident management practices, lead major incident response, and drive blameless postmortems that result in meaningful system improvements

Improve System Resilience: identify reliability risks and lead efforts around redundancy, failover, capacity planning, and graceful degradation

Optimize CI/CD & Deployment Reliability: partner with engineering teams to ensure deployments are safe, observable, and reversible; improve rollout strategies and reduce operational risk

Partner with Product & Engineering Teams: collaborate early in the development lifecycle to influence system design, scalability, and reliability tradeoffs

Reduce Toil & Improve Developer Experience: automate operational tasks, improve runbooks, and build tooling that reduces manual work and accelerates safe execution

Drive Root Cause Resolution: guide teams through deep debugging of reliability issues, ensuring fixes address underlying causes rather than symptoms

Influence Reliability Culture: promote reliability-first thinking, strong operational hygiene, and shared ownership of production systems across engineering

Mentor & Level Up the Team: coach engineers on reliability principles, incident handling, infrastructure design, and operational best practices

Qualification

Site Reliability EngineeringAWSKubernetesObservabilityPythonTypeScriptDockerCI/CDInfrastructure as CodeLeadershipCollaborationMentoring

Required

Have 7+ years of experience in site reliability engineering, infrastructure engineering, or backend software engineering

Have designed and operated highly available, production-grade systems supporting rapid product iteration

Are fluent in Python and/or TypeScript, and comfortable building automation and tooling to support reliability goals

Have a deep experience with AWS, Kubernetes (EKS), Docker, and cloud-native architectures

Have implemented and evolved observability stacks (metrics, logs, traces) and know how to create high-signal alerting

Understand how to design, measure, and enforce SLOs, SLIs, and error budgets

Have supported systems built with modern stacks such as FastAPI, Vue.js, PostgreSQL (RDS), and event-driven architectures

Have improved reliability and operational maturity in environments using CI/CD pipelines, infrastructure as code, and modern deployment workflows

Can balance reliability, velocity, and cost — making pragmatic tradeoffs that serve customers and the business

Enjoy collaborating across Product, Backend, Frontend, and Infrastructure teams to improve system health

Thrive in a role that blends deep technical execution, system design, and leadership influence in a fast-moving environment

Benefits

Medical, dental & vision insurance coverage for you

401(k) & Match

Equity

Flexible PTO

Parental Leave

Company

Stuut

Stuut provides an AI platform that automates accounts receivable work for companies.

Founded in 2024

New York, New York, USA

11-50 employees

https://www.stuut.ai

Funding

Current Stage

Early Stage

Total Funding

$35.5M

Key Investors

Andreessen HorowitzValley VenturesActivant Capital

2025-11-20Series A· $29.5M

2025-04-24Seed

2024-11-01Seed· $6M

Leadership Team

Tarek Alaruri

CEO & Co-founder

Miraj Mohsin

Co-Founder | Chief Design Officer

Recent News

Business Insider

Silicon Valley has a new cool kid agency and it's behind TBPN's branding

2025-12-07

FinTech Futures

AR start-up Stuut raises a16z-led $29.5m Series A

2025-11-25

alleywatch.com

The AlleyWatch Startup Daily Funding Report: 11/24/2025

2025-11-25

Company data provided by crunchbase