SIGN IN
Staff AI Engineer | Agentic Systems jobs in United States
cer-icon
Apply on Employer Site
company-logo

Machinify · 1 week ago

Staff AI Engineer | Agentic Systems

Machinify is a leading healthcare intelligence company with expertise across the payment continuum, delivering unmatched value, transparency, and efficiency to health plan clients across the country. The role involves building production-grade agentic systems that audit medical claims end-to-end, requiring the ability to translate vague business problems into concrete technical solutions and set the technical direction for problem areas.
Big DataArtificial Intelligence (AI)HealthcareHealth InsuranceSaaSAnalyticsBusiness IntelligenceHealth CareMachine LearningPredictive Analytics
badNo H1Bnote

Responsibilities

Drive vague business problems to closure. Sit with clinical leads, product, and ops to understand what's actually broken, where the money is, and what "good" looks like. Translate that into a concrete technical problem statement with a measurable target — and push back when the framing is wrong
Define the metric before you build the system. Decide what you're optimizing (recall on overpayments? appeal-survival rate? cost per case? agreement with senior coders?), how it will be measured, what the baseline is, and what number constitutes shipping. Build the eval harness that produces it. No metric, no project
Scope and sequence the work. Break an ambiguous initiative into a phased plan with explicit decision points, kill criteria, and dependencies. Decide what's in scope, what's deferred, and what's not worth doing — and communicate that crisply to non-technical stakeholders
Set the technical direction for a problem area. Choose the agent topology, the context strategy, the model mix, the evaluation regime, the deterministic guardrails. Own the architectural call and the tradeoffs behind it. Other engineers — including senior ones — should be able to build against the foundation you set
Raise the bar on agent engineering. Lead by example on context engineering, structured outputs, citation grounding, eval discipline, and cost/latency control. Review designs and PRs from other engineers on the team and leave the codebase and the patterns sharper than you found them
Be the technical interface to the business. Present results to clinical, product, and executive stakeholders. Defend the methodology when findings are challenged. Know the domain well enough to argue with a senior coder about why a code is or isn't supported
Use AI tooling like a force multiplier. A meaningful fraction of your day will be spent driving Claude Code, Codex, and similar tools to plan, scaffold, refactor, debug, and evaluate. We expect you to be dramatically faster with these tools than most engineers are without them, and to teach the rest of the team to be the same

Qualification

Machine learningArtificial intelligenceLarge language models (LLM)Agent engineeringPython programmingAsync programmingType discipline in PythonTested code developmentOpenAI Agents SDKAnthropic SDKClaude-agent-sdkLangGraphClaude CodeCodexVS CodeGitMetric definition and ownershipEvaluation harness developmentLong-context citation-grounded systemsHealthcare domain knowledgeLegal domain knowledgeFinance domain knowledgeCaching for LLM workloadsObservability for LLM workloadsCost control for LLM workloadsOCRLayout-aware modelsTable extractionVision-language modelsMultimodal retrieval

Required

6+ years of applied ML / AI / software engineering experience with a Bachelor's in CS, Math, Engineering or equivalent — or 4+ years with a Master's / PhD in a similar program
At least two production systems you owned end-to-end from ambiguous problem statement through measured impact, ideally including at least one LLM- or agent-based system
A track record of driving vague problems to closure. You can point to initiatives where the brief was a paragraph, you scoped it, defined the metric, ran the work, and shipped a result that moved the business — not just a model or a PR
Strong stakeholder fluency. You can sit with non-technical domain experts (clinicians, coders, ops leads, product), extract what they actually mean, translate it into a technical problem, and translate technical tradeoffs back into terms they can decide on
Deep, hands-on agent engineering. You've designed agent loops from scratch, decided between single-agent and multi-agent topologies, engineered context (system prompts, tool surfaces, structured outputs, citation grounding), and debugged failure modes that other engineers couldn't
Eval-first instincts. You don't ship without an eval; you don't believe a number you can't reproduce; and you've built eval harnesses that other engineers on the team now depend on
Strong Python engineering. Clean abstractions, type discipline, async, tested code — at a level where junior and mid engineers learn from your PRs
Hands-on experience with at least one major agent SDK — OpenAI Agents SDK, Anthropic SDK / claude-agent-sdk, LangGraph, or equivalent — with strong opinions on the tradeoffs and the scars to back them up
Fluency with Claude Code / Codex as a power user — able to plan, execute, and debug non-trivial engineering tasks with these tools, including reading their source when needed
Solid command of VS Code and git — branches, rebases, worktrees, conflict resolution, PR workflows. Not optional

Preferred

Experience defining and owning a metric that the business actually trusts — precision/recall against expert ground truth, dollar-weighted impact, appeal-survival rate, or equivalent — including the data pipeline behind it
Prior work on long-context, citation-grounded systems where the model must point to evidence, not just answer
Healthcare, legal, finance, or any other domain where 'mostly right' is unacceptable and where findings get challenged by domain experts
Experience setting technical direction for a small group of engineers (formal or informal tech lead), including reviewing designs, mentoring on agent patterns, and being accountable for an area's quality
Familiarity with reasoning models (o-series, Claude extended thinking, Gemini thinking) and a sharp sense of when they earn their cost
Production experience with caching, observability, and cost control on LLM workloads at scale
Document understanding (OCR, layout-aware models, table extraction)
Vision-language models, multimodal retrieval
Experience presenting technical results to executive or external (customer / regulator) audiences

Benefits

Hybrid role — we have a strong preference for in-office collaboration, with flexibility for exceptional candidates.
Top Medical / Dental / Vision offerings.
FSA / HSA.
Tuition reimbursement.
Competitive salary, 401(k) with company match.
Unlimited PTO.
Meaningful equity.
A flexible, trusting environment where you'll be empowered to do your best work.

Company

Machinify

linkedincrunchbase
company-logo
Machinify is a SaaS platform that enables non-technical enterprises to build AI-powered products and processes.

Funding

Current Stage
Late Stage
Total Funding
$12.79M
Key Investors
Battery Ventures
2025-01-10Acquired
2025-01-01Series Unknown
2018-10-08Series A· $10M

Leadership Team

leader-logo
Prasanna Ganesan
CTO & Board Member
linkedin
Company data provided by crunchbase