Job Description
Fingerprint empowers enterprises to detect and stop online fraud with the world’s most accurate device intelligence. We lead our industry with bleeding-edge identification capabilities and work on turning new ideas and discoveries in the fraud detection space into reality.
Fingerprint is a globally dispersed, 100% remote company. We were named on the 2026 Forbes Best Startup Employers list and ranked #803 on the 2026 Inc. 5000 list of America’s fastest-growing private companies. We have raised $77M and are backed by Craft Ventures ( Tesla, Facebook, Airbnb ), Nexus Venture Partners ( Postman , Apollo.io, MinIO , Druva) and Uncorrelated Ventures ( Redis, Rollbar, Gradle ).
What you'll do
- Make reliability measurable: define SLIs and SLOs for Fingerprint's critical request paths (identification, events, server APIs, client agents) with the teams that own them; make them visible, reviewed, and tied to decisions. Introduce error budgets as the mechanism for balancing reliability investment against feature work, and coach EMs and Staff engineers on using them. Own the reliability metrics that leadership uses to judge progress; be the No Nonsense voice on whether we are actually getting better.
- Raise the operational bar: Strengthen the incident lifecycle end to end: detection, response, communication, postmortem quality, and follow-up completion. Make the postmortem the most useful document a team writes. Close the "customers find out before we do" gap: drive alert quality, correctness anomaly detection, and escalation design across teams, working with Cloud Platform on shared tooling. Lead reliability reviews for high-risk changes and new services (production readiness, capacity, failure modes, rollback), teaching through review rather than gatekeeping. Introduce deliberate failure testing (game days, chaos exercises) in staging first, then production, to discover gaps and safe limits before customers do.
- Build the SRE mindset in teams: Embed with teams for time-boxed engagements; pair on their hardest reliability problems, leave behind better practices and a stronger owner, then move on. Develop Staff and Lead engineers as reliability leaders in their own groups — the goal is that every team has someone who thinks like an SRE. Codify practices that stick: production readiness checklists, on-call standards, runbook quality, change safety norms. Make them lightweight enough that teams choose to use them. Partner with the Architect and tech leads so reliability is designed in, not retrofitted. Stay hands-on: Dig into production during incidents and investigations. Write tooling, dashboards, and reference implementations. Be credible with the engineers you are asking to change how they work. Lead AI adoption in reliability practice: set norms for AI-assisted incident investigation, postmortem analysis, runbook authoring, and observability tooling across teams, and shape our runbooks, alerts, and operational data so AI agents can safely help diagnose and operate our systems alongside engineers.
- What you’ll do (continues): Lead AI adoption in reliability practice: set norms for AI-assisted incident investigation, postmortem analysis, runbook authoring, and observability tooling across teams, and shape our runbooks, alerts, and operational data so AI agents can safely help diagnose and operate our systems alongside engineers.
Who you are
- 10+ years of engineering experience, with 3+ years as an SRE, production engineer, or reliability-focused Staff engineer operating across multiple teams — you have owned reliability for a platform, not just for a service you built.
- Deep experience with SLI/SLO design and error budgets in practice, including the hard part: getting product teams to adopt and act on them.
- Strong incident leadership: you have run incident response and postmortems for high-severity, customer-facing incidents and materially improved how an organization learns from them.
- Hands-on depth in distributed systems failure modes — cache/database saturation and cascading failure, retry storms, capacity limits, degradation and load shedding — in a high-throughput, low-latency environment.
- Fluent in Kubernetes, AWS, and modern observability tooling (Datadog or equivalent).
- Comfortable reading and writing production code (Go, TypeScript, or similar) and infrastructure as code. You can ship a fix, not just recommend one.
- Track record of leading through influence: you have changed how teams you did not manage operate, and can explain how adoption actually happened.
- Teacher's instinct: you have coached engineers into owning reliability and can point to practices that persisted after you stepped back.
- Exceptional written communication: you make incidents, risks, and trade-offs legible to engineers and executives alike, and you default to async, documented decision-making.
- AI-native by default: you use AI tools as a normal part of how you investigate incidents, analyze telemetry, write runbooks and postmortems, and build tooling — and you have opinions, from experience, about where they accelerate reliability work and where they don't yet.
- Operate for an AI-assisted org: you think about how runbooks, alerts, dashboards, and operational data should be structured so that both humans and AI agents can diagnose and act on them safely — legible signals, clear ownership, strong guardrails on automated change.
- Pragmatism over purity: you know that reliability competes with delivery, and you can make the case for the right investment at the right time — and say when a risk is acceptable.
Nice to have
- Experience in fraud detection, identity, payments, or other adversarial, real-time domains.
- Multi-region, cell-based, or failure-isolation architecture experience.
- Experience with Elasticsearch, Redis, DynamoDB, or Kafka at scale, including their failure modes.
- Familiarity with FinOps and the reliability/cost trade-off in cloud infrastructure.
Compensation
For US-based employees, the cash compensation range for this role is $177,000 – $240,000. We share salary ranges on all job postings, specific to the hiring location. Offers vary based on experience, education, certifications/licenses, skills, training, and market conditions.
Additional notes
- Remote, global team with diverse locations; visa sponsorship not provided.
- We encourage applications from underrepresented groups in tech.
- Important notices: California CCPA, EU GDPR notices as applicable.



