Fingerprint logo
Fingerprint

Senior Site Reliability Engineer

$115,000 – $195,500estimatedRemoteFull-timeSeniorWorldDevelopment
Share

Senior Site Reliability Engineer

Fingerprint empowers enterprises to detect and stop online fraud with the world’s most accurate device intelligence. We lead our industry with bleeding-edge identification capabilities and work on turning new ideas and discoveries in the fraud detection space into reality. Fingerprint is a globally dispersed, 100% remote company.

What you'll do

  • Own the reliability of core production systems end to end — instrument them, set targets for them, operate them, and be accountable for how they behave under real traffic.
  • Define and maintain SLIs and SLOs for the critical paths you own, wire them into dashboards and alerts, and use error budget burn as the evidence base for what gets fixed next.
  • Drive alert quality: raise signal, kill noise, and close the gap where customers notice a problem before our monitoring does. Anomaly and correctness detection matter as much as uptime.
  • Take a lead role in incident response — investigate systematically across service boundaries, restore service, and write postmortems that produce follow-ups people actually complete.
  • Build secure, resilient, and cost-efficient infrastructure, with explicit attention to failure modes: timeouts and retries, backpressure and load shedding, graceful degradation, and blast radius containment.
  • Do capacity and performance work with real data — load testing, profiling, saturation analysis, and headroom planning ahead of growth rather than after an incident.
  • Improve change safety: progressive delivery, automated rollback, meaningful pre-production signal, and deployment practices that make shipping boring.
  • Manage infrastructure through code and configuration (we primarily use Terraform), consistently applying patterns that align with our overall service architecture.
  • Design, write, and ship software and developer-facing tooling that reduces toil and makes operating services straightforward for the engineers who own them.
  • Run deliberate failure testing — game days and chaos exercises, staging first — to find the gaps and safe limits before customers do.
  • Partner with product engineering teams on production readiness for new and high-risk services: capacity, failure modes, rollback plans, runbooks, and on-call handoff.
  • Teach through review rather than gatekeeping. Participate in the on-call rotation, and improve it: better runbooks, clearer escalation, less pager fatigue for everyone in it.
  • Approach all engineering work with a security lens — actively looking for vulnerabilities in your own work and in peer reviews.
  • Act as the go-to person for hard production problems in your area, and mentor engineers through code review, pairing, and design feedback so operational knowledge doesn't silo.

Qualifications

  • 6–10 years of experience in SRE, production engineering, infrastructure, or backend engineering within primarily cloud-based environments (AWS preferred), with meaningful time spent responsible for systems in production.
  • A track record of owning a system end to end — you've designed something significant, shipped it, operated it, and lived with the consequences when it misbehaved.
  • Hands-on experience defining and operating against SLIs, SLOs, and error budgets — targets adopted and acted on.
  • Strong incident skills: led or primary responder on high-severity, customer-facing incidents, and improved learning from them.
  • Depth in distributed systems failure modes in high-throughput, low-latency environments — cache/database saturation, cascading failure, retry storms, capacity limits, degradation and load shedding.
  • Depth in cloud infrastructure fundamentals: networking, load balancing, containerization (EKS/Kubernetes), and distributed systems.
  • Strong hands-on experience managing infrastructure through code and configuration (Terraform or equivalent).
  • Solid programming skills in Go, Python, or a comparable language — production-ready software that ships fixes.
  • Fluency with observability tooling (Datadog, Prometheus, Grafana, OpenTelemetry, or similar), including instrumenting systems yourself.
  • Hands-on experience operating Redis/ElastiCache in production — cluster/shard management, failover behavior, memory eviction policies, and scaling strategies.
  • Fluency with software engineering best practices: source control, code review, comprehensive test coverage, and safe deployment.
  • A high level of personal ownership and autonomy, with real experience working without clearly defined requirements.
  • Pragmatism over purity — ability to balance reliability with delivery and assess risk appropriately.
  • Strong written and verbal communication in English — design docs, PR reviews, incident updates, and postmortems that include engineers outside your team.
  • AI-native by default — comfortable using AI tools in investigations, telemetry analysis, runbooks, and tooling, with informed opinions on their use.

Compensation & Benefits

For US-based employees, the cash compensation range for this role is $152,000 – $205,000. We share salary ranges on all job postings, with location-specific adjustments. Offers vary based on experience, education, skills, and market conditions. Fingerprint is an all-remote company and does not sponsor visas. We are committed to an inclusive workplace and encourage applications from underrepresented groups in tech.

Ready to apply for this role?

Apply Now →

Please mention “I found this job at Remocate!”

Get fresh remote & relocation jobs in your inbox

A short digest of new openings. No spam, unsubscribe anytime.

Related jobs

Apply Now →