Site Reliability Engineer

Woonsocket, RI
Date Posted:29-Sep-2026
Work Type:Hybrid
Job Number:502122

Job Description

Job title: Site Reliability Engineer
Location: hybrid based out of Woonsocket, RI
Duration: 03 Months
 
Years of experience required: 8+
 
Required Qualifications
  • 8+ years of Senior Software engineering experience in SRE, DevOps, platform engineering, or related production-systems roles in distributed systems at production scale with active on-call responsibility.
  • Demonstrated experience as an on-call Incident Commander (IC) for P1 or P2 incidents — structured leadership updates, not just participant involvement.
  • Experience tuning and validating time-series anomaly detection models in a production observability context anomaly-based detection is a core function of this role.
  • Strong programming proficiency in Python, React, and Java at production quality — capable of writing operational tooling that other engineers will rely on.
  • Hands-on experience designing SLIs, SLOs, and managing error budgets for customer-facing or business-critical services.
  • Deep observability platform experience: Prometheus, Grafana, OpenTelemetry, and at least two of the log aggregation solutions (Loki, Splunk, Elasticsearch).
  • Fleet-scale deployment awareness: familiarity with progressive rollout strategies, blast radius management, and configuration drift as a reliability risk in large unattended node deployments.
  • Strong cloud platform expertise in Google Cloud Platform (GCP) and Rancher K3s.
  • Advanced Kubernetes operational experience: debugging, resource management, networking policies, and workload failure modes.
  • Experience with AI-assisted tooling and development.
  • Experience diagnosing and resolving workflow orchestration issues, batch processing failures, scheduler performance problems, and building observability on data pipelines: Apache Airflow and Tidal.
 
Preferred Qualifications
  • Experience owning Production Readiness Reviews or service launch gates.
  • Strong proficiency in transforming large-scale operational and telemetry data into actionable business insights using SQL-based analytics and reporting frameworks: Google BigQuery, PostgreSQL.
  • Hands-on chaos or fault injection experience.
  • TIC (Technical Incident Commander) certification or equivalent structured incident command training.
  • Experience operating distributed systems in retail, pharmacy, healthcare, or other operationally sensitive environments where failures have direct patient or customer impact.
  • LLM integration for operational use cases (alert summarization, runbook suggestion, incident triage assistance) — design or implementation experience.
  • Experience with streaming data platforms: Kafka.
  • Experience with service mesh and traffic management: Istio, Envoy.
  • Infrastructure-as-code proficiency at production scale: Terraform or Ansible. Position Summary
  • Own and drive the end-to-end reliability, availability, and performance of critical retail and pharmacy technology platforms across hybrid cloud and on-premises environments.
  • Establish and maintain SLI/SLO health, alerting strategies, observability standards, and business-aligned monitoring for the assigned application domain.
  • Lead production incident response as Incident Commander, drive root cause analysis, postmortems, and continuous reliability improvements.
  • Partner with engineering, product, and operations teams to embed reliability, resiliency, scalability, and operational readiness into system design and delivery.
  • Build and optimize automation, self-service capabilities, and operational tooling to eliminate toil, improve efficiency, and reduce manual intervention.
  • Design and execute proactive reliability initiatives, including production readiness reviews, dependency risk assessments, fault injection, and chaos engineering exercises.
  • Mentor engineers, champion SRE best practices, and enable teams to independently detect, respond to, and learn from production issues with minimal SRE involvement.
  • Influence organizational adoption of SLO-driven engineering, observability, incident management, and reliability practices through collaboration, credibility, and measurable outcomes.
 

Applicant Notices & Disclaimers
  • For information on benefits, equal opportunity employment, and location-specific applicant notices, click here


At SPECTRAFORCE, we are committed to maintaining a workplace that ensures fair compensation and wage transparency in adherence with all applicable state and local laws. This position's pay range is $50.00/hr – $50.85/hr.