SRE II
BrightMoney
Let the right jobs find you
Get personalised suggestions from verified company career pages, matched to your role, location, level, and skills.
Overview
Position Type
Full Time
Experience
Not specified
Job Description
What you will do?
-
Infrastructure Operations: Own and operate core production infrastructure on cloud platforms, ensuring high availability, observability, and scalability.
-
Incident Management: Lead incident response and author formal Root Cause Analysis (RCA) reports. Support the rollout of our new incident management platform and automated runbooks.
-
CI/CD & Automation: Design and optimize robust CI/CD pipelines. A major H2 goal is standardizing pipelines across all services following our Python upgrade and containerization tracks.
-
Security & Compliance: Implement infrastructure security compliance, including IAM roles, SCPs, and our upcoming Identity Platform (Teleport) rollout.
-
Observability: Maintain monitoring dashboards and alerting. Support the revamp of our VictoriaMetrics HA stack and ELK log optimization.
-
FinOps: Lead cloud cost analysis and resource tagging tracks to maintain efficient architecture.
-
Disaster Recovery: Lead the build-out of cross-region replicas and failover procedures to meet agreed RPO/RTO targets per service tier.