Staff Site Reliability Engineer
Los Angeles, United States · Hybrid · Full-time
- Posted 1mo ago
- From Crunchyroll’s careers page
- Location
- Los Angeles, United States
- Work mode
- Hybrid
- Type
- Full-time
- Level
- Senior
- Experience
- 12+ years
- Department
- Information Technology
Opens the listing on boards.greenhouse.io
Let the right jobs find you
In your inbox every Wednesday and SaturdayPersonalised suggestions from verified career pages, matched to your role, location, level and skills.
About the role
About the role
We are hiring a Staff Site Reliability Engineer (SRE) to join the Center for Data & Insights (CDI) in the US and play a critical role in advancing the reliability, scalability, performance, and security of Crunchyroll's consumer-facing data platforms. As a senior technical leader, you will partner closely with Engineering, Data, Infrastructure, Product, and Security teams to design and operate resilient cloud-native systems that power critical business and customer experiences. You will drive initiatives across observability, incident management, automation, capacity planning, disaster recovery, and operational excellence while helping teams adopt modern SRE practices such as SLIs, SLOs, and error budgets.
Core Areas of Responsibility
- Reliability Engineering: Define, measure, and continuously improve the reliability, availability, and performance of CDI platforms through SLIs, SLOs, and error budgets.
- Operational Excellence: Establish and drive best practices for incident management, root cause analysis, postmortems, and service ownership across engineering teams.
- Observability & Monitoring: Build and evolve comprehensive monitoring, logging, tracing, and alerting capabilities to enable proactive issue detection and rapid resolution.
- Automation: Identify operational inefficiencies and develop automation, self-service capabilities, and self-healing mechanisms to improve engineering productivity.
- Platform Scalability: Design and optimize cloud-native infrastructure and services to support growing business demands while maintaining performance and cost efficiency.
- Infrastructure Engineering: Drive Infrastructure as Code (IaC), platform standardization, and deployment automation to improve consistency, reliability, and operational agility.
- Capacity Planning & Performance: Lead capacity planning and performance optimization initiatives to ensure platforms can scale predictably and efficiently.
- Disaster Recovery & Resilience: Develop and regularly validate disaster recovery, backup, and business continuity strategies to ensure platform resiliency.
- Security Operations (SecOps): Partner with Crunchyroll's security team to integrate security controls, operational risk management, and security best practices into platform operations and engineering workflows.
- Vulnerability Management: Own the triage and remediation of identified vulnerabilities across infrastructure, platform, container, and application security vulnerabilities through established Crunchyroll vulnerability management processes.
- Penetration Testing & Security Remediation: Support penetration test scoping activities by providing technical context on CDI platforms. Own the triage, prioritization, and remediation of resulting findings to drive timely resolution and strengthen platform security posture.
- Cloud & Kubernetes Security: Implement and maintain secure cloud, container, and Kubernetes environments following least-privilege, defense-in-depth, and Zero Trust principles.
- Cross-Functional Leadership: Collaborate with Engineering, Data, Product, Infrastructure, and Security teams to drive reliability, scalability, and security initiatives across CDI.
- Mentorship & Engineering Excellence: Mentor engineers and champion a culture of operational excellence, reliability, ownership, continuous improvement, and security awareness.
About You
We get excited about candidates like you, because…
- 12+ years of experience in Site Reliability Engineering (SRE), Platform Engineering, Infrastructure Engineering, or related disciplines, with a proven track record of operating and scaling production-critical systems.
- Deep expertise in Kubernetes and GCP, including the design, deployment, and operation of highly available, cloud-native platforms at scale.
- Strong Infrastructure as Code (IaC) experience, preferably with Terraform, and a commitment to automation, standardization, and operational efficiency.
- Solid foundation in Linux systems administration, networking, and distributed systems, with the ability to troubleshoot complex production issues across multiple layers of the technology stack.
- Proficiency in one or more programming and scripting languages, such as Go, Python, Java, or Shell, with a focus on automation and platform engineering.
- Hands-on experience with modern observability platforms and practices, including Prometheus, Grafana, OpenTelemetry, Datadog, or equivalent monitoring and telemetry solutions.
- Demonstrated expertise in incident management, service reliability, capacity planning, performance optimization, and operational excellence, including the implementation of SLIs, SLOs, and error budgets.
- Strong understanding of cloud and platform security, including container security, Kubernetes security, CI/CD security, vulnerability management, and secure infrastructure operations.
- Good to have knowledge of security frameworks and best practices, including OWASP Top 10, Identity and Access Management (IAM), secrets management, Secure Software Development Lifecycle (SSDLC), and security-by-design principles.
- Excellent collaboration, communication, and technical leadership skills, with experience influencing architectural decisions, driving cross-functional initiatives, and mentoring engineers in reliability and operational best practices.
Skills they ask for
Pick one to see other roles that ask for it.
About Crunchyroll
Anime streaming and fan experiencesCrunchyroll is an anime entertainment brand offering streaming content and experiences for anime fans worldwide.
See all 48 roles at CrunchyrollMore roles at Crunchyroll
See all 48- Senior Manager, Rights ManagementHyderabad · Senior · HybridBusiness Operations · Senior · HybridHyderabad, India15h
- Manager, Rights ManagementHyderabad · HybridBusiness Operations · HybridHyderabad, India15h
- Senior Manager, User Acquisition – Paid SocialLos Angeles · Senior · HybridMarketing · Senior · HybridLos Angeles, United States1d
- Senior Manager, Internal Communications, Programs & EventsLos Angeles · Senior · HybridCommunications and Public Affairs · Senior · HybridLos Angeles, United States1d
Let the right jobs find you
In your inbox every Wednesday and SaturdayPersonalised suggestions from verified career pages, matched to your role, location, level and skills.