Sr Engineer, SRE TechOps CICD (Remote)
United States · Remote · Full-time
- Posted 1mo ago
- From CrowdStrike’s careers page
- Location
- United States
- Work mode
- Remote
- Type
- Full-time
- Level
- Senior
- Experience
- 10+ years
- Department
- Engineering
Opens the listing on crowdstrike.wd5.myworkdayjobs.com
Let the right jobs find you
In your inbox every Wednesday and SaturdayPersonalised suggestions from verified career pages, matched to your role, location, level and skills.
About the role
About CrowdStrike:
As a global leader in cybersecurity, CrowdStrike protects the people, processes and technologies that drive modern organizations. Since 2011, our mission hasn't changed. We're here to stop breaches, and we've redefined modern security with the world's most advanced AI-native platform. We work on large scale distributed systems, processing almost 3 trillion events per day and this traffic is growing daily. Ready to join a mission that matters? The future of cybersecurity starts with you.
About the Role:
CrowdStrike's internal SRE team owns the automation, reliability, and observability of the internal developer platform that thousands of CrowdStrike engineers rely on to build and deploy software rapidly, efficiently, and at scale. We provide the resilient infrastructure and operational rigor that let product, platform, and application teams ship with confidence.
We're looking for a Senior Site Reliability Engineer on the Technical Operations team to bring deep, hands-on expertise across load balancers, relational and non-relational databases, message queues (Kafka, Pulsar, RabbitMQ, RedPanda), and caching layers (Redis/Valkey, Varnish), paired with strong SLI/SLO instincts. This is a technical-anchor role: you're the person the team routes hard problems to, and you set the bar for operational excellence through the quality of your own work and the judgment you bring to design and incident reviews.
What You'll Do:
- Own the availability and health of key services within the CICD environment, maintaining a holistic view of system health across the platform.
- Build software and systems to manage platform infrastructure and applications, and drive automation for service deployment and operational workflows.
- Carry on-call responsibility for owned services; drive incident response and blameless postmortems to root cause.
- Gather and analyze metrics from operating systems and applications to support performance tuning and root cause analysis.
- Lead system design discussions, production readiness reviews, and capacity planning exercises.
- Evaluate and integrate agentic and AI-assisted workflows into existing team processes, and help teammates adopt them.
- Mentor mid-level and junior engineers through code review, design pairing, and incident retrospectives.
- Investigate and evaluate emerging technologies, and provide recommendations that support future roadmap goals.
- Build and maintain automated reporting on service health and compliance.
- Provide technical feedback and guidance on projects outside your core area of ownership, helping raise the bar across the broader engineering organization.
- Partner with peer senior engineers and engineering leaders to drive cross-team reliability improvements.
- Contribute to the Embedded SRE model, helping strengthen partnerships between SRE and the services teams.
What You'll Need:
- Must be eligible for CJIS clearance (requires U.S. citizenship or Green Card/permanent resident status).
- 10+ years of experience working in large-scale production SRE or infrastructure environments.
- 3+ years of experience leveraging and integrating AI-assisted workflows to increase engineering efficiency.
- Bachelor's degree in computer science or another highly technical, scientific discipline, or equivalent work experience.
- On-premise and cloud expertise deploying and operating CI/CD tools (Bazel, Jenkins, GitLab CI, GitHub Actions), IaC provisioning (Ansible, Chef, Puppet, Salt, Terraform), source code management (Bitbucket, GitHub, GitLab), and monitoring/observability platforms (Datadog, Grafana, Humio/LogScale, Honeycomb, New Relic, Prometheus, Splunk).
- Experience creating, deploying, operating, and scaling applications on Kubernetes.
- Extensive experience deploying and managing data infrastructure at scale (Cassandra, Postgres, MySQL, MongoDB, OpenSearch, Kafka, Redis/Valkey).
- Proficiency in common scripting languages (Python, Go, Bash, PowerShell).
- Experience with storage technologies (SAN, NAS, NFS, Object Storage).
- Experience architecting and deploying big data systems.
- Security-first mindset with a working understanding of cybersecurity principles.
- Proven ability to make well-informed, timely decisions under ambiguity.
- Ability to balance short-term operational needs against long-term strategic goals.
- Self-directed learner who takes initiative in fast-moving environments.
- Must be able to work with a distributed team across multiple time zones.
- Meticulous attention to detail.
- Proven experience utilizing AI technologies to enhance decision-making, streamline workflows and processes, improve efficiency and drive business outcomes.
Bonus Points:
- Knowledge of networking patterns and general network troubleshooting (Load balancers, DNS, VIPS, Routing, Firewall rules).
- Knowledge and proven operation ability across multiple cloud hyperscalers such as AWS, Azure, GCP, Oracle.
- Experience building self-service / provisioning-automation platforms that reduce operational toil.
- Experience with Active Directory / Windows Server and hybrid on-prem + cloud environments.
- Experience with data science, machine learning, and ETL principles and tooling such as Apache Airflow, Apache Spark.
Benefits of Working at CrowdStrike:
- Market leader in compensation and equity awards.
- Comprehensive physical and mental wellness programs.
- Competitive vacation and holidays for recharge.
- Paid parental and adoption leaves.
- Professional development opportunities for all employees regardless of level or role.
- Employee Networks, geographic neighborhood groups, and volunteer opportunities to build connections.
- Vibrant office culture with world class amenities.
- Great Place to Work Certified across the globe.
Compensation:
The base salary range for this position for all U.S. candidates is $140,000 - $215,000 per year, with eligibility for bonuses, equity grants and a comprehensive benefits package that includes health insurance, 401k and paid time off. Placement within the pay range depends on factors including relevant work experience, skills, certifications, job level, supervisory status, and location.
Additional Information:
This role will require the candidate to periodically undergo and pass additional background and fingerprint check(s) consistent with government customer requirements. CrowdStrike is an equal opportunity employer and participates in the E-Verify program.
Skills they ask for
Pick one to see other roles that ask for it.
About CrowdStrike
AI-powered cybersecurity and threat responseCrowdStrike provides cloud-delivered cybersecurity products for detecting and responding to threats across endpoints and other environments.
See all 114 roles at CrowdStrikeMore roles at CrowdStrike
See all 114- Sr. Software Engineer - Sensor - Cloud Runtime Protection (Hybrid)Sunnyvale · Senior · HybridSoftware Development · Senior · HybridSunnyvale, United States1d
- Sr. Engineer - Risk Platform (Hybrid)Sunnyvale · Senior · HybridSoftware Development · Senior · HybridSunnyvale, United States1d
- Product Manager - Advanced Detection (Remote)United States · RemoteProduct Management · RemoteUnited States1d
- Staff Technical Marketing Manager (Staff TMM) – Browser Security (Remote)United States · Staff · RemoteMarketing · Staff · RemoteUnited States1d
Let the right jobs find you
In your inbox every Wednesday and SaturdayPersonalised suggestions from verified career pages, matched to your role, location, level and skills.