Software Dev Senior Engineer
Pune, India · Hybrid · Full-time
- Posted 1w ago
- From SonicWall’s careers page
- Location
- Pune, India
- Work mode
- Hybrid
- Type
- Full-time
- Level
- Senior
- Experience
- 5+ years
- Department
- Engineering
Opens the listing on job-boards.greenhouse.io
Let the right jobs find you
In your inbox every Wednesday and SaturdayPersonalised suggestions from verified career pages, matched to your role, location, level and skills.
About the role
Job Description
As a Software Dev Senior Engineer, you will own the reliability, scalability, and operational excellence of our Cloud-based services. You will define and enforce reliability standards, drive the adoption of SRE practices across engineering teams, and build the systems and tooling that keep our production infrastructure healthy. We follow a DevOps model: Development and Operations teams are integrated, and the SRE function acts as the reliability layer — setting Service Level Objectives, managing error alerts, and continuously reducing toil through engineering.
Key Responsibilities:
- Lead on call response and triage – help design and lead the 24x7 response team for triage and rally engineering to address service degradation, facilitating process and communications along the way
- Define, publish, and continuously refine Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Service Level Agreements (SLAs) for all critical services, partnering with product and engineering leadership.
- Own the error management practices, including monitoring and alerting, and surfacing notable errors to the development team
- Lead the design and implementation of comprehensive observability platforms — metrics, structured logging, and distributed tracing — to ensure full visibility into production systems.
- Drive toil reduction initiatives by identifying and automating repetitive, manual operational work, targeting measurable reduction in operational burden across teams.
- Design and execute chaos engineering programs to proactively uncover reliability weaknesses in our infrastructure and services before they impact customers.
- Lead blameless postmortem culture: facilitate incident retrospectives, extract systemic learnings, and track corrective action items to completion.
- Build and improve on-call incident response processes, runbooks, and escalation paths; manage and optimize on-call rotation health to prevent burnout.
- Collaborate with platform and application developers to bring new features and services into production using production-readiness reviews and launch checklists.
- Champion reliability engineering best practices across the organization, embedding SRE principles into the software development lifecycle.
- Mentor team members on SRE philosophy, technical decision-making, code reviews, and cloud engineering best practices.
- Participate in roadmap planning, identify areas of improvement, and perform technology evaluation and selection.
Required:
- 5+ years of experience in database scaling performance and reliability, connection pool optimization
- 5+ years of experience in database HA design and failover strategies
- 3+ years of hands-on Site Reliability Engineering experience, including ownership of SLOs and error budget management.
- 4+ years of experience with Cloud Platforms, GCP CloudSQL required.
- Experience with database capacity planning, load testing, and performance benchmarking at scale.
- Experience managing other database systems — PostgreSQL, Redis, Etc.
- CS Degree or equivalent experience.
Preferred:
- Familiarity with Google SRE principles and the concepts outlined in the Google SRE Book
- Understanding of the HTTP protocol and experience diagnosing distributed system latency issues.
- Experience with instrumentation and management of automated deployments.
- Experience resolving customer-facing production issues under time pressure.
- Experience working with distributed, cross-functional teams
- Experience in infrastructure as code (Terraform, AWS CDK)
- Experience in scripting using Python, Shell, or a similar language.
- Experience with orchestration technologies, including Kubernetes.
- Understanding of networking, including routing, naming, security, network performance, and network failure modes.
- Experience designing and operating observability stacks (e.g., Prometheus, Grafana, Datadog, OpenTelemetry, Jaeger, or equivalent).
- Experience with incident management platforms and on-call tooling (e.g., PagerDuty, OpsGenie).
Skills they ask for
Pick one to see other roles that ask for it.
- Database scaling performance and reliability
- Connection pool optimization
- Database ha design and failover strategies
- Sre
- Cloud kms platforms
- Gcp cloud SQL
- Capacity planning
- Load testing
- Performance benchmarking
- Postgresql
- Redis
- Sre principles
- Http protocols
- Low latency distributed systems
- Instrumentation and management of automated deployments
About SonicWall
Unified cybersecurity for better outcomesSonicWall provides a unified cybersecurity portfolio for SMBs, managed service providers and IT teams across network, endpoint, cloud and threat response.
See all 27 roles at SonicWallMore roles at SonicWall
See all 27- Product Design Engineer- 4+ years of experienceBengaluru · HybridDesign · HybridBengaluru, India1d
- Senior Territory Manager - RemoteSeattle · Senior · RemoteSales · Senior · RemoteSeattle, United States5d
- Senior Solutions Engineer - RemotePittsburgh · Senior · RemoteSales · Senior · RemotePittsburgh, United States5d
- Senior Solutions Engineer - RemotePhiladelphia · Senior · RemoteSales · Senior · RemotePhiladelphia, United States5d
Let the right jobs find you
In your inbox every Wednesday and SaturdayPersonalised suggestions from verified career pages, matched to your role, location, level and skills.