Technical Lead-Cloud & Infra Engg
Birlasoft
Full TimeNot specifiedPosted about 5 hours ago
Let the right jobs find you
Get personalised suggestions from verified company career pages, matched to your role, location, level, and skills.
Overview
Position Type
Full Time
Experience
Not specified
Job Description
Area(s) of responsibility
SRE DevOps/Middleware
Remote
Fulltime
Tools: AWS, Nginx, Akamai, Weblogic, Tomcat, Apache, Splunk, New Relic, Shell scripting, Service Now, Bitbucket, Jenkins, Terraform, Docker, Kubernetes, UDeploy and Jira.
Key Responsibilities:
Site Reliability Engineering (SRE)
- Ensure high availability, performance, and resilience of production systems.
- Implement SRE best practices: error budgets, SLIs/SLOs, capacity planning, chaos testing, runbook creation.
- Drive automation to reduce manual operational tasks and improve MTTR.
- Conduct post-incident reviews (PIRs) and implement long-term corrective actions.
Middleware & Application Platform Management
- Manage deployments, rollbacks, and environment synchronization across Dev, QA, UAT, and Production.
- Install, configure, upgrade, and maintain WebLogic, Tomcat, Apache, and Nginx servers. Perform JVM tuning, thread pool optimization, connection pool management, and performance tuning.
- Troubleshoot middleware issues related to memory leaks, thread contention, SSL, certificates, and clustering.
Monitoring, Logging & Observability
- Configure dashboards, alerts, and performance insights using New Relic and Splunk.
- Develop log-based monitoring strategies and anomaly detection.
- Implement proactive monitoring to reduce downtime and improve reliability.
CDN & Edge Platform Management (Akamai)
- Configure Akamai caching rules, WAF policies, edge redirects, and performance optimizations.
- Troubleshoot CDN-related latency, caching, and routing issues.
- Collaborate with Akamai support for advanced troubleshooting.
Incident, Problem & Change Management
- Lead major incident bridges, coordinate cross functional teams, and provide timely updates.
- Manage problem tickets, root cause analysis, and preventive action plans.
- Ensure compliance with ITIL processes for change, release, and incident management.
Leadership & Stakeholder Management
- Lead and mentor a team of SRE/DevOps engineers.
- Provide technical guidance, training, and performance feedback.
- Act as a customer facing technical SME for escalations and production issues.
- Collaborate with product, QA, development, and business teams to ensure smooth delivery.
Behavioral Competencies
- Ownership & Accountability: Takes responsibility for production stability and issue resolution.
- Leadership: Guides team members, manages workload, and drives operational excellence.
- Communication: Clear, structured communication with customers and internal teams.
- Problem Solving: Strong analytical skills and ability to troubleshoot complex issues.
- Collaboration: Works effectively across engineering, QA, product, and business teams.
- Calm Under Pressure: Handles critical incidents with composure and clarity.