Engineering Manager
Bengaluru, India · Remote · Full-time
- Posted 1w ago
- From GitLab’s careers page
- Location
- Bengaluru, India
- Work mode
- Remote
- Type
- Full-time
- Level
- Senior
- Experience
- 3+ years
- Department
- Engineering
Opens the listing on job-boards.greenhouse.io
Let the right jobs find you
In your inbox every Wednesday and SaturdayPersonalised suggestions from verified career pages, matched to your role, location, level and skills.
About the role
An overview of this role
You'll lead the globally distributed Observability team. The team builds and operates the metrics, logging, alerting, and capacity planning platforms that GitLab engineers use to understand GitLab.com and GitLab Dedicated. You'll help determine how the team collects, stores, queries, and acts on telemetry, balancing reliable signals with scale and cost.
What you’ll do
- Lead, hire, onboard, and develop a distributed engineering team working asynchronously.
- Set priorities with Site Reliability Engineering, Product Engineering, and GitLab Dedicated teams, and help the team deliver observability services iteratively.
- Own the reliability, scalability, and cost of the team's metrics, logging, alerting, and capacity planning platforms.
- Reduce noisy or missing alerts and telemetry gaps, and use SLOs, error budgets, and self-service instrumentation to help engineers maintain the health of their services.
- Guide technical decisions about time-series storage, high-cardinality metrics, log pipelines, and distributed tracing.
- Participate in the Incident Manager On Call (IMOC) rotation, coordinating the response to high-severity incidents affecting GitLab.com.
- Keep the team's on-call rotation sustainable through coverage across time zones, useful runbooks, better alerts, and follow-through on post-incident actions.
- Use AI tools and agents to support engineering workflows and incident triage, reviewing their output while engineers retain responsibility for decisions.
What you’ll bring
- Experience leading an observability, platform engineering, or site reliability engineering team operating at scale, including supporting people in a distributed, asynchronous environment.
- Technical knowledge of metrics systems such as Prometheus and long-term storage, logging platforms such as Elasticsearch or cloud-native services, and alerting design.
- Experience using SLOs, error budgets, and capacity forecasts to make reliability and investment decisions.
- Experience operating a large software-as-a-service platform and investigating production issues such as telemetry gaps, ingestion limits, or noisy and missing alerts.
- Experience participating in and improving production on-call rotations, including incident coordination and balancing operational load with project work.
- The ability to explain technical tradeoffs to engineering partners and other stakeholders.
- Experience using AI tools or agents in engineering or management work; you can describe how you would apply them to operational problems such as incident triage.
We welcome different paths into this role, whether through practical experience, formal study, or transferable skills. If the work interests you, please apply even if you don't meet every qualification.
Skills they ask for
Pick one to see other roles that ask for it.
About GitLab
One platform for software development and deliveryGitLab provides a software development platform covering source control, CI/CD, security, and AI-assisted software workflows.
See all 93 roles at GitLabMore roles at GitLab
See all 93- People Systems EngineerBengaluru · RemoteInformation Technology · RemoteBengaluru, India3h
- People Technology Analyst - WorkdayUnited States · RemoteHuman Resources · RemoteUnited States3h
- Lead Legal Counsel, CorporateUnited States · Lead · RemoteLegal and Compliance · Lead · RemoteUnited States4h
- Benefits AnalystBengaluru · HybridHuman Resources · HybridBengaluru, India9h
Let the right jobs find you
In your inbox every Wednesday and SaturdayPersonalised suggestions from verified career pages, matched to your role, location, level and skills.