Operations Specialist
OEC Group
Full Time3+ yearsPosted 23 days ago
Let the right jobs find you
Get personalised suggestions from verified company career pages, matched to your role, location, level, and skills.
Overview
Position Type
Full Time
Experience
3+ years
Job Description
About the Role:
OEC is in the middle of a multi-year journey to modernise its infrastructure and applications, transforming how we build, deploy, and operate software at global scale. The Operations Specialist (L2) is a core member of the Monitoring Team within the Enterprise Operations Team, responsible for the configuration, maintenance, and continuous improvement of OEC’s observability stack, primarily powered by Datadog.
Key Responsibilities & Duties (essential to the job)
Monitoring & alerting
- Design, configure, and maintain Datadog monitors, composite alerts, and notification channels for production and non-production environments.
- Own the L2 alert triage process — investigate, diagnose, and resolve escalated alerts from L1, ensuring timely root cause identification.
- Tune alert thresholds and suppression rules to minimise noise and reduce false positive rates.
- Manage SLO (Service Level Objective) definitions, tracking, and reporting across assigned services.
- Participate in the on-call rotation for critical monitoring alerts and P1/P2 incident response.
Dashboards & observability
- Build and maintain Datadog dashboards covering infrastructure health, application performance, log analytics, and business KPIs.
- Configure and manage Datadog APM (Application Performance Monitoring), distributed tracing, and error tracking for key services.
- Implement and manage log ingestion pipelines, parsing rules, and log-based monitors.
- Develop and maintain Synthetic monitors (API and browser tests) for uptime and user experience validation.
Service onboarding & integrations
- Onboard new services and infrastructure components into the Datadog monitoring framework in collaboration with DevOps, Cloud, and Engineering teams.
- Configure and maintain Datadog integrations with cloud platforms (AWS, Azure, Rackspace), Kubernetes, containerised workloads, and third-party tools.
- Support the migration of legacy monitoring tooling (SCOM, Dynatrace, Grafana, Prometheus, Pingdom, Apigee, Redgate, IDERA) into Datadog.
Incident response & runbooks
- Act as L2 escalation point during incidents — perform deep-dive investigations using metrics, logs, traces, and dashboards.
- Create, maintain, and improve runbooks for common alert scenarios, incident response procedures, and post-incident remediation steps.
- Contribute to post-incident reviews, documenting root cause findings and identifying monitoring gaps to prevent recurrence.
- Alerts system owners, stakeholders, and management of degraded system status and Priority 1 and Priority 2 incidents; issues tickets for incident and problem escalation.
AI & automation
- Leverage Datadog AI capabilities including Watchdog, anomaly detection, outlier detection, and Bits AI to improve proactive issue detection.
- Implement forecast-based monitors for capacity management (CPU, memory, disk).
- Use AI-based alert correlation and deduplication features to reduce alert fatigue.
- Contribute to automation of repetitive monitoring tasks using Datadog’s API and scripting tools.
Standards & governance
- Adhere to and actively contribute to monitoring standards, tagging strategies, and naming conventions.
- Participate in regular alert quality reviews, dashboard audits, and SLO compliance checks.
- Maintain accurate documentation of monitoring configurations, integrations, and architectural decisions.
- Creates and maintains knowledge articles to be used internally and/or externally for training, best practices, solutions, or processes relating to applications, environments, and related technologies.
Infrastructure operations & support
- Analyses and troubleshoots Microsoft and UNIX/Linux server configurations and processes.
- Performs moderately complex database administration.
- Monitors data centre networks, infrastructure bandwidth, application health, servers, and other infrastructure; coordinates and communicates with vendors.
- Diagnoses and researches (using knowledge base) application incidents, monitoring alerts, and service requests; provides assistance and guidance to associate operations specialists.
- Adheres to all incident and service request processes and procedures in accordance with established Service Level Agreements (SLAs).
- Diagnoses, researches, and resolves Level-2 technical hardware and software incidents, monitoring alerts, and service requests.
- Works on ad-hoc projects to support Infrastructure Engineering or other departments, as requested.
- Demonstrates a flexible and adaptable approach to work and adjusts to shifts in priorities as the needs of the business change.
- Collaborates with DevOps, Cloud, and application teams to align monitoring coverage with service requirements.
Required Skills
DatadogIt Cloud InfrastructureContainerized EnvironmentsObservabilityScripting LanguagesMonitoring And Alerting ToolsItsm ConceptsMicrosoft And Unix Linux Server ConfigurationsAnalytical Problem SolvingCommunication SkillsCollaborative Team PlayerContinuous ImprovementFlexible And Adaptable