Senior Automation & Observability Engineer
United States · Remote · Full-time
- Posted 3w ago
- From Ensono’s careers page
- Location
- United States
- Work mode
- Remote
- Type
- Full-time
- Level
- Senior
- Experience
- 7+ years
- Department
- Information Technology
Opens the listing on ensono.com
Let the right jobs find you
In your inbox every Wednesday and SaturdayPersonalised suggestions from verified career pages, matched to your role, location, level and skills.
About the role
About the role and what you'll be doing:
We are seeking an experienced IoT / Observability Engineer responsible for monitoring, managing, automating, and optimizing enterprise infrastructure, applications, IoT platforms, and enterprise telemetry ecosystems. The ideal candidate will possess strong expertise in observability platforms, monitoring technologies, automation frameworks, and operational support processes to ensure high availability, reliability, and performance of business-critical systems.
Key Responsibilities
Monitoring & Observability
- Design, implement, and maintain enterprise monitoring and observability solutions.
- Develop and maintain dashboards, alerts, and visualizations using Grafana.
- Monitor infrastructure, applications, middleware, and IoT services using IBM Instana, Grafana, SolarWinds, and related observability tools.
- Configure and manage data collection using Telegraf, Prometheus, and monitoring agents.
- Analyze metrics, logs, traces, events, and telemetry data to identify performance bottlenecks and service degradation.
- Support SLO, SLA, and operational health monitoring initiatives.
- Perform Root Cause Analysis (RCA) and troubleshooting for infrastructure and application issues.
FOAK & Enterprise Logging/Telemetry
- Support onboarding, monitoring, and operational management of FOAK (First Office Application Kit) services and enterprise applications.
- Configure, validate, and troubleshoot Enterprise Logging & Telemetry (ELT) integrations across infrastructure, middleware, applications, and cloud platforms.
- Monitor telemetry pipelines, log ingestion, event correlation, and data quality to ensure complete observability coverage.
- Collaborate with engineering teams to improve telemetry standards, monitoring effectiveness, and proactive incident detection through ELT and observability frameworks.
- Support FOAK application integrations with Grafana, Instana, Prometheus, and enterprise monitoring platforms.
Infrastructure & Platform Monitoring
- Monitor and support:
- Linux Servers
- Windows Servers
- VMware Infrastructure
- Citrix VDI Platforms
- DNS Services
- Proxy Services
- Middleware Platforms
- Integration Services
- Enterprise Applications
- IoT Platforms
Additional Responsibilities
- Investigate performance issues, recurring alerts, and infrastructure anomalies.
- Validate monitoring platform health and monitoring coverage.
- Monitor capacity, availability, CPU, memory, storage, and service health metrics.
- Support platform upgrades, maintenance, and operational readiness reviews.
Database & Data Management
- Configure and maintain InfluxDB time-series databases.
- Manage data retention policies, performance tuning, and capacity planning.
- Develop operational dashboards and reports for infrastructure and application performance insights.
- Support telemetry data ingestion, storage optimization, and historical trend analysis.
Event & Incident Management
- Monitor operational alerts, events, notifications, and incidents from enterprise monitoring platforms.
- Acknowledge, investigate, troubleshoot, and resolve assigned incidents.
- Coordinate with Infrastructure, Network, Cloud, Security, Application, and Service Delivery teams during incident resolution.
- Participate in major incident bridges, DR exercises, and 24x7 operations support activities.
- Follow escalation procedures, SOPs, operational runbooks, and ITIL processes.
- Support Problem Management activities and contribute to RCA documentation.
Instana & APM Operations
- Administer and support IBM Instana monitoring environments.
- Monitor application, API, middleware, and microservices performance using Instana.
- Validate Instana agent health following server patching and maintenance activities.
- Support application onboarding and APM configuration standards.
- Configure alerts, baselines, and performance thresholds.
- Raise and track incidents related to Instana platform availability and performance.
Automation & Scripting
- Develop automation solutions using:
- Python
- PowerShell
- Shell Scripting (Bash)
- VBScript
Configuration Management & Infrastructure Automation
- Implement Infrastructure as Code (IaC) and automation using:
- Ansible
- Puppet
Skills they ask for
Pick one to see other roles that ask for it.
About Ensono
IT modernization, cloud and managed servicesEnsono provides managed IT, cloud, mainframe, data, AI, modernization, and security services to help organizations operate and transform technology environments.
See all 54 roles at EnsonoMore roles at Ensono
See all 54- Client Engagement Director- Preferred candidates Chicago land area / MidwestUnited States · Director · RemoteSales · Director · RemoteUnited States4h
- Client Engagement Director- Preferred location New York, Chicago, MidwestUnited States · Director · RemoteSales · Director · RemoteUnited States4h
- Applications Support SpecialistPune · HybridInformation Technology · HybridPune, India17h
- Senior Project ManagerChennai · SeniorProject and Program Management · SeniorChennai, India18h
Let the right jobs find you
In your inbox every Wednesday and SaturdayPersonalised suggestions from verified career pages, matched to your role, location, level and skills.