Senior Systems Engineer
Graphcore
Full Time7+ yearsPosted 8 days ago
Let the right jobs find you
Get personalised suggestions from verified company career pages, matched to your role, location, level, and skills.
Overview
Position Type
Full Time
Experience
7+ years
Job Description
Job Summary
We are seeking a Staff Hardware Engineer to provide advanced operational, diagnostic, and engineering support for Graphcore’s Arm-based hardware platforms across lab and data center environments.
This role focuses on supporting hardware bring-up, validation, and troubleshooting of complex AI compute platforms, including server blades, racks, and rack-scale infrastructure. The successful candidate will collaborate closely with engineering, platform, and data center teams to ensure the reliability and performance of next-generation AI systems.
Responsibilities and Duties
- Lead advanced break-fix troubleshooting for server blades, motherboards, power systems, and rack-scale infrastructure.
- Support engineering bring-up activities, including component validation and firmware interaction testing.
- Diagnose system-level failures involving thermal behavior, power anomalies, network configuration, and BIOS/BMC issues.
- Collaborate with server engineering teams to perform root cause analysis and propose corrective actions or design improvements.
- Support deployment and rollout of next-generation hardware platforms through structured validation and qualification cycles.
- Interface with facilities and infrastructure teams to understand environmental factors impacting system reliability.
- Develop and maintain standard operating procedures (SOPs), troubleshooting guides, and validation documentation.
- Provide guidance and mentorship to junior technicians and engineers on troubleshooting methodologies and hardware diagnostics.
- Participate in on-call rotations or off-hours support during critical engineering milestones or hardware bring-up phases.
Candidate Profile
Essential
- Bachelor’s degree in Electrical Engineering, Computer Engineering, Computer Science, or related discipline.
- 7 years experience with server hardware architectures and board-level debugging.
- Experience analyzing system logs, hardware telemetry, and power/thermal metrics to isolate hardware failures.
- Hands-on experience with HPC systems, AI compute platforms, or rack-scale infrastructure.
- Strong collaboration skills and ability to work effectively in fast-paced engineering environments.
- Excellent written and verbal communication skills.
Desirable
- Experience supporting prototype or pre-production hardware bring-up.
- Familiarity with data center facilities, including liquid cooling and power distribution systems.
- Experience using Python, Bash, or automation tools for hardware validation or troubleshooting.
- Exposure to structured failure analysis and reliability engineering methodologies.
Required Skills
Hardware ArchitectureBoard Level DebuggingSystems AnalysisHardware TelemetryPower Thermal MetricsHpc SystemsAi Compute PlatformsRack InfrastructureCollaborationCxo CommunicationPythonBashAutomation ToolsStructured Failure AnalysisSre