Performance Architect, CPU Cluster
Santa Clara, United States · Remote
- Posted 3w ago
- From Tenstorrent’s careers page
- Location
- Santa Clara, United States
- Work mode
- Remote
- Department
- Engineering
Apply on Tenstorrent’s site
Opens the listing on job-boards.greenhouse.io
Let the right jobs find you
In your inbox every Wednesday and SaturdayPersonalised suggestions from verified career pages, matched to your role, location, level and skills.
About the role
Who You Are
- You have a strong foundation in CPU microarchitecture and performance, with an understanding of how cores, caches, interconnects, and memory systems interact.
- You enjoy using modeling, simulation, and workload analysis to understand cluster-level performance bottlenecks and scaling challenges.
- You’re comfortable moving between detailed microarchitecture and broader system-level questions around latency, bandwidth, quality of service, utilization, and scalability.
- You’re analytical and hands-on, with the ability to turn large amounts of performance data into clear architectural recommendations.
- You’re a strong communicator who enjoys working across architecture, RTL, compiler, software, and system teams.
What We Need
- Analyze and optimize CPU cluster performance, including cache hierarchies, interconnects, coherence, and memory access behavior.
- Build and use performance models and simulation environments such as Gem5 or equivalent to evaluate architectural concepts and identify performance opportunities.
- Study real workloads to understand bottlenecks related to cache misses, coherence traffic, memory bandwidth, latency, contention, and core-to-core communication.
- Lead architectural tradeoff studies across performance, scalability, bandwidth, latency, power, and implementation complexity.
- Collaborate with CPU, cache, interconnect, memory, software, and system architects to translate performance analysis into concrete design decisions.
What You Will Learn
- How to architect, model, and optimize a multi-core CPU cluster from early performance exploration through implementation and silicon.
- How cache coherence, interconnects, memory hierarchy, and core architecture interact to determine overall system performance.
- How to model and reason about scaling across cores, including contention, bandwidth limits, latency, and workload behavior.
- How architectural decisions at the cluster level translate into real-world application performance.
- How Tenstorrent approaches scalable CPU and chiplet-based system architecture across compute, interconnect, and memory.
Skills they ask for
Pick one to see other roles that ask for it.
About Tenstorrent
Compute for every scaleTenstorrent develops AI computing systems, including superclusters and workstations for running AI workloads.
See all 38 roles at TenstorrentMore roles at Tenstorrent
See all 38- Sr. Staff Engineer, IP RuntimeAustin · Staff · HybridSoftware Development · Staff · HybridAustin, United States4h
- ATE Test Development EngineerAustin · HybridEngineering · HybridAustin, United States1d
- Solution Architect/ Field Application EngineerBengaluru · Senior · RemoteSoftware Development · Senior · RemoteBengaluru, India2d
- IP Customer Program ManagerSanta Clara · Senior · HybridBusiness Operations · Senior · HybridSanta Clara, United States3d
Share this role
Let the right jobs find you
In your inbox every Wednesday and SaturdayPersonalised suggestions from verified career pages, matched to your role, location, level and skills.