Datacenter & Agentic AI Workload Performance Optimization Engineer
Santa Clara, United States · Remote
- Posted 4d ago
- From Tenstorrent’s careers page
- Location
- Santa Clara, United States
- Work mode
- Remote
- Department
- Engineering
Opens the listing on job-boards.greenhouse.io
Let the right jobs find you
In your inbox every Wednesday and SaturdayPersonalised suggestions from verified career pages, matched to your role, location, level and skills.
About the role
Roles and Responsibilities:
Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities.
Tenstorrent is looking for a Workload Performance Optimization Engineer to help optimize the software workloads that run on our next-generation RISC-V platforms. You’ll work across modern datacenter and agentic AI workloads—including Java, Python, PHP, Node.js, Lua, Go, and Rust—to identify performance bottlenecks and develop software, compiler, runtime, and hardware-aware optimizations that improve throughput, latency, and efficiency. This role sits at the intersection of software runtimes, compilers, CPU microarchitecture, and RISC-V silicon. You’ll bring up and tune major runtimes, profile real-world applications, investigate memory and concurrency behavior, and explore optimizations using RISC-V Vector/Matrix capabilities and custom instructions. You’ll also work with AI-assisted development and automated optimization workflows to accelerate the performance engineering process. Your work will directly influence both the RISC-V software ecosystem and the architecture of future Tenstorrent CPUs.
Who You Are
- You’re a performance engineer who enjoys getting deep into runtimes, compilers, applications, and CPU microarchitecture to understand why software is fast—or slow.
- You have hands-on experience optimizing software on RISC-V or another modern CPU architecture, with a strong understanding of the hardware/software boundary.
- You’re comfortable profiling complex systems, finding bottlenecks, forming hypotheses, and iterating through optimizations using data.
- You’re excited about emerging agentic AI development workflows and using AI tools to automate profiling, coding, benchmarking, and optimization.
- You’re a strong technical collaborator who can work across compiler, runtime, systems software, hardware, and performance modeling teams.
What We Need
- Master’s or PhD in Computer Engineering, Electrical Engineering, Computer Science, or a related field, with strong experience in performance optimization, computer architecture, compilers, or systems software.
- Hands-on experience with runtime or compiler optimization, such as OpenJDK/JIT, LLVM, GCC, V8, Python, or equivalent systems.
- Strong understanding of CPU performance, memory hierarchies, concurrency, garbage collection, vector/SIMD optimization, and RISC-V architecture.
- Expertise with performance profiling and analysis tools such as Linux perf, runtime profilers, QEMU, tracing tools, and performance modeling environments.
- Strong programming skills in Java, Python, C/C++, and RISC-V assembly, with the ability to work effectively across multiple software layers.
What You Will Learn
- How to optimize modern software stacks from application and runtime all the way down to CPU microarchitecture and silicon.
- How RISC-V Vector, Matrix, and custom ISA capabilities can be used to accelerate real-world datacenter and AI workloads.
- How runtime, compiler, memory, and concurrency decisions impact performance at datacenter scale.
- How to build automated and AI-assisted performance optimization workflows that continuously profile, analyze, modify, and benchmark software.
- How to influence future CPU architecture by connecting real workload behavior and software optimization opportunities to hardware design decisions.
Skills they ask for
Pick one to see other roles that ask for it.
About Tenstorrent
Compute for every scaleTenstorrent develops AI computing systems, including superclusters and workstations for running AI workloads.
See all 38 roles at TenstorrentMore roles at Tenstorrent
See all 38- Sr. Staff Engineer, IP RuntimeAustin · Staff · HybridSoftware Development · Staff · HybridAustin, United States3h
- ATE Test Development EngineerAustin · HybridEngineering · HybridAustin, United States1d
- Solution Architect/ Field Application EngineerBengaluru · Senior · RemoteSoftware Development · Senior · RemoteBengaluru, India2d
- IP Customer Program ManagerSanta Clara · Senior · HybridBusiness Operations · Senior · HybridSanta Clara, United States3d
Let the right jobs find you
In your inbox every Wednesday and SaturdayPersonalised suggestions from verified career pages, matched to your role, location, level and skills.