Senior Software Engineer, Orchestration Platform
San Francisco, United States · Full-time
- Posted 1w ago
- From Scale AI’s careers page
- Location
- San Francisco, United States
- Type
- Full-time
- Level
- Senior
- Experience
- 5+ years
- Department
- Engineering
Apply on Scale AI’s site
Opens the listing on job-boards.greenhouse.io
Let the right jobs find you
In your inbox every Wednesday and SaturdayPersonalised suggestions from verified career pages, matched to your role, location, level and skills.
About the role
Roles and Responsibilities:
- Lead the architecture, design, implementation, and operation of Scale’s core orchestration platform.
- Build durable workflow infrastructure using technologies such as Temporal, Cadence, Kubernetes, and cloud-native systems.
- Define platform primitives and APIs for scheduling, retries, state management, task execution, observability, and workflow lifecycle management.
- Partner with product, infrastructure, data, and application teams to understand workflow needs and turn them into reusable platform capabilities.
- Improve the reliability, scalability, security, and developer experience of services that run critical company workflows.
- Establish technical standards and best practices for distributed workflow development, deployment, testing, and incident response.
- Drive cross-functional technical decisions and communicate platform direction clearly to engineers and stakeholders.
Ideal Candidate:
- 5+ years of full-time software engineering experience, with a focus on backend, infrastructure, and distributed systems.
- Experience building and operating production systems with strong requirements for reliability, availability, and scale.
- Deep familiarity with workflow orchestration platforms such as Temporal, Cadence, AWS Step Functions, Kubernetes, or similar systems.
- Deep familiarity with Kubernetes and containerized production environments, and familiarity with Terraform, Docker.
- Strong knowledge of distributed-systems concepts, including asynchronous execution, retries, idempotency, fault tolerance, state management, and observability.
- A track record of leading technically complex projects from design through rollout and ongoing operation.
- Excellent communication skills and the ability to collaborate effectively with platform consumers and non-technical stakeholders.
Nice to Haves:
- Experience with data warehouses (Snowflake, Firebolt) and data pipeline/ETL tools (Dagster, dbt).
- Experience with authentication/authorization systems (Zanzibar, Authz, etc).
- Experience scaling products at hyper-growth startups.
- Excitement to work with AI technologies.
Skills they ask for
Pick one to see other roles that ask for it.
About Scale AI
Reliable AI systems for critical decisionsScale AI develops data and AI systems used to support reliable AI and critical decisions across the AI stack.
See all 35 roles at Scale AIMore roles at Scale AI
See all 35- Systems Engineering Integration & Test Lead, Public SectorWashington · LeadEngineering · LeadWashington, United States15h
- Enterprise AI Development StrategistNew York · Entry LevelSales · Entry LevelNew York, United States2d
- Product Operations Lead, Generative AISan Francisco · SeniorBusiness Operations · SeniorSan Francisco, United States1w
- Revenue Operations ManagerSan Francisco · SeniorBusiness Operations · SeniorSan Francisco, United States1w
Share this role
Let the right jobs find you
In your inbox every Wednesday and SaturdayPersonalised suggestions from verified career pages, matched to your role, location, level and skills.