← All roles

Join Orchestra

Staff Infrastructure Engineer

Build the systems behind reliable model serving, routing, and training.

The role

Orchestra helps teams put the right model to work, then improve it against the tasks that matter to their business. That depends on infrastructure people can trust: a request has to reach the right model, an evaluation has to be reproducible, and a rollout has to preserve a working fallback.

You’ll build the systems behind model serving, routing, and the path from a promising experiment to production. This is a hands-on role for someone who can make consequential architecture decisions and follow them through to working software.

What you’ll work on

  • Build and improve the inference gateway, including streaming, provider integrations, timeouts, retries, and failure handling.
  • Make model routes observable. Help teams understand latency, errors, usage, and the cost of completing a task.
  • Build reliable execution for evaluation, optimization, and specialist-training workloads, with clear state and recoverable failures.
  • Make it safe to introduce a new route through comparison, bounded rollout, and fallback.
  • Protect customer data through careful access boundaries and explicit handling of traces and model artifacts.
  • Find the bottlenecks that matter, measure them, and simplify the systems around them.

What you’ll bring

  • Experience designing, shipping, and operating distributed systems that other people depend on.
  • Strong judgment about reliability, performance, and the tradeoffs between building a durable foundation and shipping the next useful capability.
  • The ability to diagnose a problem across application code, infrastructure, and third-party services.
  • Clear technical communication: you can explain a decision, make its risks visible, and help others work effectively in the system.
  • Ownership from design through operation, including the less glamorous work of debugging and recovery.

Experience with inference serving, GPU workloads, model providers, or asynchronous job systems would be useful. We’re equally interested in deep infrastructure experience from another domain and evidence that you learn quickly.

Tell us what you’ve built.

Tell us about a system you built or operated, a difficult failure you worked through, and a technical decision you would make differently today. A short write-up, code sample, or relevant project is welcome. Please omit confidential details from past employers or customers.

Express interest