← All roles

Join Orchestra

AI Research Scientist

Develop methods for evaluation, post-training, and learning from expert feedback.

The role

Orchestra helps teams turn recurring work and expert judgment into better evaluations, model routes, and specialist models. We start with a real task and a clear quality bar, then compare the changes that could improve it: a better prompt, a different model, or targeted training.

You will develop and test methods for learning from production examples and expert corrections. The work spans evaluation design, post-training, and model selection. Your research should make it easier to answer a practical question: which intervention improves this workload, and what evidence do we need before deploying it?

What you'll work on

  • Design task-specific evaluations that measure useful outcomes, including structured outputs, tool use, and actions that change downstream state.
  • Turn expert feedback into rubrics, training examples, and reward signals; investigate where those signals are incomplete or easy to exploit.
  • Develop and compare prompt optimization and post-training approaches for specialist models, using appropriate frontier and open-model baselines.
  • Build reproducible experiments with held-out data, versioned artifacts, and clear accounting for quality, latency, and cost per task.
  • Study failure modes, distribution shifts, and the conditions under which a specialist should defer to another model.
  • Work with engineering to bring promising methods into the product, with evaluation criteria that survive the move from experiment to production.
  • Write up findings clearly, including negative results, limitations, and what remains unproven.

What you bring

  • Experience designing and running machine learning experiments that other people can reproduce and scrutinize.
  • Depth in one or more of language-model evaluation, fine-tuning, reinforcement learning, preference learning, or model selection.
  • Strong Python skills and the ability to debug data pipelines, training runs, and evaluation harnesses.
  • Good judgment about data leakage, misleading metrics, confounded comparisons, and the difference between a benchmark improvement and a useful product improvement.
  • A habit of moving between research questions and working code, and of explaining technical tradeoffs to people outside your specialty.

Research publications are useful evidence, but so are rigorous independent experiments, open-source contributions, and methods you have put into practice. We care about the quality of your reasoning and work.

Tell us what you’ve built.

Tell us about an experiment or system you helped build, what you learned from it, and what you would investigate at Orchestra. Share a paper, repository, write-up, or another example of your work if you have one. Please leave out confidential data and anything you do not have permission to share.

Express interest