PB✓
PBridge

About this role

We’re big believers in the power of IRL, so for most roles we ask Campers to work from their local Culture Amp office an average of 2 days a week to unlock connection, pace and culture together.

Join us on our mission to make a better world of work. 

Culture Amp is the world’s leading employee experience platform, revolutionizing how 25 million employees across more than 6,000 companies create a better world of work. Culture Amp empowers companies of all sizes and industries to transform employee engagement, drive performance management, and develop high-performing teams. Powered by people science and the most comprehensive employee dataset in the world, the most innovative companies including Canva, On, Asana, Dolby, McDonalds and Nasdaq depend on Culture Amp every day.

Culture Amp is backed by leading venture capital funds and has offices in the US, UK, Germany and Australia. Culture Amp has been recognized as one of the world’s top private cloud companies by Forbes and most innovative companies by Fast Company.

For more information visit cultureamp.com .

How you can help make a better world of work

Shipping an AI product is only the beginning. The harder challenge, one few teams have mastered is continuous production evaluation: diagnosing performance shifts in real-time, decoding 'why' they occur, and driving measurable quality improvements at scale. We are looking for a Staff level Applied AI Scientist with a strong AI Engineering background to solve this problem for our Coach AI system, establishing the observability and evaluation frameworks that turn early production releases into robust, high-performance production products and then to make this sustainable by enabling the rest of our engineering org to do the same.

As part of this team of amazing humans,

You will

• Own the end-to-end feedback loop: establish a rigorous cycle of prompt engineering, evaluation at scale, and continuous improvement. You will build LLM-powered analysis tools that diagnose performance shifts, provide deep-dive insights, and automate recommendations for prompt or system-level enhancements.

• Contribute to Context engineering:  design and optimise what actually enters the model's context: retrieval, memory across sessions, context assembly and compression, and managing context budget in long or multi-turn agentic flows. Validate each change against eval rather than opinion or adhoc testing.

• Design and run evals: sampling, LLM-as-a-judge, and labelling systems over de-identified production traces (for example, with Langfuse) to build longitudinal evaluation monitoring and alerting.

• Eval-driven agentic orchestration: contribute to the agent architecture (planning, tool use, routing, decomposition, verification/critique steps) and let eval findings drive structural changes — e.g. when a failure mode surfaces, add a self-check step, change tool selection, or re-route.

• Model and provider selection: make and own model/routing decision

Tired of applying one by one?

Our Career Success Team finds roles in Australia that fit you, tailors your CV to each, and submits the applications — tracked end to end. You just show up to interviews.

We apply, you interview →