PB✓
PBridge

About this role

PagerDuty (NYSE:PD) is a leader in Digital Operations Management. In an always-on world, organizations of all sizes trust PagerDuty to help them deliver a perfect digital experience to their customers, every time. Teams use PagerDuty to identify issues and opportunities in real time and bring together the right people to fix problems faster and prevent them in the future. Over 13,000 organizations (including 60 of Fortune 100) rely on PagerDuty to succeed with Digital Transformation, Cloud Migration, and DevOps Modernization. Notable customers include GE, Cisco, Genentech, Electronic Arts, Cox Automotive, Netflix, Shopify, Zoom, DoorDash, Lululemon and more. We are expanding rapidly as a platform for Digital Operations Management using AI/ML and Automation and growing our adoption by Development, IT, Customer Service, Security, and other teams across the organization.

About the role

PagerDuty’s Operations Cloud runs on a platform that ingests billions of signals and turns them into real-time action for thousands of customers. We’re looking for a Senior AI/ML Engineer who lives at the intersection of two disciplines: large-scale distributed systems and applied AI.

In this role you will design and ship AI systems that run in production at PagerDuty’s scale — powering Incident Management AI Agents, event intelligence, and the LLM-powered capabilities embedded across our platform. You’ll own the full lifecycle, from framing the problem to serving reliably at scale.

We are looking for a candidate who is genuinely passionate about building with modern AI — LLMs, agents, and retrieval — but grounded in the realities of building resilient, high-throughput systems.

What you’ll do

• Design and build AI-powered features — LLM agents, retrieval, and event intelligence — that operate on high-volume, real-time event streams, from problem framing through production deployment and monitoring.

• Architect and own the systems behind them: agent and prompt orchestration, retrieval pipelines, tool/API integrations, and low-latency inference and evaluation at scale.

• Reason about consistency, throughput, fault tolerance, and cost across services that must stay reliable under bursty, unpredictable load.

• Take AI features from prototype to production, establishing the evaluation, guardrail, observability, and improvement loops that keep them accurate and trustworthy over time.

• Partner with platform, product, and applied-research teams to define what “good” looks like and to integrate AI cleanly into existing services.

• Raise the bar through example, reviews and mentorship, and help shape the team’s technical direction.

What you’ll bring

• 5+ years of software engineering experience, with meaningful time spent building and operating production distributed systems (high-throughput services, streaming/event-driven architectures, or large-scale data platforms).

• Hands-on experience building and shipping AI systems in production — LLM-powered application

Tired of applying one by one?

Our Career Success Team finds roles in Portugal that fit you, tailors your CV to each, and submits the applications — tracked end to end. You just show up to interviews.

We apply, you interview →