About this role
P-1377
Mission
As a Senior ML and AI Technical Solutions Engineer, you play a critical role by helping customers debug and maintain stable GenAI and ML Workloads with AI agent systems using the Databricks Platform. You will develop product expertise end-to-end by advising a broad set of customers and use cases across the space - including products such as Agent Bricks, Vector Search and Model Serving. You will collaborate cross-functionally with other teams - whether that’s working with engineering to improve the product or interacting directly with the account team on a specific customer issue. TSEs have proven production troubleshooting and optimisation experience to help our customers’ workloads run smoothly and to achieve their strategic objectives with ML/AI technology with Databricks. Additionally, you are an early adopter of GenAI technology to improve your own efficiency and amplify the team's output. Reporting to a TSE manager - you will be part of a world class global support engineering organization for Databricks, known for your technical depth and delivering impeccable customer service.
The Impact You Will Have
• Act as senior technical solution expert for complex issues spanning data pipelines, ML pipelines and/or AI applications, applying deep expertise in distributed systems.
• Analyse and troubleshoot production workloads at the code level, optimise for performance, reliability, latency, and cost.
• Diagnose and support Machine Learning and/or Large Language Model deployments, including real-time and batch inference, autoscaling, monitoring, logging, and alerting. Serve as a Subject Matter Expert guiding customers on experiment tracking, model registry, versioning, evaluation, labelling, tracing, and lifecycle observability.
• Provide high-quality support by guiding customers in leveraging Databricks AI to solve generative AI use cases & challenges, leveraging LLMs, MCP, AI Agents, RAG/Agentic RAG, APIs, vector embeddings, semantic search, Vector Search/Lakebase databases, context orchestration, memory management, and prompt engineering.
• Collaborate with internal teams to influence roadmap, product improvements and support business growth.
• Develop expertise in productionizing systems in Databricks and share your knowledge by contributing to wikis and other technical documentation, or by teaching our AI systems new skills, which will be used internally and externally by customers and partners.
What We Look For
8+ years of experience designing, building, and scaling Data, Machine Learning, and AI systems on-premises and in the cloud using Python, Scala, and Java in production environments, with expertise in Machine Learning and/or generative AI. Experience with cloud platforms (AWS, Azure, or GCP); familiarity with Databricks is a plus. Proficient in data engineering necessary for orchestrating end-to-end machine learning training pipelines, ideally with experience processin