About this role
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators.
At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there.
A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone.
With Roblox Ads & Discovery business growing at a rapid rate, we are building large scale ads machine learning infrastructure to deliver more value to our users and our advertisers.
As a Machine Learning Infrastructure Engineer, you’ll build scalable, reliable, and high-performance infrastructure that powers ML systems across our organization. You’ll operate at the scales of hundreds of billions of engagements, and redefine how we deliver performance ads to hundreds of millions of users.
You will:
• You will co-design models and systems, working at the intersection of model architecture and ML infrastructure, partnering closely with core modelers, data and AI infrastructure engineers, and product teams to push the boundaries of large-scale training and serving. Your work will span recommendation, search, and agentic applications, including large transformer architectures, LLMs, generative rankers, and efficient offline and online content-understanding systems.
• You will investigate model, data, and systems tradeoffs end to end—from data pipelines and distributed training to low-latency inference and production serving. This includes designing efficient KV-cache strategies, applying pruning and quantization, optimizing GPU utilization and memory efficiency, and developing custom kernels where needed.
• You are comfortable working across modern ML systems technologies such as FSDP, vLLM, SGLang, CUDA, distributed training frameworks, inference engines, and GPU kernels, while remaining tool-agnostic and focused on achieving step-function improvements in model quality, throughput, latency, reliability, and cost.
• Lead strategic planning and roadmap execution of scalable production-ready ML systems including model training, data pipelines, feature engineering and model inference.
• Own the architecture, establish engineering best practices of scalability, reliability, and cost-effectiveness of ML infrastructure (e.g., training, serving, feature).
• Work closely with data scientists, ML engineers, platform teams, and product stakeholders to design, implement, and operate robust ML platforms that accelerate model development and deployment.
• Stay ab