About this role
Reddit is a community of communities. It’s built on shared interests, passion, and trust, and is home to the most open and authentic conversations on the internet. Every day, Reddit users submit, vote, and comment on the topics they care most about. With 100,000+ active communities and approximately 130 million daily active unique visitors, Reddit is one of the internet’s largest sources of information. For more information, visit www.redditinc.com .
Reddit has a flexible workforce! If you happen to live close to one of our physical office locations our doors are open for you to come into the office as often as you'd like. Don't live near one of our offices? No worries: You can apply to work remotely in any country in which we have a physical presence.
About the Role
We’re building a scalable feature platform that powers Ads ML by making high-quality features and training datasets easy to build, share, and maintain. Our small but growing team works on projects like batch & realtime feature management platform, training set generation platform, sequence features platform and, agentic and automated ML workflows for feature lifecycle management.
We are looking for an engineer with deep experience in building high-scale data infrastructure and ML platforms to help evolve and scale our feature management systems.
What You’ll Do
• Design and build data infrastructure that supports large-scale feature and training set computation, transformation, and storage.
• Develop frameworks for batch and real-time features with a focus on reliability, scalability, and ease of use.
• Build platform capabilities for feature governance, including lineage tracking, validation, drift detection, anomaly monitoring, reproducibility, and versioning
• Partner with ML engineers to ensure smooth integration of feature engineering workflows into ML production systems.
• Build systems that support agentic ML workflows, including automated feature discovery, feature quality evaluation and feature lifecycle management
• Drive architecture and technical direction for high-scale ML platform infrastructure powering ads ranking, targeting, and optimization systems.
• Contribute to operational excellence through observability, performance tuning, reliability engineering, and cost optimization initiatives.
What You Bring
• 5+ years in infrastructure/platform engineering or large-scale distributed systems.
• 2+ years of hands-on experience building or operating ML platform infrastructure and production ML systems
• Proficiency with large-scale feature computation frameworks (Spark, PySpark, or Scala).
• Expertise in distributed systems (scaling, partitioning, fault tolerance, caching).
• Experience building intelligent automation or agentic workflows for ML systems is a strong plus, including areas such as automated pipeline management, feature recommendations, evaluation systems, or AI-assisted developer tooling.
• Experience with ML infrastructure and MLOps w