As a member of the GPU AI/HPC Infrastructure team, you will provide leadership in the design and implementation of groundbreaking GPU compute clusters that powers all AI research across NVIDIA. We seek an expert to build and operate these clusters at high reliability, efficiency, and performance and drive foundational improvements and automation to improve researchers productivity. As a Site Reliability Engineer, you are responsible for the big picture of how our systems relate to each other, we use a breadth of tools and approaches to tackle a broad spectrum of problems. Practices such as limiting time spent on reactive operational work, blameless postmortems and proactive identification of potential outages factor into iterative improvement that is key to both product quality and interesting dynamic day-to-day work. SRE’s culture of diversity, intellectual curiosity, problem solving and openness is important to our success. Our organization brings together people with a wide variety of backgrounds, experiences and perspectives. We encourage them to collaborate, think big and take risks in a blame-free environment. We promote self-direction to work on meaningful projects, while we also strive to build an environment that provides the support and mentorship needed to learn and grow.
What you’ll be doing:
In this role you will be building and improving our ecosystem around GPU-accelerated computing including developing large scale automation solutions. You will also be maintaining and building deep learning AI-HPC GPU clusters at scale and supporting our researchers to run their flows on our clusters including performance analysis and optimizations of deep learning workflows. You will design, implement and support operational and reliability aspects of large scale distributed systems with focus on performance at scale, real time monitoring, logging, and alerting.
What we need to see:
Ways to stand out from the crowd:
#LI-Hybrid
Join the Revolution at Leonardo.Ai! Leonardo.Ai, an Australian tech startup, is on a transformative mission to democratise design and ignite...
How to applyNO NEED TO PREDICT THE FUTURE. YOU CAN CREATE IT. SHARE YOUR PASSION. World-leading technologies don’t make it into a...
How to applyScientist – Machine Learning At ABB, we are dedicated to addressing global challenges. Our core values: care, courage, curiosity, and...
How to applyEvery day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in...
How to applyBackgroundWe are seeking a motivated and talented graduate student to undertake a master thesis project on the topic of “Generative...
How to applyScale works with the industry’s leading foundation model labs to provide high quality data and accelerate progress in machine learning...
How to apply