Principal Machine Learning Infrastructure Engineer
This role involves building and operating scalable machine learning infrastructure for training and serving large physics models, with a focus on distributed systems, high-performance computing, and efficient data pipelines. The engineer will work closely with research scientists and ML engineers to optimize training workflows on NVIDIA DGX hardware, improve data I/O for mesh-based datasets, and develop reliable model serving solutions for customer deployment. A strong systems background and experience with GPU clusters, PyTorch, and Kubernetes are central to the position.