This includes a pipeline that transforms real-world scans into usable 3D scenes through Gaussian Splatting, as well as a browser-based streaming platform built on NVIDIA Omniverse and Isaac Sim. Both technologies work today, but neither is yet operated as a product. Your mission is to change that.
The first step is internal standardization. We want to replace individually maintained environments and hand-crafted deployments with one reproducible, documented solution that teams can use independently: scene in, session out, without requiring the Simulation team to operate every setup.
The second step is external productization. The same capabilities need to be packaged so that partners and customers can operate them in their own environments.
E.g. with versioned releases, stable interfaces, tenant isolation, diagnostics, upgrade paths, and documentation.
You will join the Simulation team and focus on the infrastructure and operational foundations of these systems. You will not be building them in isolation: Wandelbots has an established infrastructure team responsible for shared infrastructure and platform capabilities. You will work closely with that team while owning the simulation-specific workloads, deployment patterns, and operational requirements.
This is a hands-on engineering role with room to shape architecture and make technical decisions in collaboration with the teams involved.
What you will work on
Build and operate the GPU infrastructure required by our simulation workloads, including NVIDIA GPU Operator, device plugins, driver lifecycle, and GPU sharing through MIG or time slicing
Work with the infrastructure team to integrate simulation workloads into our shared Kubernetes and cloud infrastructure
Make deployments declarative and reproducible using Infrastructure as Code and GitOps, and help consolidate the tooling used across today’s environments
Own CI/CD and the container and image lifecycle for CUDA- and NGC-based workloads
Design scheduling and autoscaling for two distinct workload profiles: latency-sensitive interactive streaming sessions and compute-intensive Gaussian Splatting jobs
Turn the Gaussian Splatting pipeline into a reproducible workflow, from data capture and training to OpenUSD assets
Operate and extend our Omniverse and Isaac Sim streaming platform, including Kit App Streaming, session lifecycle, and tenant isolation
Establish the platform as an internal standard through self-service workflows, golden paths, templates, onboarding, and documentation
Prepare the solution for operation by partners and customers through packaged deployments, versioned releases, upgrade paths, and actionable diagnostics
Support deployments across cloud, on-premises, and partner-managed environments
Build meaningful observability using metrics, logs, GPU telemetry, and actionable alerting
Define the simulation-specific security and access model, including ingress, TURN and STUN for WebRTC, RBAC, secrets management, SSO, and tenant boundaries
Work closely with our robotics, product, and platform teams to ensure that the solution supports real development and customer workflows
