NVIDIA
Teams at NVIDIA
Recently posted jobs
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
Lead major incident response and set technical direction for site reliability engineering across NVIDIA’s enterprise platforms. Design and operate distributed, Kubernetes-based cloud infrastructure; build automation, observability, self-healing systems, and AI-assisted incident tooling. Drive root cause analysis, SLOs, error budgets, and systemic reliability improvements. Partner with Cloud, Platform, Security, and AI/ML teams, mentor engineers, influence architecture, and communicate with executives during critical incidents.
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
Support enterprise customers deploying NVIDIA AI Enterprise across cloud and datacenter environments. Troubleshoot complex software issues, reproduce failures, collect diagnostics, and partner with engineering on fixes. Build Python automation, diagnostics, test harnesses, Kubernetes deployment assets, patches, documentation, and runbooks. Work with Linux, containers, GPUs, AI frameworks, inference services, distributed systems, and cloud platforms. Own customer escalations through resolution and participate in a monthly weekend Sev1 on-call rotation.
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
Lead NVIDIA’s graphics systems team in developing and productionizing an agentic software engineering framework. Responsibilities include defining strategy and transformation roadmaps, applying agents across requirements, implementation, validation, review, and release, maintaining graphics and display product commitments, growing engineering talent, and packaging the resulting orchestration, tools, evaluations, and operating practices for adoption across NVIDIA.
