NVIDIA Logo

NVIDIA

Manager, System Software Engineering - Local AI

Posted 21 Days Ago
Be an Early Applicant
In-Office
Pune, Maharashtra, IND
Senior level
In-Office
Pune, Maharashtra, IND
Senior level
Lead and grow a team building an on-device AI inference platform for RTX/ DGX GPUs. Drive cross-functional alignment, architect inference runtimes, optimize models and pipelines for low-latency, memory-efficient local deployment, and establish processes for system-level debugging, performance/accuracy trade-offs, and production readiness. Mentor engineers and deliver roadmap-aligned, high-quality releases.
The summary above was generated by AI

NVIDIA has continuously reinvented itself for over two decades. The invention of the GPU in 1999 propelled the growth of PC gaming, redefined modern computer graphics, and revolutionized parallel computing. More recently, GPU deep learning has ignited modern AI, positioning NVIDIA as a leading AI computing company. There is a growing focus on delivering AI models locally, closer to the source of data. This reduces latency, improves real-time processing, and addresses privacy concerns by minimizing data transfer to centralized servers. As technology advances, client-side AI (local execution) will play a key role in crafting digital experiences.  

The Local AI team is seeking a System Software Manager to lead development of an efficient on-device AI software stack. The software stack will support RTX, RTX Pro, and DGX-class systems. This role focuses on high-performance local inference, agentic workloads, low latency, efficient memory use, scalable infrastructure, practical deployment on resource-constrained platforms, and delivering a streamlined out-of-box experience for developers and end users.  

What you'll be doing:

  • Lead and grow a team building the on-device AI inference platform for RTX, RTX Pro, and DGX GPUs, with accountability for execution, technical direction, delivery quality, and roadmap alignment.  

  • Drive cross-functional alignment with NVIDIA’s software, research, architecture, and product teams, along with industry partners and open-source communities, to build strategy and strengthen the AI ecosystem across RTX and DGX platforms. 

  • Provide technical leadership for the architecture and evolution of modern inference runtimes and execution stacks across frameworks such as Llama.cpp, vLLM, PyTorch, WinML, DXCGC, and TensorRT-RTX, spanning workloads including LLMs, vision-language models, TTS, ASR, and diffusion models.  

  • Mentor engineers, develop technical leaders, and foster a high-performance team culture centred on innovation, collaboration, and operational excellence. 

  • Coordinate end-to-end optimization of AI models, data pipelines, and inference runtimes to improve performance across current and next-generation GPU architectures. 

  • Drive adoption of model optimization techniques such as quantization, pruning, sparsity, and distillation to enable efficient deployment of large models on local and edge devices. 

  • Establish team processes for system-level debugging, performance optimization, and performance-accuracy trade-off analysis, including infrastructure for performance and accuracy sweeps, gap analysis, and production-readiness improvements. 

What we need to see:

  • 5+ overall years of industry experience and 2+ years of engineering leadership experience, combined with a Bachelor’s, Master’s, or PhD in Computer Science, Software Engineering, Mathematics, or a related field. 

  • Proven experience leading high-performing engineering teams in systems software, AI infrastructure, inference runtimes, or related domains. 

  • Strong technical foundation in C++ software development, debugging, data structures, algorithms, and machine learning systems. 

  • Extensive background in AI inference pipelines and Deep Learning frameworks like Llama.cpp, vLLM, PyTorch, WinML, DXCGC, and TensorRT. 

  • Deep understanding of inference backends and runtime internals, including scheduling, memory management, KV-cache behavior, graph execution, quantization, and hardware-aware optimization. 

  • Strong analytical and problem-solving skills, with the ability to balance technical depth, execution speed, and organizational priorities in a fast-paced environment. 

  • Excellent written and verbal communication skills, with proven ability to collaborate across engineering, product, research, and executive collaborators. 

Ways to stand out from the crowd :

  • Strong understanding of modern machine learning, deep neural networks, and generative AI, along with contributions to notable open-source projects. 

  • Demonstrated success building teams, setting technical vision, and scaling execution through periods of rapid growth. 

  • Track record of delivering end-to-end products with geographically distributed teams in multinational product organizations. 

  • Experience in lower-level systems or GPU programming, including CUDA and high-performance systems development. 

  • Contributions to open-source inference runtimes, model tooling, or performance infrastructure as well as practical experience working with frameworks and APIs including Llama.cpp, PyTorch, TensorRT, Vulkan, DirectX, and vLLM. 

We're a top employer known for innovation and growth. We are an equal-opportunity employer and value diversity at our company. With competitive salaries and a generous benefits package, we are widely considered to be one of the world’s most desirable employers of technology. We have some of the most forward-thinking and hardworking people in the world working for us and, due to unprecedented growth, our best-in-class engineering teams are rapidly growing. If you're a creative and autonomous engineer with a real passion for technology, we would like to hear from you. 

NVIDIA Pune, Mahārāshtra, IND Office

Survey No.144 145, Commerzone No.5, Off, Airport Rd, Yerawada, Pune, Maharashtra, India, 411006

Similar Jobs

53 Minutes Ago
Hybrid
Senior level
Senior level
Financial Services
Lead design, develop, and maintain secure, high-quality production Java full-stack applications (React front-end). Drive adoption of AI-assisted engineering practices, enforce code quality and SDLC, improve operational stability through automation, evaluate vendor/architecture decisions, and lead engineering communities of practice.
Top Skills: Ai-Assisted Development ToolsAWSCheckstyleCi/CdEclipseHibernateIbatisJ2EeJavaMule EsbNoSQLOracle 10GOracle 9IPmdReactSonarSpringSQLTomcatWeblogic
8 Hours Ago
Hybrid
Pune, Maharashtra, IND
Senior level
Senior level
Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Lead ML engineering for Operational Intelligence: design and productionize multimodal transformer and agentic systems, build scalable RAG/Graph-RAG and LLMOps/MLOps pipelines, implement inference services and memory/state subsystems, apply traditional and deep learning methods, enforce Responsible AI, and drive research-to-production innovation.
Top Skills: AutogenAWSAws NeptuneAzureBertClipCrewaiDatabricksFeature StoreGCPGraph-RagHugging FaceKnowledge GraphLanggraphLlavaMlflowNeo4JPower BIPythonPyTorchRagSQLT5TableauTensorFlowVector StoreWhisper
8 Hours Ago
Hybrid
Pune, Maharashtra, IND
Senior level
Senior level
Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Lead a team to architect and deliver production-grade generative AI: multi-agent systems, multimodal transformers, RAG/Graph-RAG, LLM fine-tuning and LLMOps on cloud platforms. Build Python inference services, manage agent state/memory, implement responsible AI and governance, and integrate traditional ML for interpretable hybrid systems while researching and productionizing frontier models.
Top Skills: AnthropicAutogenAws BedrockAws NeptuneAws SagemakerAzure OpenaiBertClipCrewaiDatabricksFalconFastapiFeature StoresGcp Vertex AiGeminiGpt-4OHugging FaceLangchainLanggraphLlamaLlavaLoraMistralMlflowNeo4JOpenaiOpensearchPeftPgvectorPineconePythonPyTorchRlhfSQLT5TensorFlowWeights & BiasesWhisper

What you need to know about the Pune Tech Scene

Once a far-out concept, AI is now a tangible force reshaping industries and economies worldwide. While its adoption will automate some roles, AI has created more jobs than it has displaced, with an expected 97 million new roles to be created in the coming years. This is especially true in cities like Pune, which is emerging as a hub for companies eager to leverage this technology to develop solutions that simplify and improve lives in sectors such as education, healthcare, finance, e-commerce and more.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account