NVIDIA•2h ago
LinkedIn
Senior HPC Cluster Engineer
Bengaluru, Karnataka, India
Senior Level
Full Job Description
About the Role
NVIDIA is seeking a Senior HPC Cluster Engineer for its MARS team in Bengaluru. You will lead strategic guidance and operations on large-scale High-Performance Computing (HPC) systems, managing compute, networking, and storage infrastructure essential for AI research.
Key Responsibilities:
- Provide leadership in system administration, service delivery, incident response, and reliability improvements for our global GPU fleet.
- Collaborate with global teams to deliver world-class user experiences for AI/HPC researchers.
- Own day-to-day operations of production clusters, ensuring health, efficiency, and resource optimization.
- Develop scalable automation solutions and maintain heterogeneous clusters on-premises and in the cloud using tools like Terraform, Ansible, and Docker/Singularity.
- Optimize cluster performance by analyzing job fragmentation, GPU waste, and tuning workloads for SLA targets.
- Conduct root cause analysis, lead SEV triage postmortems, and participate in on-call rotations for critical production environments.
Company
NVIDIA
NVIDIA, founded in 1993 (NASDAQ: NVDA), is a global pioneer in accelerated computing and AI infrastructure.Our invention of the GPU revolutionized PC gaming, redefined computer graphics, ignited moder...
Bengaluru, Karnataka, India
Posted on LinkedIn