Senior System Engineer
Full Job Description
About the Role
We are building the first Global Integrated Operations Command Center (GIOC) in Chennai, India. This role serves as the frontline for Vultr's systems estate across compute, storage, operating systems, and virtualization.
The Senior Systems Engineer acts as the entry point of our incident lifecycle: monitoring health signals from a consolidated dashboard, acknowledging alerts within strict SLAs, performing initial triage, executing runbooks to resolve common issues, and escalating complex matters. You will work in a fast-paced, alert-driven environment supporting 24x7 global coverage.
Key Responsibilities
- Alert Monitoring & First Response: Monitor systems-estate health signals across compute, storage, OS, and virtualization platforms using our single-pane-of-glass dashboard. Acknowledge and triage alerts within defined SLA targets by reviewing recent changes and checking the CMDB.
- Triage & Severity Classification: Classify incidents based on severity and customer impact, confirming ownership (Systems/Network/Security). Route tickets accurately to correct towers or escalation paths while maintaining high accuracy standards. Reassess severity as impacts evolve.
- Runbook-Driven Resolution: Match alert signatures to documented runbooks and execute permitted Level 2 actions to resolve common issues end-to-end. Document every step with supporting evidence in the incident ticket.
- Escalation & Handoff: Escalate unresolved or out-of-scope issues to L2/L3 teams with a clean handoff, including symptom evidence and current state. Flag missing runbooks for improvement. Support major-incident (Sev-0/1) bridges.
- Operational Documentation & Communication: Maintain accurate shift logs, ticket updates, and incident timelines. Provide clear status communication to stakeholders and customers per severity requirements. Produce thorough handover notes.
- Continuous Improvement: Identify recurring alerts for tuning/automation, contribute to runbook improvements, participate in post-incident reviews (PIR), and trend analysis.
Qualifications
A graduate or Engineer with a B.E./B.Tech. We are looking for 5-8 years of experience in systems administration, NOC/SOC, or IT operations with hands-on Linux/Windows server management. You must have knowledge of OS processes/services/log analysis, virtualization concepts (VMs/hypervisors), and incident-management fundamentals.
Preferred Qualifications
- ITIL V4 Foundation certification or equivalent ITSM knowledge
- Linux certifications (RHCSA/LFCS) or cloud-fundamentals credentials
- Familiarity with scripting languages like Bash, Python, or PowerShell for automation
- Experience in 24x7 NOC/SOC environments using JIRA, Confluence, PagerDuty, and observability tools.
About the Shift Model
This is a rotational shift role supporting follow-the-sun coverage. You must be comfortable working nights, weekends, and holidays to ensure 24x7 global operations continuity.
Company
Vultr
Vultr is a leading global cloud infrastructure provider, trusted by hundreds of thousands of customers across 185 countries for its scalable Cloud Compute, GPU, Bare Metal, and Storage solutions.Why V...