EnterOne
EnterOne3h ago
LinkedIn

AI Systems & Deployment Engineer

India
Remote
Senior Level

Auto Apply to 50+ AI Matched AI Systems & Deployment Engineer Jobs

Use Auto Apply Agents to Bulk Apply jobs with ATS Optimised Resumes, find verified Insider Connections for jobs at EnterOne

Full Job Description

About the AI Systems & Deployment Engineer Role (Full-Time, Engineering, AI/ML Infrastructure) Role Overview: We seek an AI Systems & Deployment Engineer to lead the full lifecycle of large language model deployments—from model selection and fine-tuning to production scaling and continuous optimization. You will bridge cutting-edge ML research with enterprise infrastructure, building reliable, low-latency AI services that power real-world products. Key Responsibilities:
  • Deploy and scale open-source and proprietary LLMs using modern inference engines (e.g., vLLM, TGI, TensorRT-LLM, ONNX Runtime).
  • Establish and maintain CI/CD pipelines for ML models in production.
  • Fine-tune and optimize models for low-latency inference, including quantization, distillation, and throughput benchmarking.
  • Design advanced Retrieval-Augmented Generation (RAG) pipelines and integrate enterprise vector databases (e.g., Pinecone, Weaviate, pgvector).
  • Build multi-agent systems orchestrating tool use, memory, and inter-agent communication for autonomous workflows.
  • Collaborate with Infrastructure Engineers to optimize GPU clusters, reduce costs, and meet uptime SLAs.
  • Leverage NVIDIA’s software stack (CUDA, TensorRT, Triton Inference Server) for hardware efficiency.
Required Qualifications:
  • 3+ years in AI/ML software engineering, with 1+ year focused on production LLM deployments, fine-tuning, and RAG architectures.
  • Hands-on experience with inference frameworks and containerized ML workloads on Kubernetes.
  • Strong Python proficiency and MLOps tooling (MLflow, Weights & Biases, DVC).
  • Deep understanding of transformer architectures and fine-tuning techniques (LoRA, QLoRA, RLHF).
  • Ability to benchmark models for latency, throughput, and quality.
Preferred Qualifications:
  • Experience with multi-agent frameworks (LangGraph, CrewAI, AutoGen).
  • Familiarity with NVIDIA Triton Inference Server and GPU profiling tools.
  • Background in distributed systems or high-performance computing.
  • Prior work with enterprise vector stores and hybrid search architectures.
What We Offer: Competitive salary and equity, access to cutting-edge GPU infrastructure, and a collaborative team valuing engineering excellence in a flexible remote-friendly environment.

Company

EnterOne

EnterOne

Discover EnterOne: A leading global IT training and consulting partner specializing in Cisco, Splunk, VMware, AWS, Microsoft, and F5 technologies. As a Cisco Platinum Learning Partner and 2023 Learnin...

India
Posted on LinkedIn