EnterOne•3h ago
LinkedIn
AI Systems & Deployment Engineer
India
Remote
Senior Level
Full Job Description
About the AI Systems & Deployment Engineer Role (Full-Time, Engineering, AI/ML Infrastructure) Role Overview: We seek an AI Systems & Deployment Engineer to lead the full lifecycle of large language model deployments—from model selection and fine-tuning to production scaling and continuous optimization. You will bridge cutting-edge ML research with enterprise infrastructure, building reliable, low-latency AI services that power real-world products. Key Responsibilities:
- Deploy and scale open-source and proprietary LLMs using modern inference engines (e.g., vLLM, TGI, TensorRT-LLM, ONNX Runtime).
- Establish and maintain CI/CD pipelines for ML models in production.
- Fine-tune and optimize models for low-latency inference, including quantization, distillation, and throughput benchmarking.
- Design advanced Retrieval-Augmented Generation (RAG) pipelines and integrate enterprise vector databases (e.g., Pinecone, Weaviate, pgvector).
- Build multi-agent systems orchestrating tool use, memory, and inter-agent communication for autonomous workflows.
- Collaborate with Infrastructure Engineers to optimize GPU clusters, reduce costs, and meet uptime SLAs.
- Leverage NVIDIA’s software stack (CUDA, TensorRT, Triton Inference Server) for hardware efficiency.
- 3+ years in AI/ML software engineering, with 1+ year focused on production LLM deployments, fine-tuning, and RAG architectures.
- Hands-on experience with inference frameworks and containerized ML workloads on Kubernetes.
- Strong Python proficiency and MLOps tooling (MLflow, Weights & Biases, DVC).
- Deep understanding of transformer architectures and fine-tuning techniques (LoRA, QLoRA, RLHF).
- Ability to benchmark models for latency, throughput, and quality.
- Experience with multi-agent frameworks (LangGraph, CrewAI, AutoGen).
- Familiarity with NVIDIA Triton Inference Server and GPU profiling tools.
- Background in distributed systems or high-performance computing.
- Prior work with enterprise vector stores and hybrid search architectures.
Company
EnterOne
Discover EnterOne: A leading global IT training and consulting partner specializing in Cisco, Splunk, VMware, AWS, Microsoft, and F5 technologies. As a Cisco Platinum Learning Partner and 2023 Learnin...
India
Posted on LinkedIn