Readyly•1h ago
LinkedIn
Voice AI Engineer: Senior Software ...
India
Full Time
Senior Level
1200000-4000000
Full Job Description
About the Role
We're building voice-first AI systems that don't just talk, they listen, reason, plan, and act in real time. As a Senior Software Engineer focused on Voice AI, you'll architect and ship conversational voice agents powered by modern LLMs, design multi-agent orchestration, and build MCP (Model Context Protocol) integrations that let voice agents use real-world tools and data. If you love building low-latency, production-grade systems where every millisecond of response time matters, this role is for you.
What You'll Do
- Design and build real-time voice agents, including streaming speech-to-text (STT), LLM reasoning, and text-to-speech (TTS) pipelines with natural turn-taking, barge-in handling, and voice activity detection (VAD)
- Build agentic workflows (multi-step reasoning, tool-using agents, autonomous task execution) using LangGraph, AutoGen, CrewAI, or voice frameworks like Pipecat and LiveKit Agents
- Build and maintain MCP servers and clients that connect voice agents to internal tools, APIs, databases, CRMs, and external services
- Integrate telephony and real-time transport (WebRTC, SIP, Twilio, WebSockets) for inbound and outbound calling and in-browser voice experiences
- Implement RAG pipelines and vector search (pgvector, Pinecone, Weaviate) so voice responses are grounded in real data without adding latency
- Evaluate and integrate speech and language models, including Deepgram, Whisper, ElevenLabs, Cartesia, and Azure Speech, alongside LLMs from Anthropic Claude, OpenAI, and Gemini, plus open-source models like Llama and Mistral
- Develop and deploy scalable web and backend applications using React.js, Node.js, and Python, including serverless workloads on AWS Lambda
- Optimize end-to-end voice latency, cost, and call quality, and debug production issues such as dropped audio, interruptions, hallucinations, agent loops, and transcription errors
- Own code quality and mentor junior engineers through design reviews and pair programming
What You Bring
- 3+ years of hands-on experience with React.js, Node.js, and/or Python
- Experience building with LLM APIs and prompt engineering (tool calling, structured outputs), ideally tuned for spoken, conversational output
- Experience with real-time or streaming systems (WebSockets, WebRTC, or audio streaming)
- Solid experience deploying serverless workloads on AWS Lambda
- AWS Certification (Solutions Architect, Developer, or equivalent)
- Understanding of agentic AI concepts: planning loops, memory, tool use, and multi-agent coordination
- Bachelor's degree or equivalent in Computer Science or a related field
Bonus Points
- Direct experience building or consuming MCP servers
- Hands-on with Pipecat, LiveKit, Vapi, Retell, or similar voice-agent platforms
- Experience with STT/TTS providers, voice cloning, or multilingual and Indic-language voice (Hindi, Tamil, etc.)
- Familiarity with telephony (SIP, Twilio, Exotel) and call-center integrations
- Hands-on with LangChain, LangGraph, AutoGen, or CrewAI
- Familiarity with RAG architectures and vector databases
- Knowledge of AI observability tools (LangSmith, Arize, Helicone) and voice-quality evaluation (WER, latency, MOS)
- Experience with open-source model hosting (Ollama, vLLM, HuggingFace)
Company
Readyly
Readyly is a technology innovator dedicated to transforming how people work through advanced AI solutions.
India
Posted on LinkedIn