← Back to jobs
KKAPTURE CX

ML Ops Engineer

KAPTURE CX

Bangalore
Full-Time
0.8 - 1.5 LPA
3–4 experience

Description

Role Name: MLOps EngineerAt Kapture CX, we are looking for an MLOps Engineer in our AI/ML Team.Kapture CX is a leading SaaS platform that helps enterprises automate and elevate customer experience through intelligent, AI-powered solutions. We partner with enterprises across industries to bring scalable automation and insight-driven efficiencies to their CX operations. Over a thousand clients across 18 countries have used Kapture's products to enhance their customer experience, including Unilever, Reliance, Coca-Cola, Bigbasket, Meesho, Airtel Payments Bank and Cathay Pacific. Kapture delivers industry-specific solutions powered by AI, tailored workflows, and seamless automation.Kapture CX is headquartered in Bangalore and we have offices in Mumbai and Delhi/NCR in India, in addition to offices in the USA, UAE, Singapore, Philippines and Indonesia.As an MLOps Engineer, you will own the infrastructure backbone of our conversational AI platform. You will design and manage high-performance model serving systems for LLM, ASR, and TTS workloads — ensuring reliability, scalability, and low-latency performance at production scale.This is a high-impact role where your architectural decisions will directly influence system speed, cost efficiency, and millions of real-time AI interactions.You will design, deploy, and maintain high-throughput, low-latency serving infrastructure for AI models across LLM, ASR, and TTS systems.You will evaluate and select inference engines such as vLLM, SGLang, LMDeploy, or TensorRT-LLM based on workload requirements and performance trade-offs.You will implement and optimise quantization strategies (INT8, INT4, FP8, GPTQ, AWQ, SmoothQuant) to maximise performance within compute constraints.You will architect distributed serving strategies including tensor, pipeline, and data parallelism.You will build containerised, reproducible deployment pipelines using Docker and Kubernetes.You will define and monitor critical performance metrics such as TTFT, latency percentiles (P50/P95/P99), throughput, and GPU utilisation.You will collaborate closely with ML Engineers during model handoffs to translate model requirements into production-ready infrastructure.You will build CI/CD pipelines to enable continuous and zero-downtime model deployments.This is a Bangalore-based role. We work five days a week from the office, as we believe in-person interactions fuel innovation and agility.

Eligibility Criteria

You have 3–4 years of hands-on experience in MLOps, AI infrastructure, or production model deployment.You have strong practical knowledge of inference engines such as vLLM, SGLang, LMDeploy, or similar frameworks.You have hands-on experience with GPU memory management, KV cache optimisation, and quantization techniques.You are proficient in Python and comfortable working with containerisation and cloud environments.You have experience with Docker, Kubernetes, and at least one major cloud provider (AWS, GCP, or Azure).

About KAPTURE CX

KAPTURE CX is an Indian CRM platform providing customer experience management solutions for SMEs, focusing on sales and service automation.

Industry: TechnologyEmployees: 1000+Website