← Back to jobs
CConvegenius

Sr. ML Engineer- Generative AI

Convegenius

Noida
Full-Time
50.0 - 70.0 LPA
2-4 years experience

Description

What you will do

  • Design, build, and ship production-grade LLM applications: RAG pipelines, agentic workflows, and tool-calling systems for education-specific use cases such as question answering, content generation, adaptive feedback, and curriculum alignment.
  • Own full system architecture: data pipelines, retrieval layers, orchestration, APIs, and service infrastructure — written as maintainable, tested production code, not notebooks or scripts.
  • Build rigorous LLM evaluation frameworks: task-specific benchmarks, regression testing, human eval loops, and automated quality gates to catch degradation before it ships.
  • Own latency and cost at the application layer: caching strategies, request batching, model/route selection, and prompt-context optimisation to keep production systems fast and affordable at scale.
  • Build data pipelines for grounding and instruction data, including multilingual and Indic language sources.
  • Assess new open-source and closed model releases for domain applicability, cost, and production readiness.
  • Define and track system performance metrics, evaluation benchmarks, and reliability targets across all shipped features.
  • Work with open-source and sovereign LLMs, integrating them into production-proven serving and orchestration frameworks.
  • Where warranted, fine-tune or adapt foundation models (LoRA/QLoRA/SFT) — a strong plus, not a gating requirement for this role.

Key Responsibilities

  • Design, build, and ship production-grade LLM applications: RAG pipelines, agentic workflows, and tool-calling systems for education-specific use cases such as question answering, content generation, adaptive feedback, and curriculum alignment.
  • Own full system architecture: data pipelines, retrieval layers, orchestration, APIs, and service infrastructure — written as maintainable, tested production code, not notebooks or scripts.
  • Build rigorous LLM evaluation frameworks: task-specific benchmarks, regression testing, human eval loops, and automated quality gates to catch degradation before it ships.
  • Own latency and cost at the application layer: caching strategies, request batching, model/route selection, and prompt-context optimisation to keep production systems fast and affordable at scale.
  • Build data pipelines for grounding and instruction data, including multilingual and Indic language sources.
  • Assess new open-source and closed model releases for domain applicability, cost, and production readiness.
  • Define and track system performance metrics, evaluation benchmarks, and reliability targets across all shipped features.
  • Work with open-source and sovereign LLMs, integrating them into production-proven serving and orchestration frameworks.
  • Where warranted, fine-tune or adapt foundation models (LoRA/QLoRA/SFT) — a strong plus, not a gating requirement for this role.

Must-Have Skills

  • Strong production Python engineering: clean, tested, maintainable code — not scripts or notebooks. Comfortable owning services in production.
  • Hands-on experience building and shipping RAG or agentic systems in production: retrieval design, orchestration (e.g. LangChain/LangGraph or equivalent), tool calling, and multi-step reasoning pipelines.
  • Demonstrated experience building LLM evaluation systems: benchmarks, regression tests, human-eval workflows, and quality monitoring — not just anecdotal "it works."
  • Experience building and operating backend systems/services: APIs, data pipelines, deployment, and monitoring in a real production environment.
  • Working knowledge of AI/ML pipelines, data preparation, and model deployment.
  • Experience working with open-source and sovereign LLMs, using standard, production-proven frameworks.
  • Understanding of cloud platforms and infrastructure for serving ML/LLM systems at scale.

Eligibility Criteria

  • Strong production Python engineering: clean, tested, maintainable code — not scripts or notebooks. Comfortable owning services in production.
  • Hands-on experience building and shipping RAG or agentic systems in production: retrieval design, orchestration (e.g. LangChain/LangGraph or equivalent), tool calling, and multi-step reasoning pipelines.
  • Demonstrated experience building LLM evaluation systems: benchmarks, regression tests, human-eval workflows, and quality monitoring — not just anecdotal "it works."
  • Experience building and operating backend systems/services: APIs, data pipelines, deployment, and monitoring in a real production environment.
  • Working knowledge of AI/ML pipelines, data preparation, and model deployment.
  • Experience working with open-source and sovereign LLMs, using standard, production-proven frameworks.
  • Understanding of cloud platforms and infrastructure for serving ML/LLM systems at scale.
  • Fine-tuning experience is a plus for this role, not a requirement.

About Convegenius

Convegenius is an Indian edtech company using AI for personalized learning and assessments, helping schools and students improve educational outcomes.

Industry: TechnologyEmployees: 200+Website