← Back to jobs
IIdeas to Impacts

Software Test Engineer – AI POC (LLM / RAG & Agentic AI)

Ideas to Impacts

Pune
Full-Time
0-3 Years experience

Description

This role is part of a time-bound AI Proof of Concept (POC) focused on LLM / RAG output evaluation and Agent-based (Agentic) AI workflows.

The selected candidate is expected to contribute independently and be productive from Week 2.

There is no training or grooming bandwidth available for this role.

Key Responsibilities:

  • Test and validate LLM-based and RAG-driven AI systems
  • Perform AI output evaluation for relevance, accuracy, consistency, and hallucinations
  • Apply standard NLP / LLM evaluation metrics during testing
  • Test Agent-based (Agentic) AI workflows, including agent-to-agent interactions
  • Write and execute Python-based test scripts for AI validation
  • Perform functional, regression, and performance testing for AI POCs
  • Identify, document, and clearly communicate AI model issues
  • Collaborate closely with AI engineers to understand model behavior
  • Provide concise, outcome-focused test reports for decision-making

Requirements

Mandatory (Non-Negotiable):

  • 2 to 2.5 years of experience (maximum 3 years) in AI / LLM / NLP testing

  • Hands-on experience testing LLM or RAG-based systems

  • Practical experience with NLP / LLM evaluation metrics, including:

    • BLEU
    • ROUGE
    • METEOR
    • BERTScore
    • COMET
  • Strong Python programming skills with independent scripting ability

  • Understanding of:

    • RAG pipelines
    • Agentic / Agent-based AI workflows
    • Model context handling and response validation
  • Ability to work independently in a POC / experimental environment

Technical (Supporting):

  • Working knowledge of Java (basic to intermediate)
  • Familiarity with AI / ML frameworks such as TensorFlow or PyTorch
  • Exposure to automation tools (Selenium, BDD Cucumber – secondary)
  • Cloud exposure (AWS / Azure / GCP) – good to have

Important Exclusions:

  • This is not a traditional manual or automation QA role
  • Not suitable for candidates learning AI fundamentals
  • Not a buffer or training position

Benefits

  • Opportunity to work on cutting-edge AI / LLM POCs
  • Hands-on exposure to LLM evaluation, RAG pipelines, and Agentic AI
  • High ownership and visibility in a delivery-focused environment
  • Potential continuation beyond POC based on performance and outcomes

About Ideas to Impacts

-

Industry: no-mentionEmployees: 74+Website