I
Software Test Engineer – AI POC (LLM / RAG & Agentic AI)
Ideas to Impacts
Pune
Full-Time
0-3 Years experience
Description
This role is part of a time-bound AI Proof of Concept (POC) focused on LLM / RAG output evaluation and Agent-based (Agentic) AI workflows.
The selected candidate is expected to contribute independently and be productive from Week 2.
There is no training or grooming bandwidth available for this role.
Key Responsibilities:
- Test and validate LLM-based and RAG-driven AI systems
- Perform AI output evaluation for relevance, accuracy, consistency, and hallucinations
- Apply standard NLP / LLM evaluation metrics during testing
- Test Agent-based (Agentic) AI workflows, including agent-to-agent interactions
- Write and execute Python-based test scripts for AI validation
- Perform functional, regression, and performance testing for AI POCs
- Identify, document, and clearly communicate AI model issues
- Collaborate closely with AI engineers to understand model behavior
- Provide concise, outcome-focused test reports for decision-making
Requirements
Mandatory (Non-Negotiable):
-
2 to 2.5 years of experience (maximum 3 years) in AI / LLM / NLP testing
-
Hands-on experience testing LLM or RAG-based systems
-
Practical experience with NLP / LLM evaluation metrics, including:
- BLEU
- ROUGE
- METEOR
- BERTScore
- COMET
-
Strong Python programming skills with independent scripting ability
-
Understanding of:
- RAG pipelines
- Agentic / Agent-based AI workflows
- Model context handling and response validation
-
Ability to work independently in a POC / experimental environment
Technical (Supporting):
- Working knowledge of Java (basic to intermediate)
- Familiarity with AI / ML frameworks such as TensorFlow or PyTorch
- Exposure to automation tools (Selenium, BDD Cucumber – secondary)
- Cloud exposure (AWS / Azure / GCP) – good to have
Important Exclusions:
- This is not a traditional manual or automation QA role
- Not suitable for candidates learning AI fundamentals
- Not a buffer or training position
Benefits
- Opportunity to work on cutting-edge AI / LLM POCs
- Hands-on exposure to LLM evaluation, RAG pipelines, and Agentic AI
- High ownership and visibility in a delivery-focused environment
- Potential continuation beyond POC based on performance and outcomes
About Ideas to Impacts
-
