Supriya Nair
AI Quality Engineering Manager · 12 years experience
Bangalore, Karnataka
[Your email] · [Your phone]
[Your LinkedIn URL] · [Your GitHub / Portfolio URL]
Professional Summary
- Twelve-year quality engineering professional who founded India's first dedicated AI Quality Centre of Excellence within a Series-D SaaS company, managing 10 specialist AI testing engineers
- Manages a team exclusively focused on LLM evaluation, RAG pipeline testing, agentic workflow validation, hallucination detection, and AI safety for production AI systems
- Owns the DeepEval and RAGAS evaluation pipeline delivering real-time hallucination, toxicity, context recall, and coherence metrics integrated into CI/CD quality gates
- Designed the company's AI Red Teaming program aligned with OWASP LLM Top 10, covering 40+ prompt injection, jailbreak, and adversarial scenario test cases
- Established company AI quality SLAs: hallucination rate below 2%, context recall above 92%, toxicity score below 0.05, maintained across all production LLM features
- Created India's first OWASP LLM Testing curriculum in partnership with a national QA training institute, reaching 800+ testers in its first year
- Leads evaluation of agentic AI systems built on LangGraph and AutoGen for tool-call accuracy, behavioral correctness, and goal completion reliability
- Subject matter expert on responsible AI testing, EU AI Act compliance requirements, and AI ethics evaluation methodology for enterprise AI deployments
Work Experience
AI Quality Engineering Manager
[Company name] · [Start date – End date]
12 years in quality engineering progressing from manual QA to SDET to AI testing leadership. For the last 3 years, served as founding manager of the AI QE Centre of Excellence at a Series-D SaaS startup. Owned end-to-end AI testing strategy for products serving 500+ enterprise clients. Managed Rs. 55L annual budget for AI testing tooling and evaluation infrastructure. Reported to the Head of Engineering.
Responsibilities
- Lead and grow a 10-person AI QE CoE covering LLM evaluation, AI safety, RAG testing, and agentic systems testing
- Define AI quality SLAs including hallucination rate, toxicity score, context recall, and answer faithfulness targets
- Own the DeepEval and RAGAS evaluation pipeline integrated into CI/CD quality gates for all AI feature releases
- Design and execute AI red teaming exercises aligned with OWASP LLM Top 10 and AI safety best practices
- Advise product and engineering leadership on AI safety testing requirements, risk assessment, and regulatory readiness
- Drive EU AI Act compliance testing readiness for AI-powered products targeting EU-regulated markets
- Lead open-source AI testing contributions and publish AI quality research to build community thought leadership
- Mentor QA engineers transitioning into AI testing roles through structured 90-day AI testing upskilling programs
- Partner with ML team on model drift monitoring using OpenTelemetry and performance degradation alerting dashboards
- Define AI testing career paths, hiring criteria, and interview processes for specialist AI QE roles
Project Experience
Enterprise LLM Evaluation Platform
Built a self-service LLM evaluation platform using DeepEval and RAGAS enabling 12 product squads to run automated LLM quality checks in CI/CD pipelines with real-time hallucination dashboards. Technologies: DeepEval, RAGAS, Python, FastAPI, GitHub Actions, LangSmith, Grafana.
AI Red Teaming Program
Designed and executed the company's first AI red teaming exercise covering 40+ OWASP LLM Top 10 attack vectors including prompt injection, data poisoning, model inversion, and jailbreak scenarios. Technologies: Garak, PromptBench, Python, OWASP LLM Top 10 Framework.
RAG Quality Framework
Established RAG pipeline evaluation standards covering context recall, answer relevancy, faithfulness, and hallucination rate across 8 customer-facing RAG-powered features with automated quality regression. Technologies: RAGAS, LangSmith, Python, OpenAI, Pinecone, Weaviate.
Agentic Workflow Test Harness
Developed behavioral test harness for LangGraph-based multi-agent systems validating tool-call accuracy, loop detection, state transition correctness, and goal completion success rates. Technologies: LangGraph, Python, DeepEval, Playwright, Pytest.