Professional Summary
- Experienced AI/ML Validator with 4 years of expertise in validating model outputs and ensuring production reliability
- Proficient in implementing A/B testing and canary deployment procedures for ML models
- Skilled in monitoring model performance in production and setting up automated alerts for drift detection
- Knowledgeable in defining success metrics, thresholds, and KPIs for ML model evaluation
- Experienced in gradual rollout strategies and rollback mechanisms for safe model deployment
- Familiar with data science tools, statistical analysis, and model interpretability techniques
- Collaborative professional working with ML engineers and data scientists on model validation
- Strong analytical skills for identifying model biases and performance degradation
Work Experience
AI/ML Validator
[Company name] · [Start date – End date]
4 years of experience as an AI/ML Validator at a leading AI company in Bangalore. Specialized in model validation, A/B testing, and production monitoring for machine learning applications.
Responsibilities
- Implement A/B testing and canary deployment procedures for ML models
- Monitor model performance in production and set up alerts for drift
- Define success metrics, thresholds, and KPIs for model evaluation
- Collaborate with ML teams on model validation and deployment strategies
- Analyze model biases and ensure fairness in production
- Set up monitoring dashboards and automated reporting
- Provide recommendations for model improvements and optimizations
- Ensure compliance with ethical AI guidelines and standards
- Validate LLM-powered AI systems: hallucination detection, RAG pipeline accuracy (RAGAS), agentic workflow reliability, prompt injection security testing, and AI safety/alignment assessment
- Monitor production LLM systems with LangSmith tracing and implement continuous evaluation pipelines for deployed AI assistants
Project Experience
ML Ranking Model Canary Testing
Implemented canary testing for an ML ranking model with gradual rollout and monitoring, reducing production risks by 50% and ensuring stable performance. Technologies: Python, TensorFlow, Monitoring Tools, A/B Testing Tools.
Recommendation Engine Validation
Validated model outputs for a recommendation engine, setting up performance monitoring and alerts, achieving 95% accuracy in drift detection. Technologies: PyTorch, Pandas, Grafana, Statistical Analysis.
Image Classification Model A/B Testing
Designed A/B testing procedures for image classification models, analyzing results and providing recommendations for model improvements. Technologies: Scikit-learn, NumPy, Jupyter Notebook, Cloud Platforms.
RAG Pipeline and LLM Hallucination Validation
Designed and implemented validation framework for an enterprise RAG-based knowledge assistant, using RAGAS to measure faithfulness, answer relevancy, and context recall. Implemented DeepEval tests for hallucination detection and LangSmith-based tracing for production monitoring. Added adversarial prompt injection tests with PromptBench. Reduced hallucination rate by 45% pre-deployment. Technologies: RAGAS, DeepEval, LangSmith, PromptBench, RAG Pipeline Testing, Hallucination Detection, Prompt Injection Testing, Python, AWS.