Rahul Nair
ML Test Engineer · 4 years experience
Kochi, Kerala
[Your email] · [Your phone]
[Your LinkedIn URL] · [Your GitHub / Portfolio URL]
Technical Skills
Professional Summary
- Specialized ML Test Engineer with 4 years of experience testing machine learning models and AI systems
- Proficient in Python-based testing frameworks for ML model validation and performance evaluation
- Expert in implementing model monitoring, drift detection, and alerting systems for production ML models
- Skilled in data validation, dataset quality assessment, and ensuring data pipeline reliability
- Experienced in building reproducible testing pipelines for ML experiments and model deployments
- Knowledgeable in ML evaluation metrics, bias testing, and fairness assessments for AI systems
- Adept at collaborating with data scientists and ML engineers on model testing strategies
- Strong focus on CI/CD integration for ML pipelines with automated testing and validation
Work Experience
ML Test Engineer
[Company name] · [Start date – End date]
4 years in ML testing focused on ensuring the quality and reliability of machine learning systems. Expertise includes model validation, data quality testing, and production monitoring across recommendation engines, computer vision, and NLP systems.
Responsibilities
- Design and implement testing frameworks for ML models and data pipelines
- Build model monitoring systems with drift detection and performance alerting
- Create data validation rules and quality assessment frameworks
- Develop reproducible testing pipelines for ML experiments
- Collaborate with ML engineers on model evaluation and validation strategies
- Implement automated testing in ML CI/CD pipelines
- Perform bias testing, fairness assessments, and security validation for AI systems
- Monitor production ML systems and troubleshoot performance issues
- Document testing procedures and maintain testing infrastructure
- Stay updated with ML testing best practices and emerging tools
- Validate LLM and agentic AI systems: hallucination detection (DeepEval), RAG pipeline accuracy (RAGAS), adversarial prompt testing (PromptBench), OWASP LLM Top 10 compliance, and LangSmith production monitoring for enterprise AI quality assurance
Project Experience
Recommendation Engine Monitoring
Implemented comprehensive model monitoring system for e-commerce recommendation engine with drift detection and alerting. Technologies: Python, MLflow, Prometheus, AWS SageMaker. Reduced model performance degradation incidents by 70%.
Computer Vision Model Validation
Built automated testing pipeline for computer vision models including accuracy, robustness, and bias testing. Technologies: PyTorch, pytest, OpenCV, Docker. Improved model reliability by 50% and ensured consistent performance across edge cases.
NLP Model Data Pipeline Testing
Developed data validation and testing framework for NLP model training pipelines. Technologies: TensorFlow, Great Expectations, Apache Spark, Kubernetes. Enhanced data quality by 60% and reduced training failures through early validation.
LLM and RAG Production Quality Framework
Extended ML testing expertise into LLM systems by building a comprehensive evaluation pipeline for a GPT-4-powered customer service LLM. Used RAGAS for RAG retrieval quality, DeepEval for hallucination and faithfulness scoring, LangSmith for production trace analysis, and PromptBench for adversarial testing. Established OWASP LLM Top 10 compliance checks in CI. Reduced critical hallucinations by 60% before production release. Technologies: RAGAS, DeepEval, LangSmith, PromptBench, OWASP LLM Top 10, RAG Pipeline Testing, Hallucination Detection, Agentic Workflow Testing, Python, AWS SageMaker.