# Naveen Kumar

Test Data Engineer | 5 years experience | Gurgaon, Haryana

Sample resume with fictional details. Replace the content with your own experience and qualifications.

Email: [Your email] | Phone: [Your phone]
LinkedIn: [Your LinkedIn URL] | GitHub / Portfolio: [Your portfolio URL]

## Professional Summary

- Skilled Test Data Engineer with 5 years of experience in managing test data for software testing and development
- Proficient in SQL, Python, and data anonymization techniques for creating compliant test datasets
- Experienced in synthetic data generation, data masking, and test data provisioning
- Knowledgeable in data pipelines, ETL processes, and database management for test environments
- Familiar with data governance, privacy regulations, and compliance standards
- Strong background in scripting, automation, and tooling for data management
- Collaborative engineer working with QA and development teams to ensure data availability
- Adept at managing version-controlled test data environments with automated refresh strategies supporting parallel testing across multiple environments

## Technical Skills

Test Data Management, Data Anonymization, SQL, Python, Synthetic Data Generation, Data Masking, ETL, Data Pipelines, PostgreSQL, MySQL, MongoDB, AWS, Azure, Snowflake, Pandas, NumPy, Faker, dbt, Airflow, Git, Docker, Bash, Excel, AI-Generated Synthetic Data, LLM Test Data, ChatGPT Data Synthesis, Conversational Dataset Generation, Privacy-Compliant AI Data, NLP Test Data Engineering

## Work Experience

### Test Data Engineer

[Company name] | [Start date - End date]

5 years of experience as a Test Data Engineer at software companies in Gurgaon. Developed test data management solutions that improved testing efficiency by 70% and ensured 100% data compliance across projects.

### Responsibilities

- Design and implement test data management strategies and frameworks
- Create and maintain anonymized test datasets compliant with privacy regulations
- Develop synthetic data generation tools and scripts for various test scenarios
- Automate data provisioning, snapshots, and refreshes for test environments
- Collaborate with development and QA teams to understand data requirements
- Implement data masking and anonymization techniques for sensitive information
- Manage database schemas, ETL processes, and data pipeline integrations
- Provide tooling and self-service portals for test data access and management
- Ensure data quality, consistency, and realism across test environments
- Monitor data usage, performance, and compliance in testing activities
- Document data management processes and provide training to teams
- Stay updated with data privacy laws and testing best practices

## Project Experience

### Multi-Tenant SaaS Test Data Orchestration

Designed and implemented test data orchestration system for a SaaS platform with 1000+ tenants. Created automated data provisioning, anonymization, and refresh mechanisms, supporting parallel testing environments. Technologies: SQL, Python, PostgreSQL, Airflow, AWS, Faker.

### E-commerce Platform Data Masking

Developed comprehensive data masking and anonymization framework for customer data in e-commerce testing. Implemented synthetic data generation for edge cases, ensuring privacy compliance and realistic test scenarios. Technologies: Python, Pandas, MongoDB, dbt, Azure, NumPy.

### Financial Services Test Data Management

Built test data management platform for banking applications, including data snapshots and environment cloning. Created self-service tools for QA teams, reducing data setup time by 80%. Technologies: SQL, MySQL, Snowflake, Bash, Docker, Git.

### AI-Generated Test Data for LLM Chatbot Testing

Built automated test data engineering pipeline for a conversational AI platform, using ChatGPT prompt-based synthesis to generate 10,000+ realistic user utterances, edge-case adversarial inputs, and multi-turn dialogue datasets. Implemented privacy-compliant PII replacement for regulated healthcare test environments and NLP evaluation datasets for LLM accuracy benchmarking. Technologies: AI-Generated Synthetic Data, LLM Test Data, ChatGPT Data Synthesis, Conversational Dataset Generation, Privacy-Compliant AI Data, Python, Faker, dbt, AWS.

## Education

B.Tech in Information Technology from Delhi Technological University, Delhi (2016-2020, 8.6 CGPA)

## Certifications

Certified Data Privacy Professional, AWS Certified Database - Specialty, SQL Server Certification

## Achievements

Awarded "Data Innovation Excellence" for SaaS data orchestration; Reduced test data setup time by 80%; Implemented GDPR-compliant data masking for 50+ projects
