- I build test automation and evaluation tools for conventional software and LLM applications.
- My professional experience spans Python automation, API and end-to-end testing, financial and e-commerce workflows, and LLM output evaluation.
- My public projects explore RAG evaluation, agent reliability, evaluation tooling, and AI-assisted QA with human review.
| Project | Engineering focus |
|---|---|
| Qaizen | AI-assisted QA connecting traceable test artifacts, Playwright and API checks, schema validation, and human review. |
| Evalstand | Python evaluation tooling with scorers, nested traces, run history, Pytest integration, and latency/token/cost tracking. In development. |
| EvalHarness | RAG and agent evaluation with DeepEval, a Claude judge, golden datasets, adversarial tests, and separate offline/live CI workflows. |
| PG Original POM | Python and Playwright automation with reusable page objects, business assertions, a local test environment, and read-only live smoke checks. |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|



