Applied AI | Python | AI Evaluation | Agentic Systems
I build applied AI systems, Python tooling and enterprise agent workflows, using AI-assisted development with hands-on testing, debugging, validation and human review. My work sits at the intersection of AI evaluation, reproducibility, agentic workflows and enterprise adoption.
View my detailed technical portfolio
Python · AI evaluation · LLM testing · AI agents · Azure AI / Azure OpenAI · prompt engineering · grounding / RAG concepts · GitHub Actions · pytest · mypy · Ruff · responsible AI
Creator and maintainer of an open-source Python toolkit for detecting semantic drift in AI evaluation inputs and task contracts using hash-only reproducibility manifests.
- Published Python package and CLI supporting Python 3.11–3.13
- Generic JSONL, Inspect AI and Harvey LAB adapters
- Deterministic scope, coverage, ordering and semantic-field drift classification
- Text, JSON and Markdown reports with CI-friendly exit codes
- Strict typing with mypy, Ruff quality gates and enforced test coverage above 90%
- Reusable GitHub Action for evaluation reproducibility checks
- Fork-side Inspect Evals case study covering 22,773 records across seven complete datasets
- Harvey LAB pinned-revision case study identifying semantic drift across 250 task contracts while preserving hash-only evidence
Latest public alpha: v0.1.0a2
Selected contributions currently in upstream review:
- UK Government Inspect Evals — draft PR #2113 — Python tooling and regression coverage for dataset-dependency reproducibility across AI evaluations, including isolated-version comparison, semantic manifests and automated validation.
- Promptfoo — open PR #10336 — safer direct navigation to specific evaluation IDs, with extensive regression coverage across CLI, server and Web UI paths.
These contributions are intentionally described as open or draft work until upstream maintainers merge them.
At MSCI I designed and built an Azure-based account-intelligence agent using Azure AI / Azure OpenAI patterns and combining Salesforce, Power BI and public-company information. The workflow was designed to identify product gaps, reporting gaps, account signals and cross-sell or upsell opportunities across a strategic portfolio of approximately $290M.
The implementation work included prompt design, structured outputs, source grounding, evidence traceability, data-quality checks, evaluation scenarios, human-review gates and iterative prototyping. My Python work is applied-AI focused, especially evaluation, data handling, automation and prototype development.
I use AI-assisted development as an engineering workflow rather than a substitute for understanding the code. I define the desired behaviour, inspect and edit generated code, run tests and static checks, debug failures, review diffs, validate edge cases and use CI evidence before treating a change as complete.
My strongest hands-on areas are:
- Python for AI evaluation, data transformation, CLI tooling, testing and automation
- LLM and agent evaluation, reproducibility and semantic-drift analysis
- Azure AI / Azure OpenAI agent prototyping and prompt engineering
- Responsible AI controls including grounding, source traceability and human-in-the-loop review
- Git, GitHub Actions, pytest, mypy, Ruff and reproducible development workflows
I combine open-source engineering with enterprise AI strategy and research on the economic and societal effects of digitalisation.
- Springer publication — Digitalization and Its Tax Implications: Evidence from the UK and Hungary
- EU-funded ODDEA project results — research contributions developed through international secondments
Python for applied AI · AI evaluation · LLM testing · reproducibility · benchmark contracts · agentic workflows · responsible deployment · enterprise AI strategy
I am interested in roles where commercial and domain judgement are strengthened by practical AI engineering, especially applied AI, AI product, AI solutions, AI strategy and customer-facing technical roles.


