GitHub Work

Frameworks, experiments and tools exploring AI product development, evaluation and product decision-making.

Visit my GitHub profile ↗

AI evaluation framework

A practical release-gate system for deciding whether an AI product is ready to ship. It converts quality, safety, grounding, latency, cost, user outcome and operational reliability into explicit pass-or-fail evidence, with reference scenarios and an offline demonstration.

View ai-eval-framework on GitHub ↗

AI user journey

A measurement model for understanding the complete AI experience, from an unresolved user problem to successful adoption and advocacy. Nine stages, four analytical lanes and 27 metrics connect user progress with model behaviour, product friction and business outcomes.

View ai-user-journey on GitHub ↗

One layer deeper

A structured reasoning skill that challenges the first explanation of a product problem. It separates symptoms from underlying causes, tests assumptions, identifies second-order effects and produces sharper questions before a team commits to a solution or roadmap decision.

View one-layer-deeper on GitHub ↗

Evolutionary crossover memory

An experimental memory architecture for retaining useful knowledge across repeated AI interactions. It explores how strong elements from previous solutions can be selected, combined and refined so later responses improve without simply copying earlier outputs.

View ECM on GitHub ↗

PM AI Copilot

An operating system for product-management work built around five specialised agents. Persistent context, Jira access through MCP and proactive workflow hooks support research synthesis, requirements, prioritisation and delivery while keeping product judgement with the human owner.

View pm-ai-copilot on GitHub ↗