Beyond POC. In Production.
Real systems delivering real value

AI Evaluation Machine — Correctness, Not Just Plausibility
- •Conceived and built end-to-end — no separate architect, no development team
- •Evaluation framework for a production RAG assistant (Top-5 global CPG client)
- •Core thesis: similarity ≠ correctness
- •Red-team & resistance batteries, retrieval-defect diagnosis, scoring rubrics
- •It outlived the engagement it was built for — now the way answer quality gets proven on other accounts
- •Now doubles as a pre-sales asset: how we prove AI quality, shown to the next client
- •Delivered inside a top-tier global technology services company
BMAD #2767 — The Bug That Deleted Requirements Without a Trace
- •Found on a live enterprise delivery: the framework’s planner compiled epic context that dropped every reference to the 28 KB design contract — behaviour kept, filenames and tokens gone
- •Triage could not reproduce it — a precise fact package (version, model, artifacts with zero design refs) turned “works for me” into a maintainer-confirmed bug
- •Closed as tracked, not unresolved: proper backlinking became a v7 requirement, with the diagnosis recorded verbatim in the closing comment and both reporters credited
- •Independently confirmed and widened by a second team on a newer version — the thread spawned a sibling issue (#2796)
- •Proposed the v7 detector: a deterministic zero-hit grep with no model in it
- •Shipped our own stopgap meanwhile — measured, not assumed: it came with its own positive control
AI HR Assistant — Governed Self-Service on Live HR Data
- •Started where these always start: discovery workshops with a client who had not scoped the problem yet
- •Eleven HR services an employee completes in conversation instead of a ticket
- •Leave balances, leave requests, documents, request status, policy answers — on live HR data
- •A leave request that spans ten system calls across five steps becomes one exchange
- •Writes are confirmed before submission; identity comes from the session, never from the client
- •API collection read operation by operation and mapped to the capabilities it can actually deliver
- •Answers measured for correctness against the labour law before anyone is allowed to rely on them
Second Brain — Persistent Memory for an AI Collaborator
- •Two-layer memory: an explicit cloud corpus plus an implicit persona layer
- •2,346 memories extracted, typed and scoped out of real work — no curated dataset, no hand-maintained prompt
- •What repeats stops being a memory and becomes a principle: 68 active, reusable across projects
- •The working personality is grown, not written — 58 of its 64 traits came from the sessions themselves
- •Same base model, measurably different behaviour — the corpus is the difference

Plumb — A Design Module That Verifies Its Own Output
- •Three models draw the same brief in parallel; a script — not a prompt — checks what came back
- •Eight layout invariants, each one earned from a real defect that shipped
- •A third verdict beyond pass and fail: “could not check”, which is not the same as clean
- •Published open source under MIT — no account, no telemetry, no npm install
Production RAG Assistant — From Demo to Dependable
- •Hardened a production 10-agent RAG assistant for a Top-5 global CPG client
- •Compliance answers over a 200+ page marketing code, across 5 user segments
- •Diagnosed the prompt-only ceiling, then architected the agent that removed 96% of tone violations
- •Security assessment: surfaced a live auth bypass and exposed credentials
- •Retrieval & prompt improvements; ReAct agent optimization (~2.9x faster)
- •Went from external assessor to first merged contributor PR

AISH — AI Software House
- •Enterprise-grade quality for small and growing businesses
- •We don't build apps. We solve business problems and drive growth.
- •Digital team members with empathy — not bots generating code
- •Communication designed on neuroscience principles

ENIA — Enterprise AI Agents Platform
- •What an enterprise buys is not the automation — it is the governed decision around it
- •Agents run the process; the operator is pulled in only to approve the exception
- •Requirements written from discovery with Oil & Gas majors, not from a spec handed over
- •Visual workflow builder for AI agents — nodes, variables, routing, MCP / API tools
- •On-prem inside the client perimeter; audit-ready lineage, approvals, decision rationale
Requirements Machine — Idea to Enterprise-Ready Backlog
- •Five agents across six integrations — Jira, Confluence, Figma, GitHub and the repositories themselves
- •Turns a feature idea into Jira Epics & User Stories, with competitor research folded into the analysis
- •Reads both repositories and the existing Jira backlog before writing anything
- •Gherkin acceptance criteria (Given/When/Then), matched to the team’s own style
- •Generates a wireframe PNG of the proposed UI from the same analysis
- •Human-in-command: proposes, waits for approval, then publishes

Job-Hunt Agent — A Forward-Deployed Agent, With Me in Command
- •Scans official ATS APIs (Greenhouse / Lever / Ashby) — 2,300+ real roles, legally
- •Transparent matching: profile fit + seniority/comp tier + relocation fit
- •Human-in-command cockpit with an in-app command console
- •Novel architecture: the agent’s brain runs on a subscription, not a paid LLM API

Atlas — CDP Analytics Orchestrator
- •Orchestrator agent: a plain-language request → a delivered dashboard
- •613× cost reduction · 54M+ records → dashboards in ~4 minutes
- •Snowflake data access via MCP · Azure DevOps work-item sub-agent
- •Human-in-command: approval required before every step

Business Case Maker — Multi-Agent Orchestrator
- •Built for pre-sales: the demo a prospect is shown, prepared in an hour instead of a week
- •Multi-agent orchestrator: a request + a domain → a ready-to-build workflow
- •Analyst agents read the target platform (nodes, APIs, UI) to ground every scenario
- •Domain researcher finds real business cases; scenario generator drafts the flow
- •Human-in-command: you pick the scenario, approve or edit, before publish
Core Expertise
AI Product Leadership (Staff / Principal)
Owning AI products end to end — 0→1 to scale, from vision to production
Multi-Agent & Agentic Systems
Orchestrated AI agents for complex enterprise workflows, with human-in-command control
Forward-Deployed Delivery
Embedded with the client and users — shipping the fix, not just the findings
AI Evaluation & Quality
Red-teaming, RAG quality, and root-cause diagnosis — correctness over plausibility
Enterprise AI Integration
RAG pipelines, MCP / tool-calling, policy gates, and on-prem deployment at scale
AI Governance & Traceability
Compliance, auditability, and decision traceability for regulated environments
Get In Touch
Have a project in mind or want to collaborate? Send me a message!