The situation
You paid for an AI application, or a team built one quickly, and it is not ready for customers. Perhaps a supplier delivered a working demo and then left. Perhaps the original developers moved on and nobody remaining can explain how it works. Perhaps it handles the scripted cases and fails on real ones, costs more per use than expected, or leaks data between users. A launch date, a customer commitment or an investor update is on the calendar.
Your team is split. Some want to patch it; others want to start again. Neither side has evidence, and both are partly guessing about what the codebase contains.
This page is for that specific decision about one application or system. It is not about whether staff use AI tools (that is adoption) or about a portfolio of ownerless pilots (that is a programme). It is about one codebase and a question: repair, rebuild or stop?
What the review covers
Can it be run at all? I rebuild the application from source in an environment you control. What cannot be rebuilt from the handover, such as missing infrastructure code, undocumented prompts, credentials held by a supplier or absent training data, is the first finding.
What is actually there? Components, dependencies, data flows, external services, model calls and where state lives. AI-generated and rushed codebases often have duplicated logic, secrets in code and no tests on the paths that matter.
Is it safe? Authentication, authorisation between users and tenants, handling of personal data, prompt-injection exposure and what the model can do with the tools it has.
Does the AI part work? I run representative cases from the intended use and review outputs with your domain experts, separating model, prompt, retrieval, data and task-design failures.
Can anyone operate it? Deployment, monitoring, cost per use, failure handling and whether your team has the skills to own it.
The signature deliverable
You receive a repair-versus-rebuild decision, critical failure inventory and bounded recovery backlog. Illustrative example of a failure inventory extract:
| # | Finding | Severity | Evidence | Repair effort |
|---|---|---|---|---|
| 1 | Users can retrieve other tenants’ documents through the chat endpoint | Critical | Reproduced with two test accounts | Moderate: tenant filter in retrieval query |
| 2 | Infrastructure created by hand; no code to recreate it | High | Supplier repository has no deployment config | Moderate: rebuild as code |
| 3 | Extraction prompt fails on scanned documents | High | 11 of 30 sample cases failed | Unknown until data is reviewed |
| 4 | No tests on billing or permission paths | High | Coverage report | Moderate |
Illustrative example showing the format, not a client deliverable.
The recommendation sets out which path costs least to reach a shippable, ownable system, what each path assumes, and which components are worth keeping whatever you decide.
How the review is judged
It is a fixed-scope diagnostic. It is complete when the sponsor has the failure inventory with evidence for each item, a recommendation with reasoning they can challenge, and a backlog specific enough for any competent team to pick up. Diagnosis is kept separate from implementation: you are under no obligation to commission the recovery from me.
Ownership and handover
Everything produced is yours: the inventory, the backlog, any small evaluation set and the notes on how to run the application. If you proceed, the backlog can go to your team, a contract engineer under your manager, another supplier or a separately scoped delivery with me. I do the work personally; any specialist help is disclosed and approved by you first.
Boundaries
If the application was built by you with an AI coding tool and mainly needs hardening for launch, take an AI-generated application into production is the closer fit. If the problem is a portfolio of pilots without owners, see AI programme rescue. If staff are not using tools you bought, see stalled AI rollout rescue. If you are taking over a system and need its operation transferred, see AI system handover.