Artificial Intelligence

The AI Portfolio Review Needs a Kill List

An AI portfolio review needs explicit scale, repair, or stop decisions so weak pilots release capital, capacity, and risk before inertia wins.

An executive removes one dark project tile from a boardroom portfolio while viable projects remain illuminated in gold, representing disciplined AI investment decisions.

A useful AI portfolio review should force every pilot into one of three decisions: scale it, repair one defined uncertainty, or stop it. Anything called “paused” needs a named dependency, an owner, and a decision date. Without those, pause is just a stop decision that no one wants to make.

That is what I mean by a kill list. It is not a quota for failed projects or a weapon against the people who proposed them. It is a visible record of work whose next dollar no longer has a defensible case. Used well, the list returns capital, scarce technical capacity, and risk budget to better opportunities.

The approach completes the loop between planning an AI budget, measuring the fully loaded economics of AI, and acting when the evidence changes. A portfolio that can start experiments but cannot stop them is not learning. It is accumulating obligations.

Continuity Is a Decision Too

AI pilots are easy to approve because each one appears small. The portfolio cost hides in the total: model and platform consumption, data work, security review, integration, evaluation, support, and the attention of people who could be solving another problem.

The review therefore has to compare the next increment of spending with the next best use of those resources. Money already spent may explain how the organization learned, but it does not make the next funding request stronger. HM Treasury’s 2026 Green Book makes the same appraisal distinction: sunk costs should not drive the next decision, while the opportunity cost of continuing to use resources still matters.

There should not be one universal six-month deadline. A document assistant and an AI-enabled clinical workflow have different evidence and assurance needs. The important discipline is to set the timebox before work begins, match it to the uncertainty being tested, and prevent the team from moving the finish line after seeing the result.

Write the Review Contract Before the Pilot

Every pilot should enter the portfolio with a short review contract. If the team cannot fill this out in plain language, it is probably not ready to spend money.

Review field What the portfolio needs to know
Business outcome What changes for a customer, employee, operation, revenue stream, or risk exposure?
Baseline How does the work perform today without the proposed AI system?
Evidence target What result would justify the next stage, and how will it be measured?
Full cost What will development, operation, review, control, and support require?
Risk boundary Which failure or residual risk would make continued use unacceptable?
Sponsor Who owns the business outcome and can make the stop decision?
Decision date When will the evidence be reviewed rather than merely reported?
Exit plan How will access, data, integrations, vendors, and infrastructure be retired?

This does not turn innovation into a paperwork contest. The contract should be proportionate to cost, complexity, and consequence. Its purpose is to stop the portfolio from inventing success criteria after the experiment has already produced an ambiguous result.

The NIST AI Risk Management Framework supports this lifecycle view. Its Manage function calls for a decision about whether an AI system achieves its intended purpose and whether development or deployment should proceed. It also calls for mechanisms and assigned responsibility to disengage or deactivate systems whose performance or outcomes conflict with intended use.

Use Three Real Outcomes

At review time, choose one of three outcomes and record why.

  1. Scale when the pilot has produced useful evidence, the operating cost is credible, the risk is within tolerance, and an accountable owner is ready to run it as a supported product.
  2. Repair when one important uncertainty remains testable through a specific, time-boxed change. Repair is not permission to restart the whole pilot with a new story.
  3. Stop when the business need has moved, the evidence misses the threshold, the economics no longer work, the risk cannot be reduced within tolerance, the sponsor is gone, or a simpler non-AI option is better.

Decision diagram showing evidence moving an AI pilot to one of three outcomes: scale when evidence, ownership, risk, and economics are credible; repair one bounded uncertainty; or stop and release capacity while preserving the learning.

The repair category is where portfolios often lose discipline. Give it one hypothesis, one owner, one budget boundary, and one new review date. If the team returns with a different use case, different users, and different success criteria, that is a new proposal rather than evidence that the old one succeeded.

Production readiness also deserves a higher bar than a convincing demo. The minimum viable AI governance model provides risk tiers and evidence gates, while Day Two AI operations makes ownership, monitoring, incident response, and recovery visible. If those obligations make the economics unattractive, the portfolio has learned something important before scaling the liability.

Count Learning Without Inventing ROI

A stopped pilot does not need a fictional return on investment to have value. Report its learning honestly:

  • future spending avoided after the decision;
  • engineering, security, data, and business capacity returned;
  • a risk or dependency retired;
  • an assumption disproved;
  • reusable evaluation data, integration work, or control patterns; and
  • a better screening question for the next proposal.

Keep avoided future cost separate from cash already spent, and do not count released staff time as savings unless the organization can show where that capacity went. The goal is a decision record that improves the next allocation, not an arithmetic trick that makes every experiment look successful.

The kill list should also preserve the reason for stopping. Over time, recurring causes can expose portfolio-level problems: use cases without baselines, missing business owners, data that is not ready, controls added too late, or unit economics that fail under realistic quality requirements. That evidence is more useful than a dashboard showing how many pilots remain active.

Make It Safe to Tell the Truth

Fear is a poor operating model. If stopping a pilot is treated as a mark against the team, people will hide weak signals, narrow the measurement, and keep asking for one more cycle. The governance design should separate the quality of the experiment from the decision to continue the investment.

The business sponsor owns the outcome. The technical team owns the integrity of the evidence. The portfolio forum owns the enterprise trade-off. That division prevents a project team from being asked to defend its existence and prevents one division from protecting a local initiative at the expense of the larger company.

Leaders set the tone by asking, “What did we learn, and what should receive this capacity now?” That is a harder and more useful question than, “Who failed?”

Run the Next Review Differently

Start with every active AI pilot, not just the ones requesting new money. Give each sponsor one page containing the review contract, actual evidence, remaining uncertainty, fully loaded next-stage cost, and recommended outcome.

Then make the scale, repair, or stop decision in the meeting. Publish the internal kill list with the reason, capacity released, reusable learning, and owner of the shutdown work. Finally, verify that credentials, data copies, vendor commitments, integrations, and infrastructure were actually retired. A decision in a slide deck is not decommissioning.

The point is not to kill more AI projects. It is to ensure that every project continues because current evidence justifies it, not because stopping feels uncomfortable.

Further Reading