ValueOps Topics
Artificial Intelligence (AI) Planning, Governance, and Portfolio Management
AI planning, governance, and portfolio management connect AI funding decisions, oversight controls, and business outcomes into a single accountable framework. By bridging the gap between rapid AI adoption and the cost, risk, and value discipline enterprises already apply to every other investment, it empowers leaders to fund the right initiatives, control runaway costs, and prove measurable ROI.
How should enterprises govern AI investments?
Enterprises should govern AI investments the way they'd govern any other capital-intensive, high-risk spending: approval scrutiny should scale with the size and risk of the investment, and a named owner should stay accountable for actual cost and outcomes after the money is spent, not just at the moment funding is requested. Treating an AI investment as a single up-front approval, with no further checkpoint until a year-end budget review, is what allows cost and scope to drift far past what was originally approved.
How should organizations prioritize and fund competing AI initiatives?
Organizations should prioritize and fund competing AI initiatives by tying every dollar of AI spend to an explicit, trackable objective before funding is approved — not by funding whichever initiative generates the most internal enthusiasm. In practice, this means running AI investment decisions through the same discipline applied to any strategic portfolio: clear objectives and key results (OKRs), a defined budget, and a named owner accountable for tracking progress against both.
A workable prioritization and funding approach has three components:
Every AI initiative should map to a specific business outcome rather than being funded as generic "innovation" spending. Initiatives that can't be mapped to a strategic theme are the first candidates to be deferred or cut.
Early in an initiative's life, the benefit case is often non-financial: solution innovation, competitive positioning, talent development, scalability testing. Funding decisions at this stage should be judged against those non-financial targets, not against a premature ROI figure. Once an initiative moves out of R&D and into production, the benefit plan should convert to financial terms (cost savings, revenue impact) and funding should be re-evaluated against that harder bar.
Competing AI initiatives need to be evaluated side by side, with real-time visibility into where resources and budget are actually going, so leaders can reallocate from underperforming initiatives to higher-value ones as evidence comes in, rather than discovering the mismatch at year-end.
What causes AI funding decisions to go wrong even with a good framework?
The most common failure is "overly passionate pioneers" driving costs up faster than a strategic case can be built to justify them. Without a defined budget and phase-gated review points, exploratory AI work expands past its original scope before anyone evaluates whether it's still worth funding.
How can organizations prevent AI investment sprawl and control runaway AI costs?
Organizations can prevent AI investment sprawl by replacing after-the-fact monthly bill reviews with real-time, granular visibility into AI consumption, because by the time a monthly invoice arrives, the budget damage has already happened. AI cost behavior breaks the assumptions traditional total cost of ownership (TCO) models were built on: most AI spend lands in operating expense rather than a predictable capital plan, and a single autonomous agent that malfunctions can generate token consumption on the scale of hundreds of thousands of user sessions in a single weekend. Traditional forecasting cannot anticipate that kind of spike.
Controlling this volatility requires two connected disciplines:
-
Real-time consumption tracking: Costs need to be visible at the department or project level as they're incurred, not reconstructed weeks later from a bill. The goal is connecting daily operational consumption directly back to the original investment intent, so a spike is caught while it's still a line item, not after it's a budget crisis.
-
Correct financial classification of token spend: Many organizations still book all AI token consumption as a flat operating expense, which understates margins and obscures the actual return on AI-driven development. Tokens consumed while building a durable software asset function analogously to direct materials in manufacturing, and should follow a comparable capitalization path where governed by the applicable accounting standard, rather than being treated identically to an electric bill. This requires granular attribution (tagging every token request to a specific project and resource), phase-gate mapping (distinguishing exploratory/research token use, which stays expensed, from application-development token use, which may be capitalizable), and consumption-based tracking that ingests usage data from AI providers with the same rigor traditionally applied to labor timesheets.
Cost Treatment: Traditional Labor-Based Development vs. AI Token-Based Development
| Category | Labor-Based Development | AI Token-Based Development |
|---|---|---|
| Primary Cost Driver | Human labor hours | Token/compute consumption |
| Capitalization Precedent | Well-established (labor hours during application development historically capitalized) | Emerging (requires deliberate mapping to comparable accounting treatment) |
| Cost Attribution Method | Timesheets tied to projects and tasks | Granular tagging of token requests to projects and resources |
| Billing Visibility | Predictable, tied to headcount and schedule | Volatile, usage can spike sharply with little warning |
| Forecasting Reliability | High (bounded by team size and working hours) | Low under traditional models (a single automated process can consume resources at a scale disproportionate to team size) |
How do you measure ROI from AI initiatives?
Measure ROI from AI initiatives by comparing a documented before-AI baseline (timelines, effort, cost, output quality) against actual post-deployment performance — and by measuring flow and outcomes, not proxy metrics like story points or the number of AI licenses deployed. A common trap is stopping at surface-level velocity metrics. AI coding assistants can measurably speed up code writing, but if that gain evaporates in downstream review queues, time-to-market never improves and the "ROI" was illusory.
Reliable AI ROI measurement rests on a few concrete practices:
According to Gartner, only 54% of AI projects make it to production, and fewer than half of those deliver meaningful results — a gap that is difficult to close if no one captured what "before AI" actually looked like. ROI has to be calculated as net benefit against pre-deployment cost and performance, not asserted after the fact.
A useful formula: Flow Efficiency = Active Work Time ÷ Total Cycle Time × 100. Industry benchmarks commonly put flow efficiency below 15%, meaning most delivery time is spent waiting in queues, not being actively worked. According to McKinsey, AI coding assistants typically reduce active development time by 20% to 40%, but if flow efficiency is already low, the highest-ROI move is fixing the wait states in the process, not buying more AI capacity.
Standard DORA metrics, particularly Change Failure Rate and Mean Time to Restore (MTTR), act as a guardrail against velocity gains that quietly increase defect density (sometimes called the "AI tax"). In organizations that have paired AI-based observability with these practices, case studies from Meta have shown MTTR reductions of 30–70%, illustrating that AI's ROI shows up as much in incident recovery as in raw output speed.
Standard project-cost models (people, process, technology, service) systematically understate AI project costs because they don't account for the cost and complexity of preparing data for AI use. A Forbes analysis identified this gap as a key reason AI ROI calculations frequently miss the mark.
What's the highest-leverage fix when AI isn't producing measurable ROI?
If a 30-day baseline analysis shows low flow efficiency, the highest-ROI lever is almost always reengineering the wait states in the delivery process. Faster code generation on top of a slow review or approval process just produces a larger backlog of unverified work; it doesn't move the outcome that ROI is actually measured against.
What is an AI agent, and how is autonomous AI changing portfolio and delivery management?
An AI agent is a technology system that receives a complex goal, reasons about how to accomplish it, devises a plan, and then autonomously executes the steps required without step-by-step human direction at each stage. This is a meaningful evolution from earlier AI assistants, which respond to individual prompts but don't independently plan or execute multi-step work. In portfolio and delivery management specifically, this shift means AI can move from suggesting a next step to actually executing a chain of them: validating demand, drafting a business case, integrating a proposed initiative into a portfolio view, and communicating status to stakeholders, largely unsupervised.
That autonomy changes where the real constraints in delivery show up:
-
The bottleneck shifts: AI can meaningfully accelerate the mechanics of producing work, but AI-generated output can also be verbose or contain subtle errors, which pushes the constraint downstream into peer review and validation. An organization that only measures how fast work gets *created* will miss this entirely.
-
Speed without guardrails compounds risk: Accelerated output without corresponding rigor in validation tends to show up as rising defect density alongside the improved velocity numbers. The right response isn't to slow AI adoption down; it's to instrument the review and validation stage as closely as the creation stage.
- AI is a multiplier, not a fix: An AI agent bolted onto a delivery process with a broken testing practice accelerates the problem, making failures happen faster. Foundational process health has to come before autonomous agents are layered on top of it.
What is the difference between an AI agent and a traditional AI copilot or assistant, and how does it impact governance and planning?
The difference between an AI agent and a traditional AI copilot or assistant is autonomy. A copilot responds to a specific prompt and waits for the next one, while an agent takes a complex goal, plans the steps needed to achieve it, and executes that plan across multiple steps without requiring a new prompt at each stage. A copilot is fundamentally reactive, while an agent is fundamentally goal-directed.
This distinction matters practically for governance and planning. Copilots carry a lower risk profile because a human is in the loop at every step, while agents require robust governance before being given autonomy. By the time an agent has acted, there may not be a human checkpoint that caught the action before it happened.
AI Copilots vs. AI Agents
| Category | AI Copilot | AI Agent |
|---|---|---|
| Autonomy Level | Responds to a single prompt at a time | Plans and executes a multi-step sequence towards a goal |
| Initiation | Human-initiated for each step | Human sets the goal, system initiates the step itself |
| Human Involvement | Continuous | Checkpoint-based |
| Task Scope | Narrow, single-task assistance | Broader workflows spanning multiple tasks and systems |
| Governance Requirement | Lower | Higher |
| Typical Use Case | Drafting content, answering questions, code suggestions | End-to-end workflows |
What does AI governance maturity look like?
AI governance maturity describes how consistently an organization can trust, audit, and control its AI systems as they scale. Most organizations are still early on that curve. As organizations move from simple AI assistants to autonomous AI agents that plan and execute multi-step work independently, informal oversight breaks down — governance has to be designed into the system.
Three structural gaps typically block mature AI governance:
AI systems are frequently perceived as "black boxes," where the path from input to recommendation is invisible to the people accountable for the outcome. The fix is explainability: surfacing the decision drivers, assumptions, and strategic-alignment factors behind a recommendation, plus a time-stamped audit trail showing how a recommendation evolved. Now AI functions as a decision advisor stakeholders can interrogate, not a decision maker they have to trust blindly.
Governance frameworks need business rules tailored to different investment types, automated anomaly detection on external data sources, and clear control over exactly what data an AI system is allowed to draw from.
As agents gain the ability to act, access has to be scoped tightly: role-based permissions aligned to a user's actual responsibilities, controls on what document/data types an AI system can process, and isolation for any model handling proprietary information.
Why does explainability matter more for AI agents than for earlier generations of business software?
Traditional software follows fixed, auditable logic. An AI agent's reasoning path can shift with new data, prompts, or model versions, so a decision that was defensible yesterday may not be reproducible today. Explainability is what lets an organization validate results, catch bias, and explain a recommendation to a non-technical stakeholder after the fact, which fixed-logic software never had to account for.
Ungoverned AI vs. Governed AI
| Category | Ungoverned AI | Governed AI |
|---|---|---|
| Decision Transparency | Produces recommendations with no visible reasoning | Surfaces its decision drivers, assumptions, and alignment factors at the point of recommendation |
| Data Quality | Draws on whatever data it’s given | Runs continuous, automated validation and anomaly detection on the data feeding its decisions |
| Audit Trail | Leaves no built-in record of how or why a recommendation changed over time | Generates a time-stamped audit trail automatically as its recommendations evolve |
| Response Time to Anomalies | Bad inputs or output surface only when someone happens to notice the downstream effect | Flags anomalies in real time through automated alerts |
| Security/Access Scope | Operates with broad, standing access to data and systems | Access is scoped narrowly per task through role-based permissions, with isolation for proprietary data |
| Scalability of Oversight | Requires proportionally more human reviewers as AI usage grows | Oversight is built into the system itself, so it scales with AI adoption without added headcount |
How do organizations progress toward AI governance maturity?
Organizations progress toward AI governance maturity along a rough three-stage path, moving from opportunistic, ungoverned AI use toward governance that's built into the systems themselves rather than applied as a manual compliance step:
-
Reactive: AI use is opportunistic and largely ungoverned. Individuals or teams adopt tools independently, and issues (cost, data quality, security) surface only after they've already happened.
-
Structured: Formal policies exist for data quality, security scoping, and cost attribution, but they're applied manually and reviewed on a fixed cycle rather than continuously.
- Embedded: Governance is built into the systems themselves. Explainability, cost attribution, and security scoping happen automatically as a byproduct of how the work is done, rather than as a separate compliance step.
Ready to Elevate Your Portfolio Strategy?
Our experts will guide you through onboarding and help you start managing value, not just projects.