
Phantom Productivity: Why Your Company Has 25 AI Agents Doing the Exact Same Thing
Picture 4 departments demoing their new custom AI agent to leadership in the same quarter.
Sales ops has an agent that tracks 10 competitors and writes a weekly summary. Strategy has an agent that tracks 10 competitors and writes a weekly summary. Product marketing has one that tracks those exact same 10 competitors and writes a weekly summary. A regional team has one too, except theirs has slightly different wording in the prompt template.
If you work in a large company, you've probably seen some version of this already.
Each team reports a 4x ROI and hours saved. On the quarterly AI portfolio dashboard it looks like an easy win: 4 agents shipped, 4 productivity gains, $240k in reported value.
In reality, the company paid 4 times to build, maintain and run the same brief.
I think this is one of the biggest blind spots in enterprise AI right now, and it's one I spend a lot of time thinking about in strategic planning and portfolio management. We fund, build and celebrate duplicate agents across departments while counting every copy as brand-new productivity.
The math says otherwise.
The hidden assumption behind every corporate AI dashboard
Corporate productivity accounting was designed for an era where building software was slow and expensive. Every application required dedicated engineering resources, formal arch reviews, and months of dev time. If an engineering group shipped a new internal tool, leadership could safely assume it added distinct operational capacity to the firm.
Generative AI broke that foundation completely.
Today, any PM, analyst, or biz ops specialist can spend an afternoon prompting an agent into existence. When the marginal cost of creating software collapses toward zero, measuring productivity by raw output volume becomes a dangerous trap.
If 10 separate teams build 10 separate agents that generate 10 identical market research briefs, traditional accounting tallies 10 units of delivered value. But the company only needed 1. The remaining 9 copies add zero incremental value to the business. They only add API consumption, maintenance tickets, and noisy Slack channels.
Traditional metrics count every output as if it were distinct. They never de-duplicate.
The core equation: Perceived vs. Real Productivity
To fix this, we need to separate what teams report from what the company actually captures. The relationship boils down to a single equation:
Here is how the terms break down:
: The productivity ratio your AI council or PMs report on slides (). Every agent is counted as if its output were entirely novel.
: The Uniqueness Ratio (ranging from 0 to 1), representing the share of total output that is genuinely distinct across the org.
: Actual enterprise productivity after all redundant, overlapping outputs are collapsed.
Notice the mathematical punchline: cost drops out of the ratio.
Uniqueness alone sets the gap between perceived and real productivity. Operating costs and token prices do not create the illusion; they only decide when real productivity drops below break-even ().
Two immediate formulas follow from this:
1. Phantom Productivity
The portion of reported productivity gains that exists only on paper:
This is the phantom gain that keeps executive decks looking green while operating margins quietly erode.
2. The Negative Zone
Real productivity slips under water (meaning total costs exceed real enterprise value) whenever:
This leads directly to a practical rule of thumb for portfolio planning:
The N-Copies Rule: If your teams report an N-fold return on something a single copy could have delivered to the entire company, N copies wipe out the entire economic gain.
For identical copies, breakeven is exactly . For copies with an average similarity , the exact number of duplicates that pushes the entire group into net losses is:
When a team reports a 4x ROI on an internal brief, 4 identical copies wipe out the gain entirely. With 90% similarity, the 5th copy tips the entire portfolio into negative returns.
Case Study: The Competitor Intelligence Agent
Let's look at an example that virtually every enterprise tech worker has seen demoed over the last 12 months: an autonomous competitor intelligence agent.
Product, sales, marketing, strategy and every regional team each build an agent to monitor the same 10 competitors and generate a weekly research brief. They pull from the same public web sources, analyze the same product launches, and synthesize roughly 90% identical output. A single company-wide brief would serve every one of them.
Per copy, per year (at today's subsidized token prices where ):
Cost Line Item | Annual Amount | Notes |
|---|---|---|
Build tokens | $300 | Multi-sprint vibe coding, prompt iteration, tool configs |
Run tokens | $2,100 | Weekly deep-research loops across 10 targets (~$40/run) |
Maintenance tokens | $600 | Prompt fixes, broken web scrapers, model upgrades |
Tokens subtotal | ≈ $3,000 | Direct API spend |
Human labor | ≈ $12,000 | Build hours, weekly prompt review, upkeep |
Total Cost () | ≈ $15,000 | Total annual cost per agent copy |
Each team values its weekly brief at $60,000/yr based on analyst hours saved and accelerated strategic reaction times. On paper, each team reports a 4x ROI ($60k value on $15k cost).
The copies share 90% similarity (). Look at what happens to real enterprise ROI as independent teams stand up their own copies:
Teams with their own agent | Uniqueness () | Perceived ROI | Real ROI | Net Real Value | Portfolio Status |
|---|---|---|---|---|---|
1 | 1.000 | 4.0x | 4.0x | +$45k | Profitable |
3 | 0.360 | 4.0x | 1.4x | +$19k | Narrow gain |
5 | 0.220 | 4.0x | 0.87x | −$10k | Underwater (Loss) |
10 | 0.110 | 4.0x | 0.44x | −$84k | Capital destruction |
25 | 0.044 | 4.0x | 0.18x | −$309k | Severe loss |
50 | 0.022 | 4.0x | 0.09x | −$683k | Value incinerator |
Here is the operational reality of that table:
At 25 teams, the quarterly AI steering committee deck announces: '25 autonomous agents shipped, 4x ROI, +$1.1M in enterprise value created.'
The true financial outcome is an ROI of 0.18x and a net loss of $309,000.
The company paid 25 separate times for build tokens, 25 separate times for weekly deep-research loops hitting the exact same URLs, and hundreds of human hours maintaining prompt files that all do the same thing.
The subsidy cliff: what token repricing does to the math
A lot of leaders assume token costs are a rounding error because current API rates look cheap. That assumption misses how consumption economics work.
Bills equal price multiplied by volume. Agentic workflows do not make a single query; they run stateful loops, spin up subagents, browse web pages, and consume hundreds of thousands of tokens per task. More importantly, enterprise teams inevitably standardize on frontier reasoning models rather than legacy commodity endpoints.
Today's enterprise pilots run on subsidized pricing as labs burn venture cash to capture market share. When AI providers reprice their models to reflect actual compute economics, the token price multiplier jumps.
When climbs, base costs rise, which drags down . That means the level of uniqueness () an org needs to stay above water goes up dramatically:
Token Multiplier | Annual Cost Per Copy | Perceived ROI | Duplicate Team That Pushes the Portfolio Negative |
|---|---|---|---|
(Today's rates) | $15,000 | 4.0x | 5th team |
$27,000 | 2.2x | 3rd team | |
$42,000 | 1.4x | 2nd team |
At a 10x token repricing, a single duplicate copy is enough to wipe out all profit and push the entire corporate initiative underwater.
Duplication turns what should be an operational buffer into extreme exposure to API pricing revisions.
Scope: where the equation applies (and where it differs)
To apply this framework responsibly, you have to distinguish between rival and non-rival deliverables.
The equation applies strictly when the output is non-rival, meaning a single centralized copy serves the entire company just as effectively as 20 separate copies. Examples include:
Competitor intelligence briefs
Market research syntheses
System architecture docs
Compliance and security audit scans
The shared agentic capability itself
For per-user utility tools (for instance, 30 regional business units each vibe coding their own slide-deck formatting assistant), duplication behaves differently. The output is rival here: each business unit genuinely needs its own decks formatted. The value delivered to employees does not vanish.
Instead, duplication inflates the cost of acquiring the capability:
For per-user tools, one shared platform tool would burn roughly the same per-query run tokens as 25 fragmented tools. The waste here is not the per-query inference. The waste is the duplicated build tokens, the 25 separate maintenance efforts, and redundant fixed infra overhead like re-indexing the same corporate knowledge bases 25 times.
The 2x2 Framework: Problem Uniqueness vs. Solution Uniqueness
Before funding an AI initiative, portfolio planners need to map every proposed artifact across two dimensions:
Same Solution Arch | Different Solution Arch | |
|---|---|---|
Same Problem | Duplication Tax | Controlled Exploration |
Different Problem | Platform Opportunity | Genuine Novelty |
Most enterprises assume their AI portfolio sits in the bottom-right (novelty) or top-right (exploration) box. Without a registry and a search-before-build habit, nobody can actually check, and uncoordinated adoption drifts toward the top-left: identical solutions aimed at identical problems.
Measuring : the 5-step operational pipeline
You cannot manage uniqueness through manual surveys or subjective team interviews. Teams will always insist their agent has a unique perspective. You have to measure uniqueness at the architectural layer.
Here is the operational pipeline to compute directly from internal code repos, agent registries, and MCP server logs:
[Agent Registries / Repos / MCP Configs]
│
▼
Step 1: Semantic Fingerprinting
(Problem p_i, Solution s_i, Data d_i, Output o_i)
│
▼
Step 2: Weighted Similarity Matrix (K_ij)
│
▼
Step 3: Per-Artifact Uniqueness Credit (U_i = 1 / Σ K_ij)
│
▼
Step 4: Value-Weighted Uniqueness Ratio (Ū = Σ U_i V_i / Σ V_i)
│
▼
Step 5: Executive Portfolio Metrics
(Effective Solutions, Duplication Tax, Real ROI)Step 1: Semantic Fingerprinting
Extract the fundamental job-to-be-done for every AI artifact across internal registries, repos, prompt files, and MCP configs:
: Problem embedding (the operational task, user persona, decision being supported)
: Solution embedding (agent architecture, prompt patterns, tool definitions)
: Set of input data sources connected (internal databases, APIs, file systems)
: Set of output entities, schemas, or metrics produced
Fingerprint the underlying functional job, not the repo title or UI wrapper. Adding a custom CSS theme or renaming an agent from 'MarketTracker' to 'CompetitorScout' does not make an agent unique.
Step 2: The Similarity Matrix
Construct an kernel matrix comparing every artifact against artifact :
where , is cosine similarity rescaled so unrelated tasks map to 0, and is the Jaccard similarity across data and output sets. Typical weights: , , , .
Step 3: Per-Artifact Uniqueness Credit
Calculate the individual uniqueness credit for artifact :
If 5 identical agents exist in the org, each one receives a uniqueness credit of . A fully unique agent receives . Real local variations naturally lower and award proportional credit without requiring binary all-or-nothing arguments.
Step 4: Value-Weighted Uniqueness Ratio
Aggregate across the entire portfolio:
Weighting by value () ensures that duplicating a mission-critical revenue engine damages portfolio health far more than duplicating an internal lunch-order bot.
Step 5: Executive Portfolio Metrics
From these values, generate the real governance numbers:
Effective Solution Count: The Vendi score of (exponential of the spectral entropy). This gives leaders an honest headcount of unique capabilities, something like: 'We funded 140 agent initiatives this year; we have 31 distinct solutions.'
Duplication Tax:, quantifying the exact dollar spend absorbed by redundant build hours, duplicate maintenance, and unneeded API calls.
The Knowable Duplication Rule: Only penalize duplication that was knowable. If Team B starts building an agent when Team A's equivalent system was already registered in prod and discoverable in the internal registry, Team B absorbs the full duplication penalty. If both were started concurrently during an active evaluation window, it counts as legitimate exploration.
Reference Implementation (Python)
Here is a reference implementation you can run against your own agent metadata. It's written for readability, not speed; for thousands of artifacts you'd want to vectorize the similarity matrix:
import numpy as np
def rescaled_cos(a, b, floor=0.25):
'''Cosine similarity rescaled so unrelated pairs map to 0.'''
c = float(a @ b / (np.linalg.norm(a) * np.linalg.norm(b)))
return max(0.0, (c - floor) / (1.0 - floor))
def jaccard(a: set, b: set):
'''Jaccard index between two categorical feature sets.'''
return len(a & b) / len(a | b) if a | b else 0.0
def similarity_matrix(arts, weights=None):
if weights is None:
weights = {'p': 0.4, 's': 0.2, 'd': 0.2, 'o': 0.2}
n = len(arts)
K = np.eye(n)
for i in range(n):
for j in range(i + 1, n):
a, b = arts[i], arts[j]
score = (
weights['p'] * rescaled_cos(a['p'], b['p']) +
weights['s'] * rescaled_cos(a['s'], b['s']) +
weights['d'] * jaccard(a['d'], b['d']) +
weights['o'] * jaccard(a['o'], b['o'])
)
K[i, j] = K[j, i] = score
return K
def effective_solution_count(K):
'''Vendi score: exponential of the spectral entropy of K.'''
eigenvalues = np.linalg.eigvalsh(K / len(K))
eigenvalues = eigenvalues[eigenvalues > 1e-12]
return float(np.exp(-(eigenvalues * np.log(eigenvalues)).sum()))
def uniqueness_report(arts, years=1.0, pi=1.0):
'''
Evaluates enterprise portfolio uniqueness and real productivity.
pi: token repricing multiplier (1.0 = today's subsidized API pricing).
'''
K = similarity_matrix(arts)
U = 1.0 / K.sum(axis=1)
V = np.array([a['value'] for a in arts])
# Cost model: build tokens + maintenance + duplicate run tokens
tok = pi * np.array([
a['tok_build'] + (a['tok_maint'] + a['tok_dup_run']) * years
for a in arts
])
hum = np.array([a['human'] + a.get('drift', 0.0) for a in arts])
C = tok + hum
u_bar = (U * V).sum() / V.sum()
p_perceived = V.sum() / C.sum()
p_real = u_bar * p_perceived
esc = effective_solution_count(K)
return {
'artifacts_funded': len(arts),
'effective_solutions': round(esc, 1),
'uniqueness_ratio': round(u_bar, 3),
'perceived_productivity': round(p_perceived, 2),
'real_productivity': round(p_real, 2),
'phantom_productivity': round(p_perceived - p_real, 2),
'under_water': bool(p_real < 1.0),
'token_duplication_tax': round(float(((1.0 - U) * tok).sum())),
'human_duplication_tax': round(float(((1.0 - U) * hum).sum())),
}Answering the common objections
Whenever I walk tech leaders through this math, I hear the same 5 arguments. Here is why they fail under scrutiny:
1. 'Token prices keep falling every quarter.'
Inference bills equal price multiplied by volume. Agentic loops make dozens of recursive calls, execute automated retries, and run multi-step tool sequences. Volume is expanding much faster than per-token prices are dropping.
Furthermore, serious enterprise systems rely on frontier reasoning models, not older commodity tiers. Subsidized rates can disappear with a single pricing update from OpenAI or Anthropic. If you have not hedged token risk through architectural consolidation, duplicate agents multiply your financial exposure across every business unit.
2. 'Each department still accelerated their own workflow.'
Individual team velocity does not equal corporate productivity. A company is an economic unit with a single consolidated P&L. If 4 regional teams spend corporate budget to generate the exact same competitive brief, the firm absorbs 4 times the capital outlay for 1 unit of unique output. Corporate productivity is measured where the money is spent, not inside the boundaries of an isolated team demo.
3. 'Our team requires distinct local context.'
Local nuances are real, and that is exactly why the algorithm uses continuous similarity scoring rather than a crude duplicate flag. If an EMEA marketing team genuinely incorporates localized regulatory constraints or distinct competitor segments, its similarity score decreases and it receives higher uniqueness credit.
However, running a bespoke agent just to change 2 sentences in the output summary does not justify a duplicate annual run budget.
4. 'Duplication is how healthy exploration happens.'
Uncoordinated duplication is organizational friction, not exploration. True technical exploration requires a structured bake-off with clear evaluation benchmarks, a predefined testing timeline, and a clear plan to decommission the losing variants. Building in isolation because you never checked internal repos to see if an identical agent exists is simply lack of discoverability.
5. 'Goodhart's Law will cause teams to game the system.'
Teams game metrics when you measure superficial attributes like repo names or user interfaces. Because this algorithm fingerprints underlying data lineage, schema inputs, and decision targets, cosmetic rebranding does not move the score.
More importantly, belongs in portfolio governance reviews with the CIO, CTO, and finance leads. Never tie it to individual engineer performance reviews, or teams will hide their builds.
What CIOs and portfolio leaders need to do on Monday
If you want to stop paying the duplication tax inside your enterprise, take these 4 concrete steps:
Establish an internal AI Registry: Require every agent, MCP server, and automated prompt pipeline to register its metadata, data sources, tool schemas, and operational decisions before entering prod.
Implement search-before-build gating: Make semantic similarity checks part of your project intake. Before a department receives budget to build a custom agent, run its job-to-be-done against the registry. Pick a similarity cutoff (0.7 is a reasonable place to start, then calibrate it against pairs your team agrees are true duplicates). Above it, the default decision is reuse or extension, not a greenfield build.
Promote high-utility agents into shared platforms: When an agent proves useful across 3 or more departments, pull it out of the individual business unit. Transition it into a centrally maintained enterprise service with governed MCP endpoints.
Report Effective Solution Count on executive dashboards: Beside the vanity metric of '140 AI agents deployed,' publish the real numbers, for example: '31 effective solutions, Uniqueness Ratio 0.22.'
When generating software costs almost nothing, raw output volume ceases to be a meaningful metric. Unique differentiation is the only thing that creates real enterprise value.
Comments (0)
No comments yet. Be the first to share your thoughts!