Most organisations don’t struggle to come up with AI goals.
They struggle to set goals that survive contact with reality.
Leadership wants visible results quickly.
Technical teams need time to understand the data, test assumptions and build systems that can survive production.
The business wants measurable return before committing more money.
At the same time, everyone talks about AI as a long-term transformation involving better data, new workflows, automation and capabilities that may take years to mature.
That creates an apparent contradiction.
Deliver something valuable now.
But don’t make decisions that prevent something more valuable later.
The usual response is to divide the roadmap into quick wins and long-term vision.
That distinction is useful.
The mistake is treating them as unrelated.
A realistic AI goal should do more than produce a short-term result. It should generate evidence or capability that makes the next strategic decision easier.
The goal isn’t simply to win quickly.
It’s to learn cheaply enough that the organisation knows where to invest next.
What Is a Realistic AI Goal?
A realistic AI goal is not simply something the technical team believes it can build.
Consider:
Build an AI recommendation system by Q3.
That’s a project milestone.
It tells you what will be built and when.
It doesn’t tell you why the system matters.
A stronger goal starts with the business outcome:
Increase repeat purchases by improving the relevance of products shown to returning customers.
Now AI becomes a hypothesis about how to achieve the outcome rather than the outcome itself.
The organisation might believe better recommendations will increase repeat purchases.
That hypothesis can be tested.
Business Problem
│
▼
Desired Outcome
│
▼
AI Hypothesis
│
▼
Evidence
│
▼
Invest / Change / Stop
This distinction matters because an AI project can succeed while its goal fails.
The recommendation model ships.
Its offline evaluation scores improve.
Customers barely change their behaviour.
Technically, the project delivered what was requested.
Strategically, the hypothesis was weak.
Setting realistic AI goals means defining success outside the AI system itself, because AI only creates growth when it changes a business constraint.
Start With the Business Change
Before setting an AI target, ask:
What should become different if this works?
For a customer service initiative, the answer might be:
- customers receive answers faster
- representatives handle more requests
- fewer tickets require escalation
- resolution costs fall
For forecasting:
- less inventory is wasted
- stockouts decrease
- procurement decisions improve
- working capital falls
For sales:
- representatives spend more time selling
- qualified leads receive attention sooner
- conversion increases
- sales cycles become shorter
Only after identifying the change should the organisation decide what role AI might play.
This prevents a common planning failure where the AI capability itself becomes the goal.
“Deploy an AI assistant.”
“Implement predictive analytics.”
“Introduce AI agents.”
“Automate with generative AI.”
Those statements describe technologies.
They don’t describe success.
Separate Outcome Goals From Capability Goals
Not every important AI investment will produce immediate revenue or cost savings.
Sometimes the organisation genuinely needs to build capability first.
It may need:
- reliable access to operational data
- an evaluation framework
- model monitoring
- secure access to language models
- better deployment processes
- reusable retrieval infrastructure
- governance for AI systems
- staff with new technical skills
These are capability goals.
They matter because they make future outcomes possible.
But capability goals need a different justification from business outcome goals.
Consider these two statements:
Reduce support handling time by 15%.
and:
Create a standard evaluation framework for customer-facing generative AI systems.
The first should eventually be judged against a business outcome.
The second should be judged against the capabilities it enables.
Does it reduce the effort required to evaluate new systems?
Does it allow teams to compare models consistently?
Does it detect regressions before deployment?
Are multiple production use cases actually using it?
Problems appear when organisations mix the two.
A platform team is expected to demonstrate immediate revenue.
A product experiment is allowed to run indefinitely because it is supposedly building “strategic capability.”
Different goals need different evidence.
Short-Term and Long-Term Goals Need Different Evidence
A long-term AI goal may take years to fully affect the business.
That doesn’t mean the organisation should wait years to discover whether it’s working.
The mistake is expecting the final outcome too early.
Suppose the long-term goal is to use AI to reduce the cost of processing insurance claims.
The eventual outcome may be:
Lower Cost per Claim
But that result depends on several earlier changes.
Lower Cost per Claim
▲
│
More Automated Claims
▲
│
Reliable AI Decisions
▲
│
Representative Data
▲
│
Validated AI Approach
Each stage creates evidence about whether the next investment is justified.
The first goal therefore doesn’t need to be:
Automate 70% of claims within three years.
It might begin with:
Determine whether the system can classify the five highest-volume claim types accurately enough for human-assisted processing.
That goal is smaller.
But it isn’t disconnected from the strategic direction.
It reduces uncertainty about it.
Use Leading and Lagging Measures
Long-term AI goals often fail measurement because organisations choose between two bad options.
Either they measure technical activity that appears quickly:
- models built
- pilots launched
- users onboarded
- prompts executed
Or they wait for final business outcomes that may take much longer to become visible.
A better approach uses both leading and lagging indicators.
Leading indicators tell you whether the operating change is beginning to happen.
Lagging indicators tell you whether that change eventually creates the business result.
Suppose an AI assistant is intended to improve customer support.
Leading indicators might include:
- percentage of representatives using it
- percentage of suggestions accepted
- reduction in time spent searching for information
- correction rate
- escalation rate
Lagging indicators might include:
- average resolution time
- cost per resolved ticket
- customer satisfaction
- repeat contact rate
The model’s evaluation score still matters.
But it belongs underneath these measures.
Business Outcome
▲
│
Lagging Indicators
▲
│
Leading Indicators
▲
│
AI Performance
This gives long-term initiatives something meaningful to demonstrate before the final ROI appears.
A Quick Win Should Answer a Strategic Question
Quick wins aren’t the problem.
Disconnected quick wins are.
A useful short-term AI project should answer something the organisation needs to know.
Can the model perform the task reliably enough?
Will employees use its recommendations?
Can the organisation integrate AI into this workflow?
Will customers accept AI-assisted interactions?
Does automation actually reduce labour once verification is included?
Is the available data sufficient?
Does the economics still work at production volume?
These questions create information.
A quick win that answers one of them can move the long-term strategy forward even if the initial deployment is small.
Connect the Two Horizons
Long-Term Outcome
▲
│
Capability Needed
▲
│
Evidence Required
▲
│
Short-Term Win
▲
│
Hypothesis
▲
│
Business Problem
The short-term win now has a job.
It isn’t merely there to prove that AI can do something impressive.
It’s reducing uncertainty about the next investment.
Why Quick Wins Take Over AI Programs
Quick wins become dangerous when organisations reward delivery without accounting for what happens after delivery.
A small AI project looks inexpensive while it is being developed.
Then it reaches production.
The data pipeline needs maintenance.
The model needs evaluation.
A vendor changes its API.
Costs need monitoring.
Users discover edge cases.
Security reviews need updating.
Someone needs to respond when the output is wrong.
The team that was supposed to move to the next strategic initiative becomes the permanent owner of the previous quick win.
Then another quick win arrives.
And another.
Eventually, most of the team’s capacity is consumed maintaining tactical systems.
Quick Win 1 ───────► Maintenance ─────┐
│
Quick Win 2 ───────► Maintenance ─────┼──► Team Capacity
│
Quick Win 3 ───────► Maintenance ─────┘
Strategic Work ───────────────────────► No Capacity
This isn’t an argument against quick wins.
It’s an argument for including operating cost in the goal from the beginning.
A realistic AI goal should consider not only:
Can we launch this?
but also:
What will owning this require if it succeeds?
A Proof of Concept Is Not a Small Production System
This distinction becomes particularly important with AI.
A proof of concept exists to test uncertainty.
Can a model extract the required information?
Can a language model answer questions from the company’s documents?
Can historical data predict the outcome well enough to justify further investigation?
A POC can legitimately use shortcuts because its purpose is learning.
Production has a different purpose.
Production must continue working when:
- inputs are incomplete
- users behave unexpectedly
- data changes
- dependencies fail
- traffic increases
- outputs are uncertain
- people rely on the result
Turning a successful POC directly into production often creates the illusion that most of the work is finished.
It may not be.
POC and Production Answer Different Questions
PROOF OF CONCEPT
Can this idea work?
│
▼
Evidence
│
┌──────┴──────┐
│ │
No Yes
│ │
▼ ▼
Stop PRODUCTION
Can we operate
it reliably?
Production isn’t simply a more polished proof of concept.
It’s a commitment to operate the capability.
That commitment should be part of the goal-setting decision.
Define the Next Decision Before Starting
One of the simplest ways to improve AI goals is to decide what will happen when the current goal is reached.
Suppose a team proposes a twelve-week AI pilot.
Before funding it, leadership should know what evidence will lead to:
Scale
The results are strong enough to justify greater investment.
Continue testing
The hypothesis remains plausible, but an important uncertainty hasn’t been resolved.
Change direction
The problem appears valuable, but the proposed AI approach isn’t working.
Stop
The evidence no longer supports further investment.
This prevents pilots from becoming permanent simply because nobody defined what ending one looks like.
A realistic goal isn’t complete until the organisation knows what decision the resulting evidence is supposed to inform.
Set Kill Criteria Before Enthusiasm Takes Over
AI initiatives are easier to start than stop.
Once a team exists, a vendor contract has been signed and leadership has publicly supported the project, evidence starts competing with sunk cost.
That’s why stop criteria are useful before substantial investment begins.
For example:
Stop if representatives need to correct more than 30% of generated responses after the initial evaluation period.
Or:
Do not proceed to production if the workflow doesn’t reduce total review time once human verification is included.
The exact threshold depends on the use case.
The important part is agreeing that some outcomes mean don’t invest more.
Kill criteria turn failure into information.
Without them, every disappointing result can be reframed as a reason for another quarter of work.
Realistic Does Not Mean Easy
There’s another trap in the word realistic.
Organisations can interpret it to mean:
What can we achieve with our existing systems, people and processes?
That produces achievable goals.
It doesn’t necessarily produce valuable ones.
A strategically important AI opportunity may require changing the constraints.
Better data.
Different workflows.
New skills.
More investment.
New governance.
A redesigned operating process.
The correct goal may therefore be difficult.
What makes it realistic isn’t that it fits comfortably inside today’s organisation.
It’s that the organisation understands which capabilities must change, what evidence justifies changing them, and how progress will be measured along the way.
That’s a much stronger definition of realism than simply choosing the easiest AI project available.
The Organisation Has a Learning Rate Too
AI teams often talk about how quickly models learn.
Organisations have learning rates as well.
A technical team can build an experiment in four weeks.
The business may need much longer to change around it.
People need to trust the output.
Managers need to redesign workflows.
Responsibilities need to change.
Policies may need updating.
Teams need to learn when the AI is useful and when human judgment should override it.
Support processes need to accommodate new failure modes.
That organisational learning can become the real constraint.
Suppose an AI system recommends which customers a sales team should contact.
The model performs well.
But representatives continue using their existing lists.
Managers still measure activity using the old process.
The CRM doesn’t integrate the recommendations into the normal workflow.
After six months, adoption remains low.
The problem isn’t necessarily the model.
The organisation hasn’t changed enough for the model to matter.
This is why realistic AI goals need to account for more than technical delivery.
The system and the organisation have to mature together.
Technical Capability
│
▼
AI System
│
▼
Workflow Change
│
▼
Behaviour Change
│
▼
Business Outcome
If workflow and behaviour don’t change, technical capability can increase without producing much value.
Adoption Is Part of the Goal
Teams often treat adoption as something that happens after implementation.
Build the system.
Deploy it.
Then encourage people to use it.
For many AI systems, that separation doesn’t work.
The behaviour of the user is part of the operating model.
A support assistant only creates value if representatives use its suggestions appropriately.
A forecasting model only matters if planners change decisions because of it.
A coding assistant only improves engineering capacity if the time saved exceeds the time spent reviewing and correcting its output.
A realistic goal should therefore include the behaviour required for the business outcome to occur.
Instead of:
Deploy an AI forecasting system.
the goal might include:
Determine whether planners use AI forecasts to make measurably better inventory decisions.
Now adoption isn’t a vanity metric.
It’s part of the hypothesis.
Long-Term AI Goals Usually Depend on Non-AI Work
This is where many AI roadmaps become unrealistic.
The organisation identifies an ambitious future:
Personalised customer experiences.
Autonomous operations.
AI-assisted decision-making across the business.
Predictive maintenance.
Intelligent supply chains.
Then it discovers that the required data is scattered across systems that were never designed to work together.
Customer identities don’t match.
Historical records are incomplete.
Important events aren’t captured.
Definitions differ between departments.
Permissions prevent systems from accessing each other.
None of those problems are particularly exciting compared with building AI.
They may still determine whether the AI strategy succeeds.
A long-term AI goal therefore needs to identify the capabilities it depends on.
Long-Term AI Outcome
▲
│
Workflow Change
▲
│
AI Capability
▲
│
Reliable Data Access
▲
│
Operational Foundations
If the lower layers are weak, the roadmap shouldn’t pretend they don’t exist.
They become part of the strategy.
Infrastructure Should Follow Evidence
The opposite mistake is building every possible foundation before proving any use case.
An organisation decides AI is strategic.
It concludes that strategic AI requires a platform.
The platform requires standardised data.
Standardised data requires new pipelines.
The pipelines require new infrastructure.
Two years later, the organisation has built an impressive technical foundation while the original business use cases remain largely hypothetical.
That’s not long-term thinking.
It’s speculative infrastructure.
The better approach is to allow real use cases to expose which foundations are repeatedly missing.
The first project needs reliable access to customer events.
The second does too.
The third needs the same identity resolution.
Now a shared capability has evidence behind it.
Let Repetition Pull the Platform
Use Case A ───► Data Problem ───┐
│
Use Case B ───► Data Problem ───┼──► Shared Capability
│
Use Case C ───► Data Problem ───┘
This doesn’t mean organisations should never invest ahead of immediate demand.
Some foundations take too long to build reactively.
Security, governance and critical data architecture may need deliberate early investment.
But the reason should be an identified strategic dependency.
Not the assumption that every conceivable AI capability will eventually be required.
Quick Wins Should Expose Strategic Dependencies
This creates another test for choosing short-term AI projects.
A useful quick win can do more than demonstrate business value.
It can expose what the long-term strategy actually requires.
Suppose a company wants highly personalised customer experiences.
Instead of building an enterprise personalisation platform immediately, it tests personalised recommendations in one customer journey.
The experiment may reveal:
- customer identity is inconsistent across channels
- behavioural events arrive too slowly
- product metadata is unreliable
- consent rules limit available data
- the existing application cannot serve recommendations efficiently
Those findings may look like problems.
Strategically, they’re useful evidence.
The organisation now knows which capabilities genuinely constrain the long-term vision.
The short-term project has done two jobs.
It tested whether personalisation creates value.
And it identified what scaling personalisation would require.
That’s a much better quick win than an isolated AI demo that succeeds without teaching the organisation anything reusable.
Avoid Pilot Purgatory
AI portfolios often accumulate pilots.
One team tests a chatbot.
Another tests document extraction.
A third tests forecasting.
A fourth experiments with agents.
Each produces a presentation.
Few become production systems.
The organisation appears active.
Nothing compounds.
This is pilot purgatory.
It usually happens because experiments are easier to approve than production commitments.
A pilot can be funded from an innovation budget.
It can use manually prepared data.
Security exceptions can sometimes be temporary.
A few enthusiastic users can participate.
Production removes those conveniences.
Now the organisation has to answer harder questions.
Who owns it?
What happens when it fails?
How much does it cost?
Who supports it?
How is quality monitored?
Does it integrate with the real workflow?
Is the value large enough to justify all of that?
Those questions aren’t obstacles to innovation.
They’re tests of whether the innovation deserves to become part of the business.
Every Pilot Needs an Exit
A pilot should end in one of four directions.
PILOT
│
Evidence Collected
│
┌────────────┼────────────┐
│ │ │
▼ ▼ ▼
SCALE CHANGE STOP
│
▼
PRODUCTION
Sometimes another experiment is justified because a specific uncertainty remains.
But “continue piloting” should not become the default state.
If the organisation cannot describe what evidence would justify production, the pilot probably began too early.
If it cannot describe what evidence would justify stopping, the pilot may never end.
Manage AI Goals as Two Connected Horizons
Short-term and long-term AI goals don’t need to compete.
They need different jobs.
The near-term horizon should reduce uncertainty and create evidence.
The long-term horizon should define the business capabilities and outcomes worth building toward.
The connection between them matters more than arbitrary dates.
For one company, near-term may mean six weeks.
For another, it may mean nine months.
A regulated financial system and an internal productivity tool should not be forced into identical planning horizons.
The useful distinction is purpose rather than duration.
Horizon One: Evidence
Ask:
- Does the problem matter?
- Can AI materially improve it?
- Will people change their behaviour?
- Does the workflow work?
- Are the economics plausible?
- What new constraints have we discovered?
Horizon Two: Capability and Outcome
Ask:
- What needs to become repeatable?
- Which foundations need investment?
- What operating model is required?
- How does this scale?
- What business outcome should eventually change?
- Which capabilities become strategically reusable?
The relationship looks like this:
LONG-TERM OUTCOME
▲
│
Strategic Capability
▲
│
Repeated Production Use
▲
│
Evidence of Value
▲
│
Near-Term Experiment
Short-term work earns long-term investment by producing evidence.
Long-term direction makes short-term experiments worth choosing.
Don’t Make Every Quick Win Strategic
There’s an important exception.
Not every useful AI project needs to contribute to a grand transformation.
Sometimes an AI tool saves a team hundreds of hours per year.
It is cheap.
Low risk.
Easy to maintain.
It doesn’t build a strategic capability.
That’s fine.
The mistake isn’t pursuing tactical value.
The mistake is pretending tactical value automatically advances the long-term AI strategy.
Keep the categories honest.
Some projects exist because they produce immediate economic value.
Others exist because they test strategic hypotheses.
Others build reusable capabilities.
The portfolio can contain all three.
They simply shouldn’t be measured as though they are the same investment.
Balance the Portfolio, Not Every Project
Trying to make every AI initiative deliver immediate ROI, build long-term infrastructure, develop organisational capability and prove strategic transformation creates impossible goals.
The balance belongs at the portfolio level.
AI PORTFOLIO
┌────────────┼────────────┐
▼ ▼ ▼
Near-Term Strategic Capability
Value Experiments Building
│ │ │
▼ ▼ ▼
Return Evidence Future
Now About Capacity
Direction
The exact allocation will depend on the organisation.
A company under severe financial pressure may need more near-term return.
A company with strong cash flow and a clear strategic opportunity may invest more heavily in capability.
A highly regulated organisation may need substantial governance and data investment before autonomous systems can scale.
There is no universal percentage.
The important thing is that the allocation is deliberate.
Otherwise urgent projects naturally consume everything.
Protect Long-Term Work Explicitly
Long-term AI investment rarely wins an informal competition for resources.
A customer-facing feature has a deadline.
A production issue is urgent.
A sales request has an executive sponsor.
Data quality work can usually wait another week.
Then another.
Then another quarter.
If strategic capability matters, capacity for it has to be protected.
That might mean:
- a dedicated platform team
- explicit roadmap allocation
- separate investment criteria
- executive sponsorship
- defined capability milestones
The mechanism matters less than the recognition that long-term work competes badly against visible short-term demand.
If leadership says data infrastructure is strategically important but repeatedly reallocates everyone working on it whenever a tactical project appears, the real strategy is visible in the allocation.
Not the presentation.
Long-Term Goals Still Need Checkpoints
Protecting long-term work doesn’t mean funding it indefinitely.
Capability programs need evidence too.
Suppose the organisation invests in a shared AI evaluation platform.
The final value may depend on many future AI systems.
But progress can still be measured.
Are teams adopting it?
Does it reduce duplicated evaluation work?
Does it catch regressions?
Does it shorten the time required to approve production systems?
Are the promised use cases actually appearing?
Long-term goals should have intermediate evidence.
Otherwise “strategic investment” can become protection from accountability.
The principle is the same at both horizons:
Every stage should create enough evidence to justify the next stage.
A Practical Framework for Setting Realistic AI Goals
Before approving an AI goal, work through the chain.
1. Define the business problem
What is currently constrained?
Avoid beginning with the AI capability.
2. Define the outcome
What should become measurably different if the initiative succeeds?
3. State the AI hypothesis
Why do you believe AI can influence that outcome?
4. Identify the nearest useful evidence
What can you learn before committing to the full system?
5. Define leading indicators
What behaviour or operational change should appear before the final business outcome?
6. Define the long-term measure
What eventual business result determines whether the initiative was worthwhile?
7. Identify dependencies
What data, infrastructure, workflow, governance or organisational changes are required?
8. Estimate the operating commitment
What will maintaining the system require after launch?
9. Define the next decision
What evidence causes you to scale, continue testing, change direction or stop?
10. Set stop criteria
What result would tell you further investment is no longer justified?
That produces a much more useful goal than:
Deploy AI by Q4.
The Complete Goal Chain
Business Problem
│
▼
Desired Outcome
│
▼
AI Hypothesis
│
▼
Near-Term Evidence
│
▼
Production Capability
│
▼
Behaviour / Workflow
│
▼
Business Outcome
│
▼
Scale / Change / Stop
Final Thoughts
The hardest part of setting realistic AI goals isn’t deciding how ambitious to be.
It’s deciding what evidence should justify the next investment.
Short-term wins and long-term vision aren’t opposites.
A good short-term initiative can test whether the long-term direction is worth pursuing.
It can expose missing data.
Reveal workflow constraints.
Test whether users change their behaviour.
Show whether the economics work.
Identify which shared capabilities are genuinely required.
Long-term strategy gives those experiments direction.
It explains which capabilities are worth compounding and which results are merely tactical.
The mistake is allowing either horizon to dominate completely.
A portfolio made entirely of quick wins can become a collection of disconnected systems consuming more maintenance every year.
A portfolio made entirely of long-term capability can spend years building infrastructure for value that never arrives.
Realistic AI goals connect the two.
They define the business outcome that matters.
They identify the nearest evidence available today.
They make dependencies visible.
They account for what production will cost.
And they specify what the organisation will do when the evidence arrives.
That’s what makes an AI goal realistic.
Not that it’s easy.
Not that it can be completed this quarter.
But that each investment creates enough evidence or capability to justify the next one.





