An AI readiness assessment scores a business across six dimensions before it spends money on an AI project: process definition, data quality in the fields the task reads, ownership, approval and risk controls, a measurement baseline, and team capacity to review output. Each dimension scores 0 to 3, for a total out of 18. A low total does not mean AI is off the table. It means the highest-scoring, narrowest task is where to start, and the lowest-scoring dimension is what to fix first.
Why score readiness before choosing a tool
Most AI adoption challenges do not show up in the software. They show up in the gap between what a vendor demo assumes and what the business actually has: a documented process, clean data, a person accountable for the result, a way to stop something that goes wrong, and a number that proves whether it worked. An assessment surfaces that gap before money moves, not after.
This is a scored self-assessment, not a vendor questionnaire. A leadership team runs it about one candidate task, not the whole business, because readiness is task-specific. A company can be highly ready for automating a weekly report and completely unready for automating customer replies, in the same week, with the same team.
Run it in a room with the people who actually touch the task: the manager, the person doing the work today, and whoever owns the number the task is meant to move. An assessment done by leadership alone, without the person doing the work, tends to score too high, because it scores the process as written rather than the process as run.
The six dimensions
Score the candidate task against each dimension from 0 to 3. Zero means the dimension is effectively absent. Three means it would hold up under a skeptical outside review. Be specific to the one task under review, not to the business in general.
- Process definition. Can you write the task's inputs, steps and output on one page, and would two people following it produce the same result? Score 0 if it lives only in one person's head. Score 3 if it is written down, current and followed consistently today.
- Data quality in the fields that matter. Pull 50 records and check only the fields this specific task reads. Score 0 if a third or more are wrong, missing or stale. Score 3 if a sample check finds the fields reliably populated and current.
- Ownership. Is one named person accountable for this task's outcome, with authority to review output and turn it off? Score 0 if it is a shared inbox or nobody. Score 3 if one person is named, reviewing, and empowered to stop it.
- Approval and risk controls. Does anything the task produces reach a customer, change a record of value, or spend money without a human checking it first? Score 0 if it goes straight out. Score 3 if every external action waits for a specific, named approval.
- Measurement baseline. Do you know, from running the task by hand, how long it takes, what it costs and how often it goes wrong today? Score 0 if nobody has measured it. Score 3 if you have a recorded baseline from the last month.
- Team capacity to review output. Does someone have the time, each week, to check a sample of the AI's output against the source and catch drift? Score 0 if the team is already at capacity with no slack. Score 3 if reviewing is already built into someone's week.
Scoring the total: what each band means
Add the six scores for a total out of 18. The bands below describe what the total tells you about this specific task, not about the business as a whole. A different task, scored separately, can land in a different band the same afternoon.
- 0 to 6, not ready. At least half the dimensions are close to absent. Do not pilot yet. Pick the lowest single score and fix that one thing before rescoring. Attempting a pilot here usually produces the failure pattern where AI makes an undefined process fail faster and more visibly.
- 7 to 11, ready for a narrow pilot. The basics exist but at least one dimension is weak, most often data quality or a measurement baseline. Run the pilot on the narrowest version of the task, fix the weak dimension in parallel, and keep the approval gate on every output for the first cycle.
- 12 to 15, ready to run with normal governance. Most dimensions are solid. Pilot with a standard approval gate, move to sampling once the pilot proves out over a full measurement cycle, and rescore before widening to a second task.
- 16 to 18, ready to widen carefully. The task is well understood, well owned and well measured. This is the profile where a second and third bounded task can be added without a full-scale program, as long as each new task gets its own score rather than inheriting the first one's.
The dimension that gets skipped: measurement and ROI
Of the six, measurement baseline is the one leadership teams most often score generously without evidence. It feels like a formality next to data quality or ownership. It is not, because it is the dimension that determines whether AI ROI is a number or an opinion.
A task without a baseline cannot show a return, because there is nothing to compare the result against. If nobody timed how long the weekly report took by hand, a claim that AI saved four hours a week is a guess dressed as a metric. Worked example: a team estimates the report takes about a day to prepare. When someone actually times it for two weeks, it is 90 minutes on a normal week and closer to four hours the week a client review is due. Without that measured range, any post-implementation number is being compared to a guess, and the guess was wrong in both directions depending on the week.
Score this dimension honestly. If the answer to how long does this take today is an estimate rather than a timed number from the last few weeks, score it 0 or 1 and spend a week measuring before scoring again.
An AI adoption framework built on the score
The score is only useful if it changes what happens next. Use it as the front end of a simple AI adoption framework: score the candidate task, fix the lowest dimension, rescore, pilot at the resulting band's level of governance, measure against the baseline, then move to the next task with a fresh score rather than assuming readiness carries over.
This keeps AI adoption from becoming one large program with one large risk profile. Instead it becomes a series of small, scored decisions, each one bounded to a task the team already understands well enough to score honestly. A business that runs this discipline on three tasks over two quarters will usually outperform one that runs a single large rollout in a month, because every step after the first is informed by a real measured result instead of a demo.
What low scores are actually telling you
A low total is diagnostic, not disqualifying. Process definition scoring 0 means the fix is writing the page down, which usually takes a day, not a quarter. Data quality scoring 0 means sampling and cleaning the specific fields the task reads, which is bounded once you know exactly which fields those are. Ownership scoring 0 means naming a person, which costs nothing but the decision to do it.
The dimension that takes longest to fix is usually approval and risk controls, because it requires agreement on what counts as an external action and who has authority to approve it. Start that conversation early, even while fixing the faster dimensions, so it is not the last blocker before a pilot that is otherwise ready.
Questions leaders ask
How long does an AI readiness assessment actually take?
About one hour for one candidate task, with the right people in the room: the manager, the person doing the work, and whoever owns the metric the task touches. Scoring six dimensions from 0 to 3 with a short discussion on each fits comfortably inside an hour once you have picked the task in advance.
Should we assess the whole business or one task at a time?
One task at a time. Readiness is not a company-wide property. A business can score 15 on drafting follow-up emails and 3 on automated customer replies in the same week, because the dimensions, especially approval and risk controls, depend on what the specific task touches.
What is the single biggest AI adoption challenge this assessment catches?
An optimistic measurement baseline. Teams routinely believe they know how long a task takes and how often it goes wrong, and are off by a wide margin once someone times it. The assessment forces that number to be measured rather than estimated before any AI is layered on top of it.
Do we need a perfect score before piloting anything?
No. A score of 7 to 11 supports a narrow pilot with a tight approval gate while you fix the weakest dimension in parallel. Waiting for a perfect score delays value with no added safety, since the pilot's own gate is what catches problems while the weak dimension gets fixed.
How does this connect to a broader AI implementation roadmap?
The assessment picks and scopes the task. The roadmap is the phased plan that documents the process, cleans the data, assigns the owner, runs the pilot behind the gate, and measures the result. Run the assessment first on each candidate task, then take the one with the strongest score into the roadmap.
Start a Diagnostic