Here is the kind of message I see in a strategy call. A winery GM forwards a screenshot from her admin console. Seventeen of twenty-two licenses showed activity in the last thirty days. Seventy-seven percent. She wants to know whether that is a good number. The screenshot tells her almost nothing about whether her AI rollout is working, and it may quietly be telling her the rollout is failing while she nods at the dashboard.
Usage rate as a percentage of licenses active in the last thirty days is the metric every SaaS admin console makes easy to read. It is also a vanity metric that hides the question a winery owner is paying to have answered. A staffer who logged in once, typed "summarize this email," closed the tab, and never came back counts the same as a staffer who uses the tool eight times a day on real work. The dashboard rolls them up to a single percentage. The percentage looks like adoption.
I argued in five things wine industry leaders should stop doing in AI rollouts that the all-hands rollout is one of those stop-doing moves. This post takes the metric layer of the same problem. If you are tracking license activation as your scorecard, you are measuring the same surface the all-hands rollout produces — visible compliance, no behavior change underneath.
What the usage-rate metric counts
The metric an admin console gives you for free counts logins. A login is a free action. A staffer can produce a login by accident, by curiosity, by manager pressure, or by once-a-month script. None of those produce the change in how work gets done that you bought the AI subscription to produce.
Microsoft's Work Trend Index research has a more useful threshold. The habit-formation point for AI tool use is three uses per week sustained for seven to eight weeks. Below that, employees bounce off and revert to the previous workflow. Above it, the use turns into something the staffer reaches for unprompted. The interesting number on a usage dashboard is not how many licenses showed activity this month. It is how many staffers hit three uses per week for eight weeks running.
Most admin consoles do not show you that number. They show you the easy number. The easy number is what a vendor wants you to celebrate because it makes the renewal conversation friction-free.
The three-tier model that survives reality
A model worth using has three tiers. Inputs, behavior, and outcomes. Most winery rollouts I see only track the input tier.
Inputs are the things the budget purchased. Licenses provisioned. Training hours delivered. Starter prompts written. Slack channels created. They are useful for confirming the spend happened. They are not useful for telling you whether anything changed. A winery can have one hundred percent of licenses provisioned, every staffer through a training, a fully populated prompt library, and zero shift in the Monday-morning workflow. The input tier is necessary and insufficient.
Behavior is what the staffer does on a Tuesday at 2 p.m. when there is a real task on her screen and no manager looking over her shoulder. Did she open Claude. Did she paste in the half-written email she would otherwise have spent twenty minutes on. Did she come back to it the next day for a different task. The Microsoft 3x/week benchmark lives in this tier. I wrote about that habit-formation number in more depth in the 3x-a-week-for-8-weeks rule — it is the single most useful behavior metric I have seen and I build sixty-day plays backwards from it.
Outcomes are the second-order effects. Time reclaimed on a specific task. Errors caught. A coworker walking over to the staffer to ask how she did the thing. Job satisfaction. The outcome tier is where ROI lives, and it is the slowest to read out. Slack and Salesforce's Workforce Index ran a survey of over five thousand workers in April and May of 2025 and found that daily AI users reported an 81% lift in job satisfaction. The satisfaction number is the most underrated outcome metric in this whole space, because it is the one that predicts whether the staffer keeps using the tool in month three when novelty has worn off.
If a winery only tracks the input tier, every rollout will look like a success until the day someone audits whether anything is different. If it tracks the behavior tier, the leader knows two months in whether the play is going to take. If it tracks all three, she has a model that survives past the honeymoon.
Day 30
At day thirty, the question is binary. Did the staffer cross the line from "I have a license" to "I have opened the tool on a real task more than once."
A useful day-thirty check is two questions, asked in a hallway conversation, not in a form.
Question one: "What is one thing you tried with Claude in the last two weeks that you would not have tried before?" If the answer is a specific task — a draft email she would have written from scratch, or a customer note pulled together from her morning — the staffer is past the activation cliff. If the answer is "I have not had a chance," she is not.
Question two: "Was anything in the result surprising or useful?" The word surprising is doing work here. If the staffer can name a specific moment the tool did something she did not expect, the curiosity loop has started and the second use is much more likely. If she shrugs, the rollout is in the failure cohort and you have thirty more days to turn it before the staffer files the tool under "I tried it, it did not do anything for me."
What you are not measuring at day thirty is volume. Volume too early is performative — it is the staffer who logs in to look good on the dashboard. Engagement at depth on one task beats ten shallow opens.
Day 60
At day sixty, the question moves from binary to behavioral. Is the staffer in the 3x/week habit yet, and can she point to one task she has reshaped around the tool.
The way to read this is a fifteen-minute check-in between the GM (or whoever owns the rollout) and the staffer, not a telemetry dashboard. The check-in covers three things. How many days last week did she reach for the tool. What was the most useful use, and roughly how much time did it save. Has anyone else asked her how she did the thing she did.
That third question is the leading indicator nobody puts on a dashboard. Klarna's published case study on its AI rollout pulled out the same pattern — staffers asking each other for setups was a stronger signal than any platform metric. When a tasting room manager turns around to the DTC coordinator and asks how she pulled that summary together, the rollout is starting to compound. When nobody is asking anyone anything, the rollout is a chain of isolated logins.
The Klarna deep-interview pattern — a fifteen-minute one-on-one with each AI-using staffer, quarterly — is the closest thing to a sane scorecard at this stage. It costs about twenty minutes of leader time per staffer per quarter. For a thirty-person winery with eight active AI users, that is under three hours a quarter of the GM's time. The trade is information about how the tool is showing up in the work, instead of a dashboard percentage that hides whatever is going on.
Day 90
At day ninety, you are looking at outcomes. Three of them.
First: self-reported hours reclaimed. Not "do you feel more productive" — too vague to trust. Instead: "Pick one recurring task you used to do without AI and now do with it. How long did the old version take, and how long does the new version take." Even rough numbers are useful here, because the rough number is what the leader uses to decide whether to put a second staffer on the same play.
Second: the shareable artifact. Has the staffer produced something — a starter folder or a refined prompt — that another staffer could pick up and use without a training session. The shareable artifact is the test of whether the staffer's knowledge is now portable. If everything she has built lives in her head, the rollout has not yet produced organizational learning.
Third: the downstream coworker. Has a second staffer asked to be set up with what the first one is doing. If yes, the rollout is starting to grow on its own and the leader's next move is to clear time for the second setup. If no, the play needs a structural change — usually around what gets shown publicly, since most staffers will not ask for help they have never seen anyone else ask for.
The job satisfaction read fits here too. The Slack/Salesforce 81% lift is a population-level number, but you can read the same signal at one staffer's desk. Has the tone of how she talks about her workload changed. Is she taking on tasks she previously deflected. These are observations a leader who is in the room every week can make and write down — they will never show up on a dashboard.
What to take off the scorecard
A few metrics I see on winery rollout plans that I would cut.
License utilization percentage. The number a vendor wants you to look at. Hides everything underneath the login count.
Total prompts run org-wide. A staffer running fifty test prompts to look engaged produces the same number as a staffer running fifty real-work prompts. Total volume without context is theatre.
Number of training sessions completed. A training session attended is a thing that happened on the calendar, not a behavior change in the work. Attendance is an input tier metric and it is almost always 100% on the day and almost always disconnected from month-three behavior. I argued the structural version of this in why all-hands AI training day kills adoption.
Self-rated "comfort with AI" survey scores. They drift up after every training day, then drift back down by month two, then up again after the next training. The score is measuring how recently the staffer was in a training, not how AI is showing up in her work.
Sentiment dashboards. Anonymous "how do you feel about AI" pulses produce data that looks like it should mean something and almost never does. The fifteen-minute one-on-one beats the dashboard pulse by a factor I would not even try to put a number on.
If your AI rollout report this month is mostly the metrics above, the report is making your rollout look healthier than it is.
One sentence to put on the rollout doc
If I were going to put a single sentence on the front of a winery's AI rollout doc to anchor the measurement question, it would be a swap of the question. Drop "how many staffers used AI this month" and replace it with "for which staffers, on which tasks, has the way the work gets done changed."
That sentence is harder to dashboard. It is also the one that maps to whether the work is paying off.
A note on the webinar
The second half of the Employee AI Engagement webinar walks through the sixty-day play that produces the day-30, day-60, day-90 reads above. The hour is built around small wineries and small SMBs — under one hundred people — and the play is the one I run with consulting clients when we are setting up a rollout we expect to still be running in month four. If your winery already has AI licenses out and you are not sure how to read whether they are working, the webinar is built for that exact situation.
Part of a 20-post series on employee AI adoption for wineries — see the full series under AI Adoption.