High weekly engagement: KPMG reported strong weekly AI adoption in its US advisory division after rolling out an internal usage dashboard. KPMG used that adoption data to link people-level AI activity to productivity and business outcomes, a rare example of measurement that ties usage to revenue-relevant results. Universities face a different pressure, with the Research Excellence Framework influencing roughly £2 billion of annual research funding, so tracking AI without harming income means measuring outcomes not raw counts. This guide explains a six-step approach organisations can use to instrument AI use, protect funding or billable work, and publish governance that stands up to scrutiny.
Why measurement matters
Organisations measure AI for two blunt reasons: to capture value, and to manage risk. Employers want higher-quality billable work, faster delivery and lower rework. Universities want faster, more accurate research administration, stronger partner matches and demonstrable impact in assessment exercises that determine public funding. The UK government’s AI for Science Strategy, published in 2025, set aside £137 million to strengthen public-private research partnerships, reinforcing pressure on higher education to show the value of AI investments.
Counting every prompt or session won't do the job. KPMG’s internal dashboard, which it reported reached weekly AI engagement by more than 90 percent of staff in its US advisory division, pairs adoption figures with productivity and outcome signals. Russ Grote, a KPMG spokesperson, said the exercise is intended to improve quality of work and free staff to focus on higher-value tasks. The difference between measuring volume and measuring value is why the question that actually decides this isn't whether staff used AI, but whether that use increased billable or funding-related value.
1. Start with clear objectives and revenue-linked KPIs
First, write a short definition of what tracking must show. For a professional services firm that will usually mean higher-quality billable work, faster turnaround on client deliverables, or lower rework rates. For a university the immediate goals will be faster and more accurate research administration, improved partner acquisition, or demonstrable research impact in REF submissions.
Translate those objectives into KPIs that directly connect AI activity to money or funding outcomes. Useful KPIs include changes in billable hours per employee, reduction in time to submit grant applications, improvements in partner conversion rates from CRM activity, or higher scores on REF submission metrics. Use public numbers to set the scale of importance: REF influences about £2 billion a year of research funding; REF2021 cost around £471 million in total and about £3 million on average per participating higher education institution. Those figures make clear why precise, defensible reporting matters for institutional budgets and reputations.
Worked example: if a consultancy wants to free up ten hours a week per partner through AI, convert that into additional billable capacity and a margin figure. If a university aims to cut administrative time on a REF submission by 20 percent, express that as staff days saved and as the fraction of annual research administration costs that will be affected.
2. Map AI touchpoints and instrument data collection
Second, catalogue every place AI appears in paid work and research workflows. That includes CRM integrations that suggest partners, generative models used to draft narratives for funders, developer tools used by technical staff, and automation embedded in research-administration systems. Decide the level of granularity required: tool-level logs, prompt metadata, session counts, token volumes or outcome-level indicators such as time saved per task.
For universities, instrument AI telemetry inside CRM and research-management systems so partner engagement and REF-related evidence are measurable across the lifecycle of a grant or submission. Engage IT, legal and privacy teams early to make sure logs are collected lawfully and in line with data-protection obligations. The University of Bristol-led study funded by Research England found that generative AI is already being used quietly in some REF processes to gather evidence, score outputs and prepare narratives, which is precisely the sort of activity you should be able to trace.
Worked example: tag every CRM record created or updated by an AI prompt, and link that tag to subsequent partner meetings and contract wins. For REF-related drafts, record whether a draft was AI-assisted and capture the time taken for human review so you can show whether AI reduced administrative cost without altering scholarly judgement.
3. Measure outcomes not just volumes
Third, build dashboards that correlate AI usage with the revenue- or funding-linked KPIs you set in step one. Frequency of use and token volumes are useful context, but they're auxiliary signals, not proof of value. Where possible measure downstream outcomes: changes in billable hours per employee, client satisfaction scores, faster completion of grant applications, or improved scoring of research outputs.
Combine quantitative telemetry with qualitative review. Peer assessments, spot audits and outcome validations help detect automated inflation of activity numbers. The Bristol-led study warned that uneven access to AI tools and in-house capabilities can create competitive disparities between better-resourced institutions and smaller ones when automated processes are applied to REF submission work. That's why pairing telemetry with human review matters: it shows whether AI helped a genuine outcome or merely accelerated paperwork.
Here's the thing, worked example: rather than report that a team processed 3,000 AI prompts in a quarter, report that the team reduced average client turnaround from five days to three days, and show how that change translated into additional billable capacity.
4. Pilot selectively and validate the signals
Fourth, run a time-boxed pilot on teams that represent the broader organisation. Make the pilot long enough to observe outcome correlations, for example a quarter of billing cycles in client services or one full REF-preparation cycle in a university. Use the pilot to test whether your chosen metrics capture meaningful change and whether staff can game them.
Adjust the metric set to eliminate noisy or manipulable indicators and replace them with measures that reflect real value. KPMG’s approach, which pairs usage metrics with programmes that reward demonstrable innovation and outcome-producing behaviour, is an instructive model: it shows the need to validate metrics against the things the organisation actually sells or reports to funders.
Worked example: pilot a dashboard that logs AI-assisted draft creation plus reviewer time. If reviewers stop spending time on oversight, that's an early warning that quality may have slipped. If reviewer time declines while client satisfaction and error rates hold steady or improve, the pilot has demonstrated a positive signal.
Fifth, design incentives to reward validated outcomes rather than raw adoption. Avoid bonus schemes or recognition that prize prompt counts or session totals. Instead tie rewards, career progression or bonus pools to improvements in validated KPIs: higher invoiceable hours, faster grant throughput, higher partner conversion rates, or demonstrable improvements in REF submission quality.
If your organisation sells time-based services, ensure productivity gains are converted into sellable capacity or improved margins. For universities, govern the use of AI in REF evidence preparation so automation doesn't create inconsistencies that open institutions to funding risk. The Bristol-led study emphasised scepticism about unsupervised generative-AI use in assessment contexts and called for oversight. That scepticism reflects a simple truth: automation that isn't auditable is a liability when funding panels or auditors ask for explanation.
Worked example: rather than pay a bonus when a team records a high number of AI sessions, pay the bonus when audit trails show AI-assisted work reduced time-to-delivery and the client satisfaction score didn't fall.
Sixth, make governance public and consistent. Publish rules that describe what telemetry you collect, how results are used, how audits are performed and how disparities in tool access are addressed. For universities, bake those protocols into REF-preparation governance so automated assistance is traceable and defensible when funders or panels request clarification.
The University of Bristol-led study, funded by Research England, highlighted widespread scepticism about unsupervised generative-AI use in assessment contexts and recommended national oversight. Institutions should respond by publishing the logic of their measurement programmes so panels, funders and external auditors can interrogate automated workflows and confirm that AI improved administration without changing substantive judgments.
Worked example: publish a one-page protocol that lists the telemetry fields collected for REF-related tasks, the retention period for logs, the human-review steps required before a draft is submitted, and the audit process for disputed submissions. Make that protocol available to funders and internal auditors.
Implementing measurement is costly and technically tricky. Measurement systems commonly miss bespoke developer tools and siloed in-house systems. Dashboards that rely on single indicators are easy to manipulate. Prompt counts and token volumes can be gamed by automated scripts and don't capture developer tool usage. Treat these signals as supportive, not definitive.
Scaling measurement across a large workforce requires investment in instrumentation and data engineering. The cost of measurement should be weighed against likely savings or revenue uplift. The Bristol-led study also found variation in how institutions use generative AI for REF work, with better-resourced higher education institutions more able to develop in-house tools that streamline submission effort. That disparity raises governance and equity questions: if only a subset of institutions can produce machine-assisted REF submissions at scale, funders and sector bodies must consider how to evaluate submissions fairly.
Technical caution: involve legal and privacy teams from the start. Logs and telemetry can contain personal data. Collect and retain only what's necessary to validate outcomes and comply with data-protection obligations.
Implementation checklist
First, set a short objective statement and link it to at least one revenue or funding KPI. Second, map every AI touchpoint and decide what telemetry to collect. Third, build dashboards that show outcome correlations not just volumes. Fourth, run a pilot long enough to validate the signals. Fifth, align incentives to validated outcomes. Sixth, publish governance protocols and audit paths.
In Short
• Measure outcomes, not only prompts or tokens.
• Convert productivity gains into sellable capacity or defendable funding claims.
• Pilot before you scale and validate against billable or REF-related KPIs.
• Publish governance so audits and funders can inspect automated assistance.
• Treat prompt and token counts as auxiliary signals, not proof of value.
Related Articles
- Deploy CrewAI agent teams on AWS Bedrock: 8 steps
- Open a UK business bank account fast: 6 steps
- 3 AI IDEs Compared: Choose by Autonomy, Models, Security
Organisations should use the time before their next funding or assessment cycle to run a short, time-boxed pilot that pairs an instrumentation dashboard with revenue- or funding-linked KPIs and a published governance protocol, and scale only after the pilot shows tracked AI use delivers measurable billable or funding-related value.
This article was created with AI assistance.