The headline number from AISI’s pilot study is 25%. Workers using AI produced output that scored 25% higher on quality. Good news. But the 500-person randomised controlled trial also found something more useful for anyone making deployment decisions: AI capabilities are “jagged.” The technology boosted productivity by 102% on information interpretation tasks whilst producing zero measurable improvement in strategic planning.

For UK businesses deciding where to invest in AI tooling, that distinction matters more than the aggregate figure.

Strategic Reality: A 25% average quality improvement masks extreme variation. Some tasks saw 102% productivity gains. Others showed no improvement at all. Your deployment strategy needs to account for this unevenness.

What AISI actually measured

The UK AI Safety Institute ran this study through its AI and the Future of Work Unit and published the results on 2 February 2026. It is one of the first rigorous experimental studies of AI productivity from a government body anywhere in the world.

The methodology was straightforward. 500 participants recruited via Prolific were randomly assigned to two groups: one with access to a large language model released in early 2025, one without. Both groups completed four tasks drawn from the O*NET occupational taxonomy, spanning information input, work output, interaction with others, and mental processes.

TaskDomainQuality gainTime gainPPM gain
Tasks 1 and 2Monitoring and technical specifications22-23%NoneModerate
Task 3Strategic planningNoneNoneNone
Task 4Interpreting informationSignificant42% faster102%

“Points per Minute” (PPM) combined quality and speed into a single productivity measure. Across all tasks, the AI-assisted group achieved 61% higher PPM. But that average is misleading if you are making investment decisions based on it.

Critical Context: This is a pilot study with acknowledged limitations. The sample size creates uncertainty around individual task estimates. The authors describe these as “preliminary indicators” requiring further research. Treat the numbers as directional signals, not benchmarks.

The jagged frontier and why it matters

The researchers describe AI capability as “jagged,” and that word deserves to enter every UK technology director’s working vocabulary. AI does not improve all work equally. It excels at structured, analytical tasks requiring close-ended answers. It struggles with subjective, open-ended work.

This matches what practitioners already see. Ask an LLM to summarise a technical document or extract structured data from an unstructured source and the results are often good. Ask it to develop a strategic plan or make a nuanced judgement call and the output tends toward the generic.

The practical consequence is that organisations deploying AI as a blanket productivity tool will see mixed results. Those mapping AI capability to specific task types will see much stronger returns.

Strategic Insight: The “jagged frontier” means AI deployment should happen at the task level, not the role level. Two people with identical job titles may benefit from AI very differently depending on their actual daily work.

What this means for UK workplaces right now

This study arrives at an interesting moment. UK businesses are past the experimentation phase with generative AI but most haven’t settled into mature deployment patterns. The gap between “we’ve tried it” and “we’ve deployed it systematically” remains wide.

AISI’s findings suggest one reason for that stall: blanket deployments produce underwhelming aggregate results because the gains are concentrated in specific task types. If you measure at the team level, the signal gets averaged away.

StakeholderWhat the data meansPriority action
CTOs and CIOsAggregate ROI calculations understate returns on best-fit tasks and overstate returns on poor-fit onesConduct task-level AI suitability audits
HR directorsJob restructuring should be task-based, not role-basedMap roles into component tasks before assessing AI impact
Operations leadersProductivity measurement must be granularTrack AI impact per task type, not per department
Frontline managersTeams need help identifying which tasks benefitBuild task-specific guidance instead of general “use AI more” directives

Implementation Note: If your organisation measures AI productivity at the department or team level, you are probably averaging out the signal. The AISI data shows that task-level measurement is where the real pattern appears.

A framework for task-level deployment

Based on the AISI findings, UK organisations should consider a three-tier approach:

Tier 1: High-confidence deployment. Tasks involving information interpretation, data extraction, summarisation, and structured analysis. These are where the 42% time savings and 102% PPM gains sit. Deploy AI tools here with confidence and measure results.

Tier 2: Quality-focused deployment. Tasks involving monitoring, technical specification writing, and process documentation. These showed 22-23% quality gains without time savings. Deploy AI as a quality check rather than a speed tool. Set expectations accordingly.

Tier 3: Human-led work. Strategic planning, creative direction, nuanced stakeholder engagement, and judgement-intensive tasks. AI showed no measurable benefit here. Protect human time for this work rather than forcing AI tools into workflows where they add nothing.

Success Factor: The biggest productivity trap is deploying AI into Tier 3 tasks and wondering why ROI is flat. The AISI data gives you permission to be selective.

Four challenges that will slow you down

The measurement problem. Most organisations measure productivity at the team or project level. AISI’s study shows that task-level measurement reveals patterns invisible at higher granularity. Retooling measurement systems takes work, but without it you cannot see where AI actually helps.

The expectations gap. A 25% quality headline creates expectations that don’t match the task-level reality. When strategic planning tasks show no improvement, leadership may conclude that AI “isn’t working” rather than recognising it was never suited to that task type. Setting accurate expectations by task prevents this cycle.

The role redesign problem. If AI dramatically improves some tasks within a role but not others, optimal deployment means restructuring how people spend their time. That is an organisational design challenge, not a technology one. Most IT departments are not set up to drive workforce restructuring.

Hidden Cost: Retraining staff to use AI on the right tasks requires knowing which tasks those are. Most organisations haven’t done the task-level mapping needed to make this work.

The pilot-to-production gap. AISI’s study used controlled conditions with recruited participants. Translating these gains into real workplace settings where people have competing priorities, established workflows, and varying levels of AI literacy is a separate challenge entirely. The 500-person RCT gives you a ceiling estimate, not a floor.

What to do with this

The AISI study’s core contribution is specificity. Not “AI improves productivity” but “AI improves these specific types of tasks by this much, whilst having no measurable effect on others.”

Three actions for UK organisations:

  1. Audit tasks, not roles. Break job functions into component tasks. Categorise them as structured/analytical or open-ended/strategic. This tells you where AI investment will pay off and where it won’t.

  2. Set task-level KPIs. Stop measuring AI ROI at the department level. Measure quality improvement and time savings per task type. The variation AISI found (0% to 102%) means averages are useless for decision-making.

  3. Protect human-led work. The absence of AI benefit in strategic planning tasks is not a failure. It is useful data. Shield strategic and creative work from pressure to adopt AI tools that won’t help, and redirect AI investment toward the structured tasks where gains are real.

Take Action: Start with your five most common repeatable tasks. Run a two-week comparison measuring quality and speed separately for AI-assisted and unassisted work. The AISI data suggests you will find stark differences between task types, and those differences should drive your deployment roadmap.

Source and methodology

This analysis is based on the AI Safety Institute’s blog post “AI and the future of work: Measuring AI-driven productivity gains for workplace tasks”, published 2 February 2026. The study was a randomised controlled trial with 500 participants using the O*NET occupational taxonomy for task selection. AISI characterises the results as preliminary indicators from a pilot study with acknowledged sample size limitations.

Resultsense provides independent analysis of UK AI developments.