Ask Anthropic, Microsoft and OpenAI which work their chatbots do, and you get three different answers. Ask American workers directly, and you get a fourth. That is the finding of a VoxEU column published on 25 September by Alexander Bick, Adam Blandin, David Deming and Tyler Schumacher, who set a US survey designed to be nationally representative against the task shares the three companies have derived from their own conversation data. Across 332 work activities, the survey’s picture correlates at 0.34 with Anthropic’s, 0.10 with Microsoft’s and 0.11 with OpenAI’s. For the UK, this is not an academic quarrel. In 2025 the government signed an agreement with Anthropic that named the company’s Economic Index as a source of “situational awareness” on AI in the economy, and in January 2026 DSIT said the metrics it would use to track AI’s effect on jobs were still being worked out. The choice of ruler is open, and this research says chat logs should not be it.
Strategic Insight: Chat logs record what people type, not who is typing or why. A labour-market measure needs to know the second and third of those things. The authors’ conclusion is that platform data can complement worker surveys but cannot replace them.
Why does it matter which data measures AI at work?
Which jobs AI changes depends on which tasks it performs. The authors set out the logic simply: tasks a technology can take over usually become less valuable, while work that complements it becomes more valuable. So anyone planning skills funding, redundancy support or graduate hiring needs a task-level view of where AI is actually being used, not just where it could be.
Two kinds of evidence have dominated so far. Exposure scores estimate, task by task, what AI could theoretically do. Chat-log studies take sample conversations from a platform, sort each into a task from O*NET (the US labour department’s occupational database) and publish the proportion of chats falling under each task. The column tests three chat-log sources: Anthropic (Handa et al. 2025), OpenAI (Chatterji et al. 2025) and Microsoft (Tomlinson et al. 2025).
The problem is who is missing from the log
The authors’ central objection is structural. A chat log generally sees the conversation but not the worker, so the classifier cannot see the user’s occupation. O*NET, though, describes work by purpose within an occupation. Without the occupation, the classifier often cannot tell what the conversation is for.
There is a second gap. How big a task looks in the chats blends two separate things: how many people perform the task, and how many of those people use AI for it. Because a log contains only people who used AI, it cannot pull those apart. A task can bulk large because it is common, even if few of the people doing it use AI, or because it is rare but most of those doing it use AI.
| Measure | Chat-log data | Worker survey |
|---|---|---|
| Correlation with survey task shares (332 activities) | Anthropic 0.34, Microsoft 0.10, OpenAI 0.11 | n/a |
| Share of all AI use taken by the single largest task | 15% to 23% | 4% |
| Share taken by the ten largest tasks | 46% to 61% | 22% |
| Correlation between the chat sources themselves | 0.08 to 0.38 | n/a |
Strategic Reality: The chat sources do not agree with each other either. Their pairwise correlations run from 0.08 to 0.38, according to the column. When three measures of the same thing disagree this much, the disagreement is itself a finding.
What did the survey actually find?
The survey behind the column is the Real-Time Population Survey, an online survey of American adults built to mirror the Current Population Survey, which has tracked generative AI use since August 2024. Since August 2025 it has asked each worker for their detailed occupation, shown them the ten most important O*NET activities for that occupation, and asked which they perform and which ones AI routinely assists with. Combining four quarterly rounds up to May 2026 gives a sample of almost 14,000 workers between 18 and 64.
Two patterns stand out. Adoption is broad: at least a fifth of workers use AI in their jobs in upwards of 80% of occupations, and for over 40% of tasks, more than 20% of the people doing them use AI. Adoption is also shallow: under 3% of tasks are AI-assisted by a majority of the people performing them, and none by more than 70%.
That shallowness has a sharp consequence. Most of the variation sits between people in the same job. Measured by occupation or by task, exposure scores account for as much as half of the differences in adoption. Measured person by person, they account for under 10%. In the authors’ words, AI use for any given task “depends far more on who performs it than on the task itself”.
Critical Context: This is the finding employers should sit up for. Most of the difference in AI use is between people doing the same work, so an occupation-level average tells you little about your own team.
How one misfiled chat becomes a misfiled job
The column’s clearest illustration involves a professor who asks a chatbot for help spotting trends in a dataset. O*NET treats that as part of the professor’s research into their specialist subject. A classifier that cannot see the occupation will file it under “analyse data to identify trends”, an O*NET task that belongs to data scientists and analysts, not academics. Because studies then infer occupations from tasks, the error travels: the task is wrong, so the occupation is wrong too.
Each chat source’s top task shows the same bias towards generic activity. For OpenAI it is editing written material, for Anthropic designing computer systems or applications, and for Microsoft gathering information. Editing accounts for over 15% of OpenAI’s work chats, yet the occupations O*NET credits with an editing task employ just 2.4% of American workers. Far more people edit, of course; O*NET simply folds editing into bigger purposes such as preparing reports. The survey’s top task, “direct organisational operations, activities, or procedures”, is part of jobs that together employ 43% of US workers, and the authors describe it as “almost invisible in chat data”.
What gets overlooked when chat logs are read as labour-market data:
- Non-users are absent. A log cannot say what share of people doing a task use AI, which is the number a displacement or skills analysis needs.
- Purpose is guessed. Without the user’s occupation, the classifier sees the activity (writing, coding, searching) but not the job it serves.
- Errors compound upwards. Task misclassification becomes occupation misclassification once occupations are inferred from tasks.
- Coarser groupings do not fully fix it. Rolling up into nine broad groups of activity lifts the Anthropic and OpenAI correlations to about 0.6. The authors conclude that some of the mismatch comes from classification noise, “but much of it is not”.
Official statisticians are already using these proxies
This is not a hypothetical risk. The column reports that the US Bureau of Labor Statistics has adopted a pair of chat-based measures for the “observed” part of the AI-exposure groupings behind its jobs projections, while itself cautioning that they “do not directly observe whether workers in a particular occupation used AI on the job”. The authors’ verdict on such proxies is short: they “deserve caution”.
The political direction of travel is towards more data, and the labs want to help supply it. Senators Mark Warner and Ted Budd put forward the Workforce Transparency Act in April 2026. It would require US statistical agencies to gather and publish task-by-task data on AI use, and according to the column, OpenAI, Anthropic, Google and Microsoft all support it.
⚠️ Warning: The BLS wrote the limitation down. Our concern is what happens next: when its categories are reused by others, the caveat does not travel with them automatically.
Where does the UK stand?
The UK has made less of a commitment than the US statistical agencies, but it has pointed the same way. The Memorandum of Understanding between the UK and Anthropic on AI opportunities, signed on 13 February 2025 by Peter Kyle, signing as science secretary, and Anthropic chief executive Dario Amodei, lists “situational awareness” among the areas the two sides intended to explore. The text describes the Economic Index as providing “unique real-world AI model usage data” to inform insight into how AI is being integrated into the economy. The memorandum is voluntary and not legally binding.
A year later, DSIT and the AI Security Institute published an assessment of AI capabilities and the UK labour market on 28 January 2026. It is candid: the available evidence “does not yet provide clear answers to many of the questions that matter most for policy”. It announced an AI and Future of Work Unit to close that gap, which we covered at the time. Its proposed monitoring indicators include individual usage patterns, occupational exposure scores and the “evolution of task content within roles”, with the caveat that “specific metrics are under development”.
The assessment already names the weakness this column quantifies. It states plainly that “exposure is not adoption”, and warns that exposure indices “have not been extensively validated against real-world outcomes”. Its adoption figures, meanwhile, come from the ONS Business Insights and Conditions Survey, which counts firms: around one in five are using or planning to use AI, and within adopting firms fewer than a third of employees use it. Those figures describe firms and the share of their staff using AI, not which tasks that staff use it for.
| Stakeholder | What the chat-log gap means | What they need | How to tell it is working |
|---|---|---|---|
| Government analysts | Task-level indicators built on lab data could misplace AI use across occupations | Worker-level data that includes non-users | Indicators that can be cross-checked against a survey |
| Employers and HR leaders | Sector benchmarks may say little about their own staff | Their own role-by-role picture of use | Adoption measured by role, not by licence count |
| Unions and professional bodies | Their occupation’s apparent exposure may be over- or under-stated | Occupation-anchored evidence to argue from | Claims that survive a change of data source |
| AI providers | Their data is increasingly used as an exposure proxy | A shared task taxonomy with official statistics | Usage data published with its limitations |
🎯 Success Factor: The test for any UK indicator is whether it can see people who never use AI. A measure built only from AI users cannot give an adoption rate, and adoption rates are what skills and transition policy are built on.
What should UK organisations and policymakers do with this?
The authors do not argue for throwing chat data away. Their own survey has limits: it asks only about each occupation’s top ten tasks and is too thin to say much about small occupations. Chat logs are strong where the survey is weak, observing use at scale, and weak where it is strong. The authors call the two “complements rather than substitutes” and argue both would benefit from a shared taxonomy.
💡 Implementation Framework: Anchor, then enrich
Phase 1: Anchor in people (next quarter)
- Measure AI use by role and task among your own staff, including non-users
- Record occupation or job family alongside every usage figure
- Treat vendor usage dashboards as activity counts, not adoption rates
Phase 2: Cross-check (six months)
- Compare your role-level picture with published exposure scores and lab task shares
- Investigate where they disagree rather than averaging them
- Note which tasks your data shows that generic chat categories would hide
Phase 3: Enrich (this year)
- Add platform or telemetry data only where it can be linked to a role
- Map internal task lists to a recognised taxonomy so results are comparable
- Revisit annually, since adoption within a role varies widely
Priority actions by starting point
For organisations just starting
- Ask before you infer: A short staff survey on which tasks AI helps with, by role, gives you what no vendor log can: the share of people who never use it.
- Distrust generic top tasks: If a report says writing or information-gathering is the main AI use, ask what those activities serve in your organisation.
- Keep job context in every figure: A usage number without a role attached cannot inform workforce planning.
For organisations already underway
- Separate breadth from depth: Track how many people in each role use AI, not just total sessions or prompts.
- Look within roles: The survey found that most variation sits between people in the same job, so compare individuals within a team before comparing teams.
- Challenge benchmark claims: When a supplier cites lab usage data to size your opportunity, ask how the underlying chats were assigned to occupations.
For advanced implementations
- Build a shared taxonomy: Map internal tasks to a public framework so your data can be compared with official statistics as they arrive.
- Link telemetry to roles: Where privacy rules allow, attach job family to usage data so it can be read the way the survey reads its respondents.
- Feed the evidence base: Share anonymised role-level findings with sector bodies building the UK’s picture.
Resource Reality: A role-based AI use survey does not need to be elaborate. A short questionnaire mapped to existing job families, repeated quarterly, captures the two things lab data cannot: who each respondent is, and whether they use AI at all.
The hidden challenges
Challenge 1: The data that is easiest to get is the data with the widest reach
Lab usage data is published by the labs themselves, from logs they already hold. The survey in this column, by contrast, needed outside funding, from the Walmart Foundation among others. The risk is that convenience decides which evidence frames policy, not quality.
Mitigation Strategy: Use platform data for what it does well, such as spotting new activities early and at scale, and pair it with a representative survey before drawing occupational conclusions. Where only chat data exists, say so when reporting results.
Challenge 2: Averages hide the worker-level story
According to the column, exposure scores explain under a tenth of the person-to-person differences in adoption. Any policy or HR decision based on occupation-level exposure will miss most of what differs between people.
Mitigation Strategy: Design training and redeployment around what individuals actually do with AI, gathered directly, rather than around job titles.
Challenge 3: Different sources, different winners
The four sources name four different top tasks. An organisation that picks one dataset to justify an investment may simply be picking the answer it wanted.
Mitigation Strategy: Before adopting any external benchmark, test whether its conclusions change when you switch data source. If they do, treat the finding as unproven.
Challenge 4: Non-public methods are hard to scrutinise
The column notes that Google’s own approach, which infers the occupation first and then assigns a task, runs the other way round from the chat studies it tested, and that Google has not made its task-level results public. Where data cannot be examined, their errors cannot be measured either.
Mitigation Strategy: Give more weight to evidence whose data and methods are open. The survey data behind this column can be downloaded, and that openness is part of its value.
Reality Check: None of this makes lab data worthless. It makes it one input among several, with a known blind spot about who is using AI and why. The problem arises only when it is treated as the whole picture.
The strategic takeaway
The UK government has said, in its own words, that it does not yet have clear answers on AI and jobs, and that its metrics are still being developed. That is the moment to choose the right foundation. This column’s evidence is that chat logs, however large, cannot see the worker, and a labour-market measure that cannot see the worker will misplace AI use across occupations. Survey data anchored in people, non-users included, should be the base. Platform data can then add scale and speed on top. For background on the lab data itself, see our analysis of Anthropic’s January 2026 Economic Index.
Three things that decide whether the measurement works
- Worker anchoring: Every figure should be traceable to a person with a known occupation.
- Non-users counted: Adoption rates need a denominator that chat logs do not have.
- Cross-checking by design: Survey and platform sources should share a taxonomy so each can test the other.
Measuring use, not conversation volume
The useful question for a UK employer is not how many conversations a tool handled, but which of its people use AI for which parts of their jobs, and which do not. Those two numbers move independently. A rising session count can coexist with most of a team never touching the tool.
Strategic Insight: Conversation volume measures a product’s popularity. Adoption by role measures change in work. Only the second tells you anything about jobs.
Your next steps
Immediate actions (this week):
- List which AI usage figures your organisation currently reports, and where each comes from
- Check whether any of them can distinguish users from non-users
- Flag any external benchmark built on lab chat data in current business cases
Strategic priorities (this quarter):
- Run a short role-by-role survey of AI use across staff
- Attach job family to any internal usage telemetry
- Compare your results with published exposure scores and note disagreements
Long-term considerations (this year):
- Map internal tasks to a public occupational framework
- Watch for the UK’s published labour-market indicators and the data behind them
- Track whether US task-level AI data collection legislation progresses
Source: Measuring what work generative AI does: Survey evidence versus chat logs (CEPR VoxEU, 25 September 2026), by Alexander Bick, Adam Blandin, David Deming and Tyler Schumacher, drawing on the authors’ CEPR Discussion Paper 21889, “What work does generative AI do?”. The authors note the views expressed are their own rather than those of the St. Louis Fed or the wider Federal Reserve System. UK context draws on DSIT’s January 2026 labour-market assessment and the February 2025 UK–Anthropic memorandum, both linked above.
This strategic analysis was written by Resultsense, a UK-focused AI news and analysis publication. We will be watching which data the UK’s labour-market monitoring is built on when its first indicators are published. Read more analysis at Insights, or get in touch.