Does generative AI improve employee productivity?
Bottom Line
What can reasonably be concluded from the available evidence?
Generative AI can improve employee productivity in specific settings, with observed gains in customer support and selected software-development tasks (S10, S20, S24). Benefits vary by task and employee experience; survey associations between greater use and higher productivity do not establish causation (S3, S4). The evidence does not establish universal or sustained organization-wide gains, and verification and coordination burdens may reduce task-level benefits (S21).
Key Claims
Evidence Strength: how strongly is the claim supported? Applicability: how relevant is the evidence to this question? Assessed separately.
Consensus & Divergence
Consensus
- Evidence converges on productivity benefits in selected customer-support and software-development tasks, rather than across all occupations (S10, S20, S24).
- Routine and structured software tasks show clearer efficiency gains than complex tasks; survey respondents also report benefits in drafting, retrieval and summarisation (S4, S20, S24).
- Greater frequency or intensity of use is positively associated with productivity in distinct survey samples, without establishing causal direction (S3, S4).
- Employee experience matters: novices and lower-skilled workers benefited most in customer support, while experienced and highly skilled workers saw minimal gains; generalization across occupations is not established (S10).
- Assistance configuration matters: in-code suggestions and chat-based prompting each improved efficiency, but combining them within a task diminished benefits (S20).
- Net benefits depend on operational integration and the burdens of verification, orchestration and resource management (S12, S21).
Divergence
- Positive task-level findings contrast with S21's account of overhead, slower experienced developers and unclear organizational delivery improvements. These findings concern different outcomes and settings, so they do not establish a direct contradiction (S10, S21, S24).
- Benefits are not uniform across outcomes: complex software tasks show more modest time savings but improved task success, qualifying any simple conclusion that complex work gains little (S24).
- Study designs differ: S3 and S4 use surveys, S10 examines a staggered introduction, S20 combines controlled sessions with natural work periods, and S24 reports an experiment. These designs support different causal interpretations; random allocation in S10 and S24 is not reported.
- Productivity measures differ: S10 measures issues resolved per hour, S24 measures task completion time and success, S4 measures self-reported productivity, and S21 discusses organizational delivery. Improvement in one measure need not imply improvement in another.
- Populations and tasks differ: S10 studies customer-support workers with differing experience, while S20 and S24 study software-development work with differing task demands.
- Evidence independence and completeness limit synthesis: S5 duplicates S4, while S1, S2, S7 and S13 provide no reported study findings in the supplied text. S12's underlying synthesis evidence cannot be independently assessed from that text.
What We Know / What We Don't Know
What We Know
- Access to a conversational assistant increased issues resolved per hour in the studied customer-support deployment (S10).
- Generative AI reduced completion time for selected software tasks, with larger time savings for routine work and improved success on high-complexity tasks (S24).
- More intensive use is associated with higher productivity in surveys, including self-reported productivity, but these associations do not establish that greater use causes improvement (S3, S4).
- Tool combinations can diminish efficiency benefits, and verification and coordination burdens are relevant cautions when interpreting net productivity (S20, S21).
What We Don't Know
- Whether observed gains persist over the longer term or translate into sustained organization-wide productivity improvements.
- How reliably findings generalize to occupations and organizations outside the studied settings.
- How frequently verification and coordination burdens fully offset task-level gains; S21 does not establish this through a reported multi-organization causal evaluation.
- Whether broad work-quality improvements or larger benefits for novices hold across occupations; the supplied objective quality and cross-occupation evidence is limited (S6, S8, S10, S24).
Strategic Implications
What could this evidence mean for decision-making?
Evidence Suggests
The evidence supports a task-specific productivity opportunity, particularly in customer support and routine or structured software work, rather than a blanket claim that adoption improves every employee's performance (S10, S20, S24).
Important Conditions
The supported interpretation depends on task demands, employee experience, assistance configuration and workflow integration. Faster task completion, output quality and organizational delivery are distinct outcomes, and overhead may change the net result (S10, S12, S20, S21, S24).
Decision Uncertainty
The principal uncertainty is whether gains in a particular deployment remain after verification and coordination costs and persist at organizational scale. The supplied evidence supports benefits in selected settings but does not resolve that broader judgment (S10, S20, S21, S24).
StratifyLens structures the evidence. The decision remains yours.
Underlying Evidence
16 sources used · 34 retrieved
Where this comes from
How StratifyLens researched this question
Follow-up
Answers use only this Investigation's retrieved sources; new sources are retrieved when needed.
Insufficient evidence found. The provided sources do not directly establish whether the productivity findings generalize to European organizations. **What may transfer:** Generative AI improved productivity in customer support and selected software-development tasks, with benefits varying by task and employee experience. However, the geography of these studies is not reported, so they cannot establish European applicability. [S10, S20, S24] **What the additional evidence adds:** A randomized field experiment at a global company found that AI improved performance on product-innovation challenges. This extends the evidence beyond routine tasks, but the participants’ geography is not reported; a global company setting alone does not establish applicability to Europe. [S29] A randomized online experiment also found larger gains among less-educated participants, but it was conducted outside firms and does not establish outcomes in European workplaces. [S34] **European-specific evidence is missing:** The UK academic survey does not report a specific employee-productivity outcome. It therefore cannot confirm or refute productivity benefits for European organizations. [S16] Overall, these findings offer plausible hypotheses for European organizations—not validated regional conclusions. Context-specific evaluation would need to distinguish task-level speed and quality from net organizational gains, including verification and coordination burdens. The sources support organizational experimentation as a way to investigate that uncertainty, rather than assuming gains will transfer automatically. [S21, S31]
6 additional sources retrieved