← Articles
Use AI, don't become it

Use AI, don't become it

The research on AI and productivity looks contradictory until you separate two modes of use. One produces measurable cognitive offloading. The other produces reasoning better than working alone — same tool, same people, opposite outcome.

12 August 2026 · 9 min read · Crocodata

First, the productivity numbers

Corporate AI investment keeps climbing. The returns are, so far, uneven.

In an NBER survey of roughly 6,000 executives across the US, UK, Germany and Australia, about 80% report no measurable productivity impact. That's perception data, but the signal repeats. The UK Government ran its own controlled Copilot trial — 300 people, three months, independent evaluation — and found users completed Excel tasks more slowly and to lower quality than the group without it.

~80%Of executives report no measurable productivity impact (NBER)
5.7% / 1.6%Time spent using AI tools versus total work time saved (St. Louis Fed)
+3h 15mExtra hours worked per week in high-AI-exposure roles (NBER)

The St. Louis Fed measured the ratio directly: workers spend 5.7% of their working time using AI tools and save 1.6% of total work time. And NBER's Extended Workday research finds employees in high-exposure roles working an extra 3 hours 15 minutes per week — the time saved is being reallocated to more tasks, not to rest or deep work.

Then the other side. A BIS/EIB study of more than 8,800 European firms found +4% short-run labour productivity among AI adopters. The Atlanta Fed estimates AI now contributing around 0.8 percentage points annually, and rising. BCG and Harvard Business School found GPT-4 users completing 12.2% more tasks, 25.1% faster, at 40% higher quality — on the right kind of task.

The technology is not failing. The adoption model is. Everyone seeing real gains changed how they work, not just which tools they opened.

Then, the brain drain

There is a second cost that doesn't show up in productivity dashboards.

MIT Media Lab (Kosmyna et al., 2025) put EEG caps on 54 people over four months. The group using LLMs to write showed the weakest neural connectivity across memory and reasoning networks — below even the group using search engines. After four months, they struggled to accurately recall essays they had themselves submitted. The researchers named the accumulating effect cognitive debt.

The study is preliminary and not yet peer-reviewed; Harvard faculty reviewing it in late 2025 summarised it fairly: the risk of cognitive atrophy is real, but preventable.

Microsoft Research (Lee et al., CHI 2025) surveyed 319 knowledge workers and analysed 936 real work examples. The more workers trusted AI output without questioning it, the less critical thinking effort they applied. Crucially, workers who actively verified and challenged outputs showed no reduction at all.

Two very different methods, one shared conclusion — and one shared caveat most coverage skips: the damage attaches to a mode of use, not to the tool.

The study that actually explains it: Gerlich (2025)

This is the one worth reading properly, because it does something the others don't. It separates the cause from the correlation, and then tests the fix.

Part one: the survey

Gerlich, at SBS Swiss Business School, surveyed 666 participants and found that frequent AI use correlated negatively with critical thinking scores. That headline travelled widely. The important detail didn't: the relationship was mediated by cognitive offloading. AI use itself did not predict weaker thinking. Handing off the thinking did. When offloading was statistically controlled for, the effect largely disappeared.

Two subgroup findings sharpen the picture. Younger participants (17–25) showed the highest offloading and the lowest critical thinking scores. Higher educational attainment was protective — people with more training were more likely to interrogate output rather than accept it, regardless of how often they used AI.

The tool was never the variable. The delegation was.

Part two: the experiment

Gerlich followed up with a controlled cross-country experiment across Germany, Switzerland and the UK (n=150), comparing unstructured AI use against a structured prompting protocol: state your own position first, ask the model for counter-arguments, require it to justify its reasoning.

The structured group didn't merely avoid the decline. Their critical reasoning scores came out above the human-only baseline. Same model, same participants, same task type. Opposite outcome.

That's the whole argument in one dataset. Passive use produces measurable offloading and weaker reasoning. Structured use produces measurably better reasoning than working alone. The variable under your control is not which model you use — it's whether you show up with a position of your own.

The 5% who redesigned their work

Google and Ipsos surveyed 4,464 US workers in February 2026 and found that only 5% qualify as AI-fluent — defined as having reorganised significant portions of their work around AI, rather than bolting it onto existing habits.

5%Of workers qualify as AI-fluent (Google/Ipsos, Feb 2026)
4.5×More likely to report higher wages
8h vs 3hMedian weekly hours saved, fluent versus casual users

What separates them is not tool access. It's three choices. They aim higher — complex analysis, research synthesis, scenario planning, not formatting and email (Asana Work Innovation Lab, n=9,000+). They stay in the loop — active guidance of outputs produces 30–35% productivity gains, while passive acceptance produces marginal gains and more fatigue (Stanford HAI). And they change the metric — not tasks completed per hour, but value created per hour of human judgment.

The other 95% added a very fast tool to an unchanged process and expected the process to improve.

Sycophancy: your AI is built to agree with you

RLHF — reinforcement learning from human feedback — trains models on human approval ratings. Agreeable answers score higher, so the model learns that agreement is rewarded. This is structural, not a bug awaiting a patch.

The effects are measurable. In Science, Cheng et al. evaluated 11 models and found AI affirmed users' actions 49% more often than human peers did — including in scenarios involving deception or clear harm. The SycEval benchmark found sycophantic behaviour in 58.2% of medical and mathematical queries, with models switching from a correct answer to an incorrect one in 14.7% of cases purely because the user pushed back. Two separate teams, different methodologies, same direction.

MIT CSAIL and the University of Washington showed in February 2026 that even a perfectly rational Bayesian reasoner becomes progressively more deluded through repeated interaction with a sycophantic model. Being smart is not protection.

In April 2025, OpenAI rolled back a GPT-4o update within 48 hours for being too agreeable. Users liked it. Researchers were alarmed. The most engaging model and the most honest model are not the same model.

How to avoid it: you have to explicitly ask for the opposite of what the training rewards. CHI 2025 research (Drosos et al.) found that AI prompted to challenge users rather than validate them produces better decision quality and more calibrated thinking. Four prompts that reliably work:

  • What is the strongest argument against this conclusion?
  • Where is my reasoning most likely to be wrong?
  • Give me three significant flaws in this approach.
  • What would my most rigorous peer reviewer say about this?

Three rules

  1. Think before you prompt. Write your own answer first — even two sentences. This is the Gerlich finding in practice: having a position of your own is what prevents offloading, and it preserves the memory encoding that builds durable expertise.
  2. Give AI the heavy work, not the light work. Point it at diagnosis, scenario analysis, synthesis, stress-testing assumptions. Formatting and email are where the fluent 5% don't spend their AI budget. Aim high enough that the output is worth arguing with.
  3. Prompt for disagreement. End every significant session with "What's wrong with this?" or "Argue the opposite." It neutralises the sycophancy RLHF builds in by design, and it's the single behaviour that moved reasoning scores above baseline in Gerlich's experiment.

The posture, not the tool

Two modes exist.

Accepting

You accept and the model decides. Your brain shifts into skimming and approving, and you accumulate cognitive debt — busy, but not sharper.

You can tell you're here when your first draft is always the model's draft.

Authoring

You frame the problem, the model executes, and you challenge the result. Your brain stays in authoring mode, and your expertise compounds.

You can tell you're here when the model surprises you and you understand why.

Nothing in the research says AI makes you worse at thinking. It says one specific way of using it does — and that the same tool, used deliberately, makes you measurably better than working alone.

The people who thrive won't be the ones who use AI the most. They'll be the ones who know how.

Sources

NBER (2025/2026) · UK Dept for Business and Trade Copilot Trial (2025) · St. Louis Fed (2025) · BIS/EIB (Jan 2026) · Atlanta Fed (Mar 2026) · BCG/Harvard Business School (2024/2025) · MIT Media Lab — Kosmyna et al., arXiv:2506.08872 · Harvard Gazette (Nov 2025) · Microsoft Research — Lee et al., CHI '25 · Gerlich, MDPI Societies 15(1),6 (2025) and MDPI Data 10(11),172 (2025) · Google/Ipsos (Feb 2026) · Asana Work Innovation Lab (2025) · Stanford HAI (2024) · Cheng et al., Science 10.1126/science.aec8352 · Fanous et al., SycEval (2025) · MIT CSAIL/UW (Feb 2026) · Drosos et al., CHI 2025

This is one of the topics we keep working through at Crocodata. If you are thinking about the same things, we would like to hear from you.

Get in touch