Category: Commentary

  • Giving outsiders a look at how people actually use Claude

    We finally got a peek at how people genuinely use an AI assistant, because Anthropic let researchers outside the lab do it.

    Until now, real-world data on Claude use has been stuck in two unsatisfying places. You had the analyses Anthropic itself publishes, which use real data but answer the company’s questions, not yours. Or you had public datasets, which anyone can study but which skew toward casual use. Neither tells you much about what most people actually do with AI day to day.

    This spring Anthropic ran a pilot to change that: it gave three independent research groups access to aggregate, privacy-preserving Claude usage data, through its internal tool Anthropic Insights. Roughly 250,000 conversations from April–May 2026. The groups designed their own studies, ran their own analyses, and kept the right to publish even results that make Anthropic look bad.

    Delegation of consequential tasks

    The first finding cuts against a tidy assumption. Earlier research suggested people hand AI the low-stakes tasks and keep the important ones for themselves. The SALT Lab at Stanford found the opposite is now common: in over half of conversations, people delegated consequential tasks to Claude, especially when seeking professional guidance on legal or financial questions. That’s a real shift, and it’s the kind of thing you only see in real usage.

    In nearly three-quarters of conversations, people set the direction and then adapted Claude’s output rather than using it verbatim. That friction isn’t wasted effort. Getting good results pushed people to clarify their own intent, which is a benefit that doesn’t show up in a demo.

    How it feels to use

    Oxford’s Human Information Processing Lab is looking at how people feel while using Claude. Early patterns are striking: when Claude is warm, people get more positive; when it refuses or disagrees, people push back; when it’s eccentric, people get more intellectually engaged. The emotional dynamics look similar to how people engage with everyday internet browsing.

    Measuring real productivity

    METR, meanwhile, is trying to quantify real productivity gains from coding agents. Its preliminary analysis suggests newer models deliver a significant speedup over older ones, and it was able to cross-check Claude’s own time estimates against known completion times from a developer study.

    The caveats

    There are plenty of caveats. The pilot ran on ~250k conversations, a snapshot that Anthropic chose. The privacy audit only covers the data shared so far. And “consequential” is a judgment call researchers had to define themselves. Still, this looks like a genuine step toward opening up data that was previously locked inside a handful of labs.

    The details matter less than the direction. Independent researchers now have a vetted path into real usage data, and they’re already finding things the company says it wouldn’t have thought to look for. If the program scales, we’re going to learn a lot more about how AI actually fits into people’s lives, whether or not the companies involved like what turns up.