The UX Researcher's Guide to Claude, Claude Cowork, and Claude Code

Which tool you need, how to set it up, and what the risks actually are

B
Brittany Hobbs · · 30 min read
The UX Researcher's Guide to Claude, Claude Cowork, and Claude Code
Overview
  • Claude, Claude Cowork, and Claude Code are not version increments of the same product — each represents a different relationship with AI that fits a different stage of UX research workflow.
  • Correction: Cowork is not covered under Anthropic's BAA in any configuration, and Claude Code is only covered with zero data retention explicitly enabled, not automatically inside an Enterprise seat. If you handle PHI, the compliant surface is chat plus an API pipeline, not the agentic tools this guide otherwise recommends.
  • Opus 5 became the default model on Max and the strongest on Pro on 24 July, with no migration step — any Claude Code brief calibrated before that date needs a spot-check, not blind trust.
  • Matching model and effort level to the task is now the biggest lever on cost and quality, ahead of prompt caching and batch processing.
  • Data privacy varies significantly by tier — Free, Pro, and Max consumer plans all default to training-on-inputs unless you opt out; Team, Enterprise, and API usage are excluded from training by contract.

Updated July 27, 2026. Claude Opus 5 shipped 24 July and is now the default model on Max, the strongest available on Pro. Self-serve HIPAA configuration shipped 14 July and was missing from the last update on July 22, 2026.

A 2025 Microsoft and Carnegie Mellon study found that knowledge workers applied zero critical thinking to roughly 40% of their AI-assisted tasks. A separate METR study of experienced developers found something stranger: practitioners using AI tools believed they were 24% faster, but measured against a control they were actually 19% slower. Productivity feel and productivity reality were inversely correlated.

Both findings land hard in UX research, where we are paid for the rigour of our judgment rather than the speed of our output. An AI tool that makes you feel more productive while quietly degrading your analytical standard is a liability you cannot see by looking at the artefact.

I have spent more than 20 years in UX research and led 300+ client engagements, and it took real trial and error before I arrived at a method I trust: treat every research engagement like a product, with a written PRD, explicit guardrails, a defined quality bar, and constraints on what the AI is allowed to do at each stage. At PH1 and AI Value Acceleration, we now help research teams set up exactly that working method. I am writing this so you can shortcut past the failures I went through, and keep the rigour and security you're paid for intact while you fold these tools into your practice.

The tools are everywhere and the expectations are building, but almost all of the guidance online assumes you're a developer building software or an executive buying it — not someone who designs studies, interviews participants, synthesises qualitative data, and owns the integrity of the finding. This article is written for that person: the three Anthropic tools you actually have access to, why each exists, how each fits a different stage of your workflow, and the privacy, security, and rigour risks most coverage skips.

The harder question, rewiring how you think to use these tools well, is the subject of Part Two: The Cognitive Shift Every UX Researcher Needs to Make. And Part Three: What UX Research Looks Like When Context Becomes the Engine covers the structural shift the role is going through as research becomes the context layer that powers everyone else's AI. Read this one first.

If you want the wider view first, The Future of UX Researchers lays out where the role is heading and the three postures toward AI that decide your place in it — this guide is the practical starting point for the posture that bets on learning these tools well. The stakes for getting that right, not just adopted, are the subject of our reporting on AI's silent failures: 79% of AI failures never surface as a complaint, exactly the kind an unreviewed AI process won't catch either — a conclusion Moritz Sudhof of Bigspin and the UX research team at Microsoft AI reached independently on our podcast.


The Job Market You're Reading This In

The pressure to learn these tools is not arriving in a vacuum. The conditions matter.

Indeed reports UX research job postings fell 73% from 2022 to 2023, one of the steepest single-year declines the discipline has seen, and postings have not recovered since. The User Interviews State of User Research 2025 found that 49% of researchers now feel negative about the future of UXR, a 26-point increase from the prior year, and 21% of surveyed organisations reported laying off UX researchers in 2025. A 2025 Measuring U analysis found that 35% of organisations reported reducing UX staff.

The stories behind those numbers are familiar: research teams cut from eight people to two, discovery studies compressed to "just the key headlines," designers absorbing research as a side responsibility, agency pipelines from 2022 considerably harder to sustain.

An analysis by UX Army found AI tools already automating entry-level research tasks: basic transcript coding, usability pattern identification, survey analysis, the work that once served as career on-ramps. The Nielsen Norman Group State of UX 2026 report describes the researcher role evolving toward strategy and synthesis while the execution layer compresses, and senior practitioners are surviving while entry-level and mid-market roles are not recovering at the same rate.

The discipline is consolidating into a different shape, recognisable to anyone who has spent time at the senior end of this craft.

Update: the more recent data reads less bleak than the numbers above. Postings jumped more than 15% in a single month this spring, and the share of organisations calling research essential to their strategy nearly tripled in a year, from 8% to 22%. The field is being redistributed, not erased — see The Future of UX Researchers for the fuller picture and what it means for where you place yourself in it.

Product teams in 2026 ship faster than ever — vibe coding, design AI, automated A/B testing, AI-driven analytics all compress the gap between idea and live product — which changes the demand profile for research dramatically. The trickle of small evaluation studies and mid-funnel usability checks that filled most UX calendars from 2018 to 2022 is being absorbed into faster product cycles, handled by AI tools or designers themselves.

What remains, and becomes more valuable per engagement, is foundation and generative research: discovery interviews that surface what an audience actually needs, behavioural research that reveals what people do rather than say, strategic synthesis that connects evidence into a defensible direction. This work is lower-incidence than it used to be and higher-value per study, because shipping the wrong product fast now costs more than shipping the right one slow.

The research function that survives works both ends of this barbell: the strategic end that justifies a dedicated function, and the execution end where AI tools handle the coding, tagging, and first-draft synthesis that used to consume most of a researcher's week. The middle is what's being compressed. Researchers who can credibly do both ends — and who use these tools well enough that execution doesn't burn their week — will be in a fundamentally different position than those who can do only one.


What Claude Code Is Actually For

The claims about Claude Code circulating right now range from accurate to absurd, and the gap between them is where UX researchers are getting hurt.

The accurate version, in research terms: you have a folder with forty interview transcripts, a coding taxonomy, and a research question. You install Claude Code, describe your project and methodology, and ask it to apply your taxonomy across all forty transcripts, outputting structured CSVs with the verbatim quote behind each code, a confidence rating, and an explicit flag on anything ambiguous. That would take a careful researcher most of a week by hand. Claude Code produces a competent first pass in an afternoon — you still review and correct it, you're still the analyst — but the execution layer that used to eat the week now runs alongside your actual analytical work.

That capability is genuinely new, which is why Claude Code reached $1 billion in run-rate revenue inside six months of launch and why "vibe coding", describing what you want in plain English and watching software get built, has entered mainstream tech vocabulary.

The absurd version is uglier: LinkedIn influencers and self-styled "AI research consultants" claiming you can replace your entire research function with prompts, run a complete discovery study in 90 minutes, or generate insights from no actual data. These claims are predatory — aimed at scared researchers, credulous leaders, and budgets looking for an excuse to cut — and produce work that looks like research without being it.

What research teams we work with are actually wrestling with, underneath the gap between those two stories, is leadership pressure on one side and unclear capability on the other: leaders judging researchers against capabilities that don't exist, researchers feeling inadequate against an imaginary benchmark and either freezing or chasing the wrong workflows. The defence against both is a precise picture of what each Anthropic tool does and doesn't do, which is what the rest of this article lays out.


The Case for Learning These Tools Now

An OpenAI productivity analysis covered by VentureBeat found a 6x productivity gap between AI power users and everyone else, with identical tools available to both groups. The State of User Research 2025 found that 88% of researchers identified AI-assisted analysis as the top trend shaping UXR in 2026.

Not every piece of AI hype is justified, particularly for research work. The case for learning is narrower and harder to argue with: the gap between practitioners using these tools with genuine rigour and those who don't is now large enough to show up in what teams can take on, what clients pay for, and who makes it through the next round of headcount decisions. Experienced researchers shouldn't have to justify their employment by mastering tools built for a different discipline — but the conditions are what they are.


What Changed, and Whether You Need to Act

Two things shipped since this guide's last update that were worth creating a revised guide a week later!

  • Opus 5 became the default model on Max, strongest on Pro, 24 July — nobody opted in. Re-validate your coding brief before your next study.
  • Self-serve HIPAA configuration went live 14 July. It surfaced that Cowork was never a compliant surface for PHI, in any configuration, and that Enterprise-seat Claude Code isn't automatically covered either. If you handle PHI, read the risks section before your next transcript.
  • Neither changes your pseudonymisation obligation. This is still on you to manage. No new action beyond what you should already be doing.

Three Tools. Three Different Relationships with AI

A common mistake is treating Claude, Claude Cowork, and Claude Code as version increments of the same product. They're not. Each represents a different relationship between you and the model, and each fits a different part of the UX research workflow.

Claude (claude.ai, Free or Pro)

A web-based chat interface. Each conversation starts fresh by default, but Claude Pro now includes Projects, which let you attach documents (briefs, taxonomies, prior reports) the model references on demand inside that Project's conversations — persistent context at the personal-account level, but only inside Projects you set up. Outside of them, you are the memory.

Best fit: desk research and source synthesis when you need to think alongside a fast reader, refining your own writing, drafting interview guides, sense-checking a research question before you brief the team, exploratory analysis you want to talk through rather than commit to a deliverable. Less suited to high-volume processing or shared team methodology.

Claude Cowork (every paid plan, Pro through Enterprise)

Claude Cowork is Anthropic's agentic product for knowledge work. It launched on the Claude desktop app for macOS and Windows, where, with permission, it reads, edits, and creates files inside folders you grant access to — delivering finished documents into your working folders rather than producing text you copy out. It launched Max-only, opened to every paid personal plan within days, and as of July 2026 is also rolling out to web and mobile; the local-file-access model described below is specific to the desktop app, and the file-handling rules on web and mobile are still settling — worth checking directly before you point either at participant data.

Cowork itself gives every paid plan Projects with shared knowledge bases and system prompts. Team and Enterprise add what a research org actually needs on top: admin controls, role-based access, a guarantee that your inputs are never used for training, and connectors to services like Google Drive and Gmail. A personal Pro or Max login gets the same Cowork experience but without those controls — and without the training guarantee, which matters more than it sounds; see the risks section below.

On the desktop app specifically: filesystem access is scoped to the directories you explicitly grant. The model runs with your user-account permissions, so the directories you authorise are the directories it can read and write. Treat the folder you point Cowork at the way you would treat a folder shared with a contractor: structure it deliberately, keep raw participant data out of it unless pseudonymised, and audit what the model has touched at the end of each session. For sensitive research, run it against a dedicated working folder rather than your home directory.

This is the right tool for most working UX researchers running active studies.

Claude Code (agentic coding tool — CLI, browser, or app)

Claude Code started as a command-line tool you install locally, and that's still the fullest version. But it's no longer only that: you can open it in a browser with no setup, or pick a task back up from the Claude app on your phone after starting it elsewhere — start at your desk, check in from anywhere, finish wherever's convenient. However you run it, it reads and writes files in your project folder, runs scripts, and executes multi-step tasks autonomously rather than waiting on chat turns: process all 47 transcripts using this taxonomy, output a structured CSV with grounded quotes and confidence ratings. That's an automated research pipeline, not a chatbot session.

The real distinction from Cowork isn't privacy — both follow the same account-level training rule, explained once in the risks section below rather than repeated here. The difference is how directly you're running things. Claude Code, installed locally, runs as a program on your own computer: it only sends the specific piece of text it's working on for each step, not your whole folder. The newer browser version needs no install and is handy for picking up a task from your phone, but it runs in a hosted space rather than your own machine — a different setup, worth confirming directly with Anthropic before you point it at anything with participant data in it.

Because Claude Code can read and write files on its own, set it up the way you'd set up a new contractor: access to one project folder, not your whole computer. Anthropic also runs it inside a built-in security sandbox by default — a locked-down space limiting what it can touch — with settings to explicitly block it from reading sensitive files like saved passwords. Keep it inside a folder built for that one project, don't hand it more access than the task needs, and glance over what it changed each session. For regulated data, treat the sandbox as mandatory, not optional. Anthropic's Claude Code sandboxing documentation has the full detail if you want it.

Setup takes real time, and the ceiling on what is possible is significantly higher than the chat surfaces.


The Model Changed Underneath You

Opus 5 became the default model on Claude Max, and the strongest available on Claude Pro, on 24 July — two days after this guide's last update, with no prompt and no migration step. If you were logged in on Max, your next Claude Code session simply ran on a different model than your last one did.

In research terms: a silent model change is an undocumented change to your instrument. For workflows you've built, or more complicated systems, you should update your Claude Code project brief for two to eight weeks until the output is consistent enough to trust — that calibration, wherever you did it before 24 July, was done against Opus 4.8. Opus 5 is a genuinely stronger model. This means it will act differently and will attempt to be more proactive in ways you may not expect.

None of that is a reason to be wary of the upgrade — it's the reason to be genuinely glad about it. Opus 5 is built to check its own work before handing it back, which is exactly the failure mode this guide spends the most time warning about: the model catching more of its own mistakes before you ever see them. It also reads charts, documents, and diagrams well enough to bring session screenshots and annotated prototypes into an analysis directly, and it can hold an entire study corpus in view at once, finding a connection between transcript 3 and transcript 38 that a smaller-context model would simply never see.

Treat any model change the way you'd treat a new coder joining a study mid-stream: competent, unproven on this specific brief, and worth a spot-check before you trust its output at the same level as before.


Match the Tool to the Stage of Your Workflow

These three tools are not stepped versions of the same product, and the question is not which one to commit to. The question is which tool fits which part of a research engagement. Most working researchers will use all three across the lifecycle of a study, just for different stages.

If you're doing desk research, refining your writing, or thinking alongside a quick reader: use Claude.ai. This is the right surface for early-stage exploration, sourcing and summarising published research, sense-checking a research question, drafting an interview guide, or sharpening language in an executive summary. The Pro Projects feature lets you attach a brief or style guide the model references on demand — the strength is fluency and speed against well-defined inputs.

Correction: this guide previously named Cowork as the default day-to-day surface for active studies with real participant data. That's wrong wherever the data includes PHI — Cowork is not an Eligible Service under Anthropic's BAA in any configuration, on any plan, at any retention setting. Claude Code fares better but not automatically: it's covered only with zero data retention enabled on a qualified account, which an Enterprise seat does not give you by default. The trap: a research lead who sees Claude Code bundled into their org's Enterprise plan reasonably assumes it sits inside the compliance perimeter. It doesn't, unless ZDR has been explicitly turned on — full detail in the risks section below.

If you're running active studies handling general participant PII, not PHI, and want a shared team methodology baked into every session: use Claude Cowork. This is where the work compounds. A Cowork Project holds your research framework, analysis taxonomy, screener, discussion guide, and quality bar in one place, and every conversation inside it starts informed; the desktop app's local file access lets the model deliver finished synthesis directly into your working folders. Any paid plan gets this. Team and Enterprise add admin controls, role-based access, a guaranteed no-training commitment, and audit visibility IT and Legal will care about. For most researchers running studies with real participant data, a Team or Enterprise login is the day-to-day surface — Pro or Max works too, but carries the training exposure covered below.

If you're processing volume or building a repeatable research-ops workflow: use Claude Code. Forty transcripts needing consistent coding, six studies to synthesise across, a quarterly operation you currently dread, a multi-stage pipeline (coding pass → thematic synthesis → recommendations → stakeholder summary) where each stage is constrained by the brief above it. The setup investment is real (two to eight weeks for the first serious workflow) and so is the security surface, but the upside is structural. Run the local install and source data stays on your machine, only excerpts going to the API. One thing changes the scale of what's worth attempting: a full study corpus — every transcript, taxonomy, screener, prior studies — now fits in a single Opus 5 session at everyday rates, making cross-quarter meta-synthesis affordable for the first time. That matters analytically, not just economically: a chunked pipeline structurally cannot find the contradiction between transcript 3 and transcript 38, because it never holds them both in view at once.

If your evidence isn't text: feed it in directly. Opus 5's vision gains now cover charts, documents, and diagrams well enough to put session screenshots, annotated prototypes, survey charts, and whiteboard affinity maps into any of the three tools instead of transcribing them first. Connectors provide the evidence; the model still only provides the reasoning — and every connector is a new data path that has to clear whatever coverage regime applies to that study.

A few honest questions to ask before any of them touch participant data:

  • Ready-made app, or comfortable setting up a project folder and CLAUDE.md yourself? The former means Cowork — nothing to configure. Claude Code rewards more hands-on setup and is reachable whenever you're ready, not necessarily day one.
  • What kind of data will you be feeding it? Named participants or PII: a Team or Enterprise login, not personal Pro or Max — pseudonymisation comes first regardless. Read the risks section before you do anything.
  • Delivering next week, or building a workflow that compounds? Next week sits in Claude.ai or Cowork. Compounding capacity sits in Cowork plus Code.

The cleanest starting point: match the tool to the stage of work, keep the rigour and security it demands, and add the next tool only when the current one stops being enough. Most teams we advise start with Cowork, layer in Claude.ai for desk research, and add Claude Code once they have a repeatable workflow worth encoding.


Getting More Out of Each Tool

Once the three tools are mapped to your workflow, the highest-leverage thing left to learn is which model and effort level to point at which task — a lever most readers are getting wrong by default, because both settings default to the most expensive option.

Match the model to the artefact, not the tool:

Sonnet Opus 5
Interview guide drafts, screener variants Cross-study synthesis across a full corpus
Formatting and cleanup passes Contradiction-hunting between participants
Tagging against an unambiguous taxonomy Evidence-weighted opportunity mapping
Single-transcript summaries Multi-quarter research archives

Then set the effort level in the same breath. Effort defaults to high on the API and in Claude Code, and for a lot of research work — tagging, formatting, single-transcript summaries — low or medium effort holds quality while cutting tokens and latency sharply. Most readers are paying frontier rates for what's genuinely a clerical pass. The trap: effort controls how much the model thinks, not how much it says. Opus 5's default responses already run longer than prior Opus models did, and getting a tighter one takes an explicit instruction to be concise, not a lower effort setting.

Caching and batching still help with volume, just less than routing does. Prompt caching cuts the cost of repeated context — your methodology brief, taxonomy, discussion guide — by roughly 90% on every call after the first. The Message Batches API gives a flat 50% discount on anything that can wait 24 hours, which covers most overnight transcript coding. Stack both under the right model and effort level, and a research-ops pipeline in Claude Code costs a fraction of what specialty UXR analysis platforms charge per seat.

The strategic implication. Most specialty AI-for-UXR tools are wrappers around the same underlying models, marked up substantially per seat. A team with the discipline to write a clear PRD and a defined taxonomy can replicate most of what those tools do with Cowork plus a well-routed, prompt-cached, batch-processed Claude Code pipeline — at materially lower cost, with your methodology as the source of truth. This is the pattern we help teams set up at PH1 and AI Value Acceleration when the answer to "should we buy [specialty tool X]?" is "let's see what your existing methodology can do first."

The technical levers above are necessary, not sufficient — the methodology and brief discipline that make any of this worth running is the subject of the companion piece.


Claude Cowork: Setup and What to Expect

Setup: 15–30 minutes

  1. Go to claude.ai → sign in or create an account → any paid plan unlocks Cowork; upgrade to Team under Settings → Plans if you need admin controls, role-based access, and a guaranteed no-training commitment for participant data
  2. Navigate to Projects (left sidebar) → New Project
  3. Name the Project for the engagement or methodology
  4. Under Project Knowledge, upload your core context: research framework, discussion guide, analysis taxonomy, relevant client brief
  5. Add a system prompt. Even one sentence changes the quality: "You are assisting a senior UX researcher. Before summarising any pattern, identify and explicitly surface disconfirming evidence."
  6. Invite team members from Settings → Team Management

Every conversation inside that Project now inherits all of that context. You stop re-explaining your methodology every session.

What it does well for UX research: applying a consistent analysis taxonomy across multiple interview sessions, generating discussion guide variants from a master template, structuring debrief notes into a standardized format, drafting synthesis with your methodology already loaded, keeping shared team methodology visible and consistently applied.

What it won't do: catch contradictions you haven't asked it to find. Default LLM behaviour produces coherent, pattern-aligned summaries, and in qualitative UX research the most important finding is frequently the one that doesn't fit. Build the disconfirmation ask into every analysis prompt as a structural requirement, not a polite gesture.

A practical evaluation criterion, borrowed from Robert Brunner (Apple's Industrial Design Group): after two weeks of real use, count the steps in your workflow before and after. If Cowork added steps rather than removed them, your brief or project context is doing too little.


What Claude Code Looks Like in Practice

Almost every description you'll find online assumes you're a developer, and that framing is exactly what's intimidating most researchers off this tool unnecessarily. Here is what it actually feels like in research practice.

You install Claude Code. It opens in a terminal window — the black box with the blinking cursor that probably gives you IT-helpdesk flashbacks. Don't let the access point stop you; the actual product is what happens after you type your instruction and press enter.

You navigate Claude Code into the folder where your research project lives. You write a plain-text file called CLAUDE.md that explains what this project is, what methodology you use, what rules you want it to follow. ("Always surface disconfirming evidence before summarising a pattern. Never invent quotes. If a code is ambiguous, flag it; do not force it.") Then you give it a task: analyse these transcripts using this taxonomy, output a structured CSV, flag low-confidence interpretations. You watch it work through the steps you would have done by hand.

Opus 5 changes what "give it a task" should mean. It's built to be more thoughtful and proactive, and part of its edge comes from checking its own work before handing it back — which supports a real shift in how you brief it: define the outcome, supply the context, and let the model plan the steps, rather than scripting every step yourself. That shift needs a guardrail in the same breath, though, because a single assignment that asks for themes, contradictions, confidence ratings, JTBD statements, follow-ups, and an executive summary comes back looking finished — and a finished-looking deliverable gets interrogated less than a patchy one, which is exactly backwards. So state the failure condition alongside the outcome:

If the transcript does not support a code, return NULL and flag it. Do not infer, do not reconstruct, do not work around missing data.

The model got better at finishing. Finishing is not the same as being right.

The pattern across research leaders who've actually integrated this is consistent: write a careful brief, test it on five transcripts (not fifty) to find where it fails, tighten the brief, repeat. They iterated for two to eight weeks before output became consistent enough to trust without heavy correction — after that, every one of them describes a transformation in what their team can take on.

Equally consistent is what's missing from those accounts: a researcher solving a real research problem on a first prompt, trusting Claude Code's analytical judgment without expert review, or succeeding by skipping the brief and hoping the model would figure out what they needed.

The honest version of "is Claude Code for me right now": do you have the time, the kind of work, and the data volume to make a setup investment worthwhile in the next two months? If yes, the Anthropic Claude Code documentation will get you started. If no, Cowork will serve you well meanwhile — the option to expand later is always there.


A Few More Features Worth Leveling Up Into

None of this is required to get value out of Claude Code or Cowork. Treat it as the next rung once the basics feel comfortable, not a prerequisite.

Skill files. A skill is a saved set of instructions Claude can reuse instead of you re-explaining the same thing every session — your coding taxonomy, your report format, your quality checklist. Ask Claude to help you build one for something you do repeatedly, and it will ask you questions until it has enough to write the file itself. Once it exists, both Claude Code and Cowork can use it, automatically or by name.

Auto mode (Cowork, beta). Instead of working a task with Cowork turn by turn, hand off the whole thing — "synthesise these 41 customer calls" — and let it run end to end, checking in from the web or your phone instead of sitting with it the whole time. Reach for it once you trust a workflow enough not to watch every step.

Dynamic workflows. For bigger jobs — auditing every transcript across a study, six prior studies at once — Claude Code can break the task into pieces that run in the background simultaneously instead of one at a time while you wait. You describe what you want; it works out the split and hands back one finished result. Worth reaching for once a task outgrows a single conversation.

Building a Repository for Claude Code

The project folder you point Claude Code at is often called a "repository," or repo for short — the folder itself, with your CLAUDE.md file sitting alongside your transcripts, taxonomy, and everything else Claude Code needs. A well-set-up repo is the difference between consistent output every time and something different every time you ask.

The highest-leverage move is the same one developers use: keep CLAUDE.md short, specific, and built from real mistakes you've caught, not a wish list of hopes. This explainer walks through a widely used approach, built from a running list of everyday failure patterns and shared by AI researcher Andrej Karpathy — a good model even outside of software. If you'd rather see it built for research specifically, Greg Isenberg's walkthrough of setting up a Claude Code repo inside an Obsidian vault is the closer match; the same approach works with a plain GitHub repo instead, if that's what your team already uses.

The practice that follows directly from a model changing underneath you: treat model identity as research provenance, not an implementation detail. Keep five hand-coded transcripts as a standing regression set, and any time the model changes, re-run and diff before touching live study data. Record the model ID — in CLAUDE.md, in Cowork Project knowledge, in the methods note of any deliverable — because "Claude" is not an answer when a client or ethics board asks in six months how a finding was derived; the model, version, and effort level are. (Enterprise admins can now control which models and effort levels their users access, if you want this enforced rather than requested.)


The Risks You Cannot Ignore

These risks are organised by what's actually in your data, not by industry — a compliance gap doesn't care whether the client calls the work healthtech. Three categories, in order of how much protection actually exists:

  • PHI has a BAA path. It's the only category here with a rail built for you, and the covered routes are narrow: the HIPAA-ready API on standard retention, and Chat on a HIPAA-ready Enterprise plan. Nothing else qualifies — see the coverage table and checklist below.
  • Children's data, biometric data, session recordings, financial data, employee data, and EU personal data sit under different laws entirely, and none of them has an equivalent rail in any Anthropic configuration. Treat each as its own compliance question, not a variant of the PHI one.
  • General participant PII — a named adult talking about a product — has no rail at all beyond your own research discipline: pseudonymise, minimise what you paste, and don't assume a plan tier is doing this work for you.

The absence of a rail for GDPR, PIPEDA, provincial health privacy law, or an IRB protocol is not the same as permission.

Is your data covered under a BAA?

Product Covered under Anthropic's BAA?
Claude Chat (HIPAA-ready Enterprise) Yes
Claude API (HIPAA-ready, standard retention) Yes
Claude Code Only with zero data retention enabled, on a qualified account — not automatic, even inside an Enterprise seat
Claude Cowork No, in any configuration
Claude Free / Pro / Max / Team No
Self-serve Enterprise, before HIPAA readiness is enabled No

For a researcher handling PHI, the compliant day-to-day surface is chat, code, or API.

Before anyone enables HIPAA readiness, know this:

  1. Only the org's Primary Owner can accept the BAA and enable it. Other Owners and Admins cannot complete the flow.
  2. The BAA is a standard agreement and cannot be modified.
  3. Enabling it resets certain settings across the organisation.
  4. The change cannot be reversed from organisation settings.
  5. A BAA signed for API usage before 2 December 2025 covers API only — it does not extend to the HIPAA-ready Enterprise plan.
  6. Team, Free, Pro, and Max cannot enable HIPAA readiness at all.

If your org runs both a HIPAA-ready API organisation and a separate ZDR organisation for Claude Code, those are two different organisation IDs, and someone needs to own making sure PHI never crosses into the uncovered one — set folder and project naming conventions for covered work before anyone starts.

None of this is legal advice. Read Anthropic's implementation guide before the Primary Owner clicks anything, and route the actual sign-off through your own counsel.

PII and participant data

Your research participants consented to being interviewed, not to having their verbatims processed by a third-party AI system. Pseudonymise before you paste anything: replace names, job titles, company names, locations, any identifier. This is sound research hygiene, applies at every tier, and it is not HIPAA de-identification, which has a defined standard and a much longer identifier list than most teams assume — a team that swapped names on a health transcript has probably not de-identified it, since dates, rare conditions, and employer details all survive. The inverse matters too: properly de-identified data falls outside HIPAA's scope entirely, and for a large share of healthcare UXR, rigorous de-identification is cheaper and faster than a BAA plus a separate organisation ID — the practical move for most readers in that position.

One more identifier teams forget: session recordings. A face and a voice identify someone as surely as a name does. Pseudonymising the transcript does nothing about the video sitting in the same folder Cowork can read.

Platform data handling differs significantly by tier

Since an August 2025 update to Anthropic's Consumer Terms and Privacy Policy, this table looks different from what most people still assume:

Account type Used for model training?
Free Used by default, unless you opt out in account settings
Pro / Max (personal) Used by default, unless you opt out in account settings
Team No
Enterprise No (per enterprise agreement)
API usage No

This table is about your account, not which product you happen to have open. Claude chat, Cowork, and Claude Code all follow whichever account you're logged into — and now that Cowork runs on every paid plan, "I'm using Cowork" no longer tells you anything about your training exposure on its own. A personal Pro or Max login carries the default-on rule inside Cowork exactly like it does in chat; only a Team or Enterprise login is exempt.

Pro does not get a training exemption the way it used to. If you've been treating a Pro subscription as the privacy-safe tier for participant work, that assumption stopped being true in August 2025 — opting out is a setting, not a default, in your account's privacy preferences. Check Anthropic's current privacy policy and commercial terms before you process any participant data; verify the current version, not what you remember reading months ago, and not what this table said when the article was first published.

GDPR and IRB implications

Sending EU participant data, even pseudonymised, to a US-based AI provider has GDPR jurisdictional implications. Your IRB protocol almost certainly wasn't written to cover LLM processing of participant data. HIPAA readiness solves none of this — it's US law covering a defined set of entities, and enabling it does nothing for GDPR, PIPEDA, provincial health privacy law, or an ethics protocol still unwritten for LLM processing; the organisation remains responsible for its own compliance either way. One more gap: Anthropic's BAA does not apply to services purchased through a third-party cloud marketplace rather than directly. If you're unsure whether your setup is compliant, treat it as non-compliant until you've verified with your ethics board or legal counsel.

Hallucination in qualitative synthesis

LLMs produce outputs that are statistically coherent given the input, and statistical coherence is not the same thing as research validity. A synthesis that smooths over an edge case is a finding-level failure, not a minor formatting issue. A 2023 Nature study found AI-assisted researchers producing more output alongside measurable convergence toward the median, and the most important UX findings are usually the ones that don't fit the dominant pattern. Build disconfirmation explicitly into every prompt.

The single most effective fix is procedural, not a smarter prompt: run a second pass, as its own step with its own prompt, in which the model re-examines its own coding against edge-case quotes and flags anything it can't support. Bundling this into the first pass defeats the point — it needs to happen as a genuinely separate look.


Deciding Is the Easy Part. Getting Value Is the Hard Part.

This article has helped you decide which tool fits which stage, how to set each up, where the cost-saving levers move the needle, and what risks to manage. That's the easy part. The harder part — whether these tools compound your capability or quietly produce more mediocre work at higher speed — happens after setup.

The aggregate data is brutal. MIT's Project NANDA found in mid-2025 that 95% of enterprise generative AI pilots produce no measurable business return, despite $30–40 billion in collective investment. Boston Consulting Group's September 2025 study of 1,250+ companies found only 5% achieving AI value at scale, while 60% reported essentially no value at all, and only 37% of executives could demonstrate clear ROI from AI initiatives even as 85% increased AI investment year over year.

These are behavioural-layer failures, not technology failures — the models work. What breaks down is the moment someone has to decide to use the tool well, abandon it, or fake using it. Organisations treated "deciding to adopt AI" as the hard decision and "getting value from it" as something that solves itself once licences are bought — the same pattern wasting tens of billions at the enterprise level shows up at the individual researcher level too. A researcher who installs Claude Code, runs a few prompts, gets mediocre output, and concludes "the tool doesn't work" is making the same mistake at a smaller scale. Using these tools well is a separate craft from the documentation, and it's the craft we work on with teams who want to get there faster than trial and error allows.

Part Two: The Cognitive Shift Every UX Researcher Needs to Make is the next thing to read, and Part Three: What UX Research Looks Like When Context Becomes the Engine follows it with the structural argument for why research is becoming infrastructure for the rest of the company's AI rather than a step that produces decks.


A Few Words Before You Go

The discipline you trained for is being reshaped under conditions that aren't fair, on a timeline that isn't humane. Most senior researchers are working this out in real time, including the ones whose LinkedIn posts suggest otherwise. The ones who come through with their standards intact treat it like learning any new method: deliberately, in paced iterations, grounded in the same rigour they apply to everything else.

If you only do one thing this week: set up a single Cowork Project on a study you're actively running, attach your methodology brief and taxonomy as Project Knowledge, and use it on two weeks of real research. Pay attention to where the output needs correction and where it doesn't — that's the first data point you have about how your methodology actually translates into a model context, and it's the same starting point we use when helping teams shift their practice.

Then read Part Two for the cognitive shift that determines whether the setup compounds, and Part Three for where the role itself is heading as context becomes the engine.

Eight further resources worth your time:

If your team wants advisory or training support setting this up internally, that is the work we do at PH1 and AI Value Acceleration.


Sources and Further Reading


Brittany Hobbs is COO and VP Research at PH1, CEO of AI Value Acceleration, and co-host of the Product Impact podcast. She has led research at Mozilla, Spotify, Google, BBVA, TELUS Health, and Schneider Electric across more than 300 engagements.


Corrected and updated July 27, 2026. This revision fixes a safety-relevant error: earlier versions recommended Cowork as the day-to-day surface for active studies without flagging that it's never covered under Anthropic's BAA, which matters if your participant data includes PHI. It also adds Claude Opus 5, which became the default model on Max and the strongest on Pro on 24 July, and the self-serve HIPAA configuration Anthropic shipped 14 July.

How helpful was this article?

Have a story to share?

0 / 500
B
Brittany Hobbs

COO, PH1 · CEO, AI Value Acceleration · Co-host, Product Impact Podcast

Latest Episodes

All episodes

Product Impact Newsletter

AI product strategy delivered weekly. Free.