Designing research for AI moderators: Lessons from running 100 interviews a week
From interview moderator to moderation designer
Rebecca Klee introduces a weekly research programme that combined around 100 spoken AI interviews with an insight report. She explains how automation moved her research judgement upstream into study objectives, moderator guidance and quality controls.
What kind of depth can AI moderation deliver?
Klee challenges the promise of qualitative depth at survey scale by comparing the evidence different research methods can produce. AI can probe volunteered explanations across many participants, but those conversations cannot establish population prevalence or reliably reveal unspoken motivations.
Where speed, flexibility and reduced social pressure help
Klee identifies recurring research, broader qualitative comparison and asynchronous participation as useful applications. Interviews with 14 previous participants reveal that convenience and reduced social pressure can support participation without improving conversational quality. She argues that suitability depends on the research question, participant context and need for human support.
Four levers for designing a better AI interview
Klee connects uneven context, excessive agreeableness and plausible inaccuracies to practical interview risks. She describes four design levers: time and depth, research goals, context and moderation guidance. Her guidance prioritises narrow objectives, recent concrete experiences, neutral questions and explicit limits on probing.
Reviewing 1,000 sessions reveals uneven improvements
Klee uses an AI-assisted transcript review to assess session conduct, question quality, probing and research coverage across 10 rounds. Structured objectives and depth controls coincide with stronger coverage and probing, while failures in conversational mechanics persist. She explains why changing studies, platform versions and review methods limit causal conclusions.
The hidden work participants do to rescue interviews
Two anonymised exchanges show moderators leaking instructions, stalling and making participants repair the conversation. Klee connects these failures to participant accounts of repeating answers, managing dead ends and withholding details to avoid unproductive tangents. Completed sessions and low visible frustration can therefore conceal missing evidence and participant effort.
Automation leaves research responsibility with us
Klee proposes judging sessions through research coverage, participant experience and trustworthy evidence together. She closes by assigning researchers responsibility for both advance design and retrospective review, including how participants compensated for failures. Platforms must also expose useful controls and meet clear standards for reliable research.
This this will be a talk on designing research for AI moderators. There are no right or wrong answers. We may challenge a few assumptions. You may leave with more questions than answers. Are you ready to get started?
For some of you, this audio may feel familiar. It adapts the introduction that participants hear at the beginning of an AI moderated research session so it seemed appropriate here too. AI moderated research is still new. The platforms and models are changing quickly and we don't yet have settled best practices. What I can share today is what I've learned from designing and reviewing these sessions over the past year where the method is useful, where it goes wrong and what I now do differently. This isn't a talk about the one correct way to use an IR moderator it's about learning to design research for one and the assumptions we may need to reconsider as more of the interview becomes automated. First some context on how I became involved.
Late last year, Askable approached me to run an InsightStream programme for one of its clients. The format was already established. Each week the client brought me a research goal. I turned it into a study, recruited participants, coordinated around 100 spoken AI moderated interviews and delivered an insight report before the next round began. I stepped into a program with that scale and cadence already in place. That distinction matters because I wasn't starting with a blank sheet of paper and asking whether AI moderation might work in theory.
I had to make it work repeatedly, at scale, inside a live research programme. The questions sometimes built on the previous round and sometimes changed completely. Either way I had to learn quickly enough to improve the following week. My role was to provide the research judgement within that system. I worked out what we were trying to learn, translated that into objectives and guidance, reviewed the sessions and decided whether the evidence supported a finding.
Although I was no longer sitting in the interview the work of moderation hadn't disappeared. Much of that judgement had moved upstream into the moderator's design and quality controls. I had moved from being the moderator to being the moderation designer. A quick caveat: these platforms and models are changing extremely quickly.
Some of the things I describe today may be fixed, others will no doubt appear. So I'd just suggest to listen less for a permanent catalogue of limitations with AI, and more for the disciplines that outlast them choosing the right method, setting boundaries and checking what actually happened. The capabilities will change the need for research judgement will last longer.
Before the design, what does AI moderation offer? You've probably heard the pitch qualitative depth at survey scale? By a show of hands who feels confident that AI can genuinely achieve depth at scale? Is unsure?
Who is skeptical? That makes sense because depth is doing quite a lot of work. Surveys give us reach but limited depth. Interviews give us depth but limited reach. AI moderation appears to offer both a conversation with every participant repeated hundreds or thousands of times.
But I don't think depth at scale is quite right. Saydab Bakshi, a quantitative researcher at OpenAI, calls this framing a fallacy because it treats methods as a single line with AI moderation finally giving us everything at once. An AI moderator can ask what changed, why something stopped feeling worthwhile, what someone did instead.
That is meaningfully deeper than an open text response. But more probing produces more words not necessarily everything we associate with a depth interview. A moderator can't yet observe behaviour, notice hesitation or investigate the gap between what someone says and what they do.
Nor can it reliably uncover reasons people don't understand or aren't willing or able to articulate. More sessions don't provide a valid measure of prevalence. If thirty out of 100 participants mention something that doesn't mean that thirty percent of the population experiences it.
AI moderation offers a particular kind of depth with particular structural limits. A more useful way to think about this is that different methods answer different parts of the research problem. Surveys tell us how many people experience something and how that varies between groups, but are largely limited to reasons we already know to ask about. Human moderated interviews reconstruct what happened, how someone behaved and how an experience felt, but their scale limits claims about prevalence.
Open texts can surface reasons we didn't anticipate, but tend to stop at a participant's first answer. AI moderation fits between them. It can take a volunteered reason and ask what sits underneath it and repeat that across many participants.
Bakshi calls this probed language at machine scale. In my program that language was spoken rather than typed, but the capability was the same. Rather than asking whether AI gives us depth at scale, I now ask what kind of evidence is this method structurally capable of producing?
That question is much more useful when deciding whether to use it. Where might AI moderation earn its place? First use case is recurring or time critical research. My programme moves from a focused question to completed sessions and an insight report each week.
Because interviews run-in parallel, evidence can inform a decision without waiting weeks for sessions to be conducted sequentially. That speed is valuable when it supports a real decision, not simply because faster is always better. The second use case is broader qualitative comparison exploring how and why experiences differ across customer groups or markets, including different language communities.
But broader comparison doesn't make the findings representative or turn mentions into percentages. Asynchronous participation can also help us reach people who are difficult to schedule, such as specialists, shift workers or those in different time zones.
But all of these are largely researcher benefits. I was curious to understand the experience and value for the person being interviewed. So I conducted human moderated interviews with 14 people who had previously taken part in AI moderated research. I heard that the clearest participant benefit was flexibility.
As one participant explained, I can do it at my own time, I don't have to sit down specifically for a session. For people with irregular work, caring responsibilities or large time zone differences that flexibility can be the difference between participating and not participating. Flexibility doesn't necessarily mean a better conversation.
All 14 participants spontaneously compared AI to human moderation. No one described AI as the better conversation. Explicitly separated convenient access from the interaction quality. Value was often not that AI provided a better interview, but that it made the interview easier to access.
Participants described another benefit reduced social pressure. As one participant explained: It's not the same pressure as having a human look at you while you're trying to formulate a response. For some the absence of a visible human reaction felt more relaxed and allowed greater directness even about mild embarrassment or frustration.
AI doesn't simply make the interview less social it removes the social dynamic. So that can support candor, but it also removes a human moderator's ability to recognise discomfort or distress and respond appropriately. That's why I see the clearest fit in everyday, lower stakes situations.
Topics with reduced social pressure can support candour without creating a real need for human support. The boundary isn't simply sensitive versus non sensitive. Mild embarrassment may benefit from the absence of a visibly reacting person, but vulnerability or trauma is a different proposition.
The value is conditional speed and breadth for researchers, flexibility and sometimes less social pressure for the participant. AI moderation earns its place when it suits the research question, is appropriate for the participant context and can responsibly produce the kind of evidence we need.
Choosing the use case is only the first design problem. The underlying model also shapes the conversation. Three characteristics are particularly relevant and I saw all three of them across the sessions: Limited and uneven context In longer conversations a moderator may lose track, repeat a question and forget where it was heading Helpfulness and agreeableness: Ask several questions explain what the participant should interpret, lead them and praise and answer.
Third, plausibility over accuracy. The next question may fit the conversation but introduce a detail or interpretation that the participant never supplied. These aren't necessarily reasons to dismiss the method, but they are behaviours that we need to anticipate and design around.
That is why much of the judgement moves upstream. A human moderator can notice drift, reframe a confusing question, decide a topic has been explored far enough or stop when someone becomes uncomfortable. With AI moderation more of that needs to be anticipated before the session begins. Four levers shape the result: time and depth, the research goal, context and moderation guidance.
The first three of those levers set the session structure before it begins. The first is time and depth longer sessions aren't necessarily deeper. As conversations become more complex, a moderator is more likely to lose track, repeat itself or pursue an unproductive thread.
The second is the research goal each objective should be specific and narrow enough to explore properly in one session. Substantially different questions are best left for a separate study. Also make sure to interrogate assumptions. If a given goal presumes a behaviour or problem, or that a particular feature is the answer to that problem, the moderator may carry that into the questions and follow ups.
The third is context. The moderator needs enough background to understand the topic but more isn't necessarily better. I keep only what helps to conduct the interview. Control of all of this does depend on the platform. Askable lets me structure objectives, adjust depth and provide context although those controls evolve during the programme, which matters later.
The principles apply across platforms using them depends on the controls and transparency we're given. The fourth lever is moderation guidance or how the moderator conducts the conversation. Researchers may control this directly, influence it indirectly or leave it to the system. All of these principles should be familiar the difference is making them explicit.
Ask one open question at a time. Models often bundle several together. Focus on concrete recent experiences what someone did, not what they might do or what people generally think. Stay neutral don't praise, validate, explain or advise even that's a great point may shape what follows.
And don't assume human like warmth automatically creates trust. One participant I spoke to described AI trying to emulate a person: This trying to be my friend thing, comfortable. Clarity and honest responsiveness may build more trust than scripted empathy.
Finally, probe meaningful problems, workarounds, trade offs, emotions or risks once or twice then move on. The participants I spoke to reinforced these principles that it was easier for them to answer questions that were anchored in recent memories and that the AI should treat I've already answered this as a stop signal, not as an invitation to rephrase.
None of this is new interviewing practice. The challenge is also encoding it clearly enough for the system to follow. That leaves an important question: Did the moderator actually follow it? During the programme I was focused on reviewing what participants told us, so I went back and examined the moderator instead.
I flipped the lens. I treated the moderator as the subject of the review. What was it doing or failing to do across sessions? Reviewed over 1,000 sessions from 10 research rounds across the four headline dimensions shown here session conduct, question quality, probing behaviour and research coverage.
This review was AI assisted. A Claude code skill applied a defined framework to each transcript. Flag behaviours such as repeated or leading questions, missed signals, broken closures and leaked instructions and linked each flag to the evidence. The framework distinguished isolated slips from recurring problems.
A session won't be rated poor because of a minor imperfection in the conversation. Using AI to review AI has obvious limitations so I don't treat the scores as fact. The value is that I was able to review a large number of sessions more consistently and check each issue against the transcript.
It also wasn't a controlled experiment the studies varied, the platform changed and my review method evolved as well. Was looking for large patterns that showed up repeatedly rather than focusing too much on small differences. One of the clearest of these appeared when I compared the three rounds before Askable introduced structured objectives and depth controls with the seven rounds that followed. The most striking change was in research coverage or how successfully the moderator covered the study's objectives.
The colours here show full, partial and insufficient coverage across each round. Before structured objectives and depth controls only 20 six-thirty 8% fully covered the study the three lowest results. After the changes full crop coverage increased to between 6991% in every round. Most of the early sessions were partial so the moderator explored much of the study but missed or compressed an objective.
The pattern suggests that explicitly structuring each objective helped the moderator distribute its attention more reliably. Over the 10 rounds question quality was consistently strong and probing improved. Poor question quality ratings stayed between 08% so question construction was the main limitation. In the three pre framework rounds ten-twenty 2% of sessions received poor probing ratings.
After the changes poor fell to between zero-eight percent and the final three rounds were 90 seven-ninety 8% good. The moderator generally asked useful questions and became more reliable at following enough to meet each objective. The bigger remaining problem wasn't what it asked, it was whether it could run the sessions successfully.
Session conduct was the most significant and persistent issue. Poor ratings ranged from two to 41%, with severe problems returning after apparently strong rounds. And the issue persisted even though its severity was intermittent. These were failures in running the conversation loops, internal state leaks, unacknowledged freezes, turn taking gaps and broken endings.
The structured framework appears to have improved coverage and probing, but conduct was less responsive suggesting these failures involved session mechanics as well as research design. These rounds span changing versions of the platform, so it's not a claim that every session behaves this way or that each issue reflects the current product. The final two rounds are encouraging, but the 10 round view shows why a clean recent result isn't enough to include that the risk had gone away.
Two anonymised exchanges may help to make this category a bit more concrete they are examples, not representations of every session. Here the participant was waiting for the interview to continue. Instead the moderator narrated its time state and shut down instructions, then moved into a conventional closing statement.
The session ended without asking the final research question. Was intermittent not typical but it shows why conduct needs to be addressed separately. Useful questions don't guarantee a reliable interview. A second example shows the burden moving to the participant.
The participant recognised a stall and asked the moderator to continue. The moderator then repeated the participant's request back as though it were an answer then interpret the complaint as their priorities the participant had to correct it restate the request and recover the conversation.
The session still ended so a completion measure might miss this but the participant was doing the moderator's work. That raised a question the transcript couldn't answer what did it cost participants to keep these interviews going? To understand that I returned to the human moderated interviews with the 14 previous participants.
The main finding wasn't dramatic frustration it was how people seemingly absorbed the failures. 10 described coping, repeating or rephrasing, prompting the moderator, shortening responses or working around the expected behaviour. Seven took responsibility for keeping the session moving, recovering from dead ends, restarting or checking that it had been captured.
Yet 11 were willing to participate again. Only two described overt frustration and just one reached anger, early termination and permanent opt out. Visible escalation is only part of the story. In a live interview an experienced researcher might notice that quieter adaptation and response. Once removed from the session, we lose that opportunity.
Repair work can look like ordinary cooperation. What participants decide not to say leaves no trace. One participant was careful about what he said because he worried the AI would disappear down an unproductive tangent. That cost is invisible.
The session may look complete, but the participant has edited his account to fit what he believes the moderator can handle. It wasn't simply a shorter or less detailed answer a genuine line of reasoning that never entered the transcript at all. An absence of visible frustration may simply mean that correcting the moderator isn't worth the effort.
Research coverage can appear successful even though the participant has narrowed the evidence before it even reaches the transcript. This has changed how I assess the quality of an AI moderated research session. Coverage matters, but it isn't enough.
Did the participant compensate for the moderation? Does the account reflect what they wanted to say rather than what the system could handle? A robust session needs all three of these: research coverage, a reasonable participant experience and evidence that we can trust. At the beginning I borrowed the introduction that participants hear in an AI moderated research session.
There are no right or wrong answers. There may be no one correct way to use AI moderation, but there are better questions to ask of it. Not simply How many interviews can we run? But What kind of evidence is this method structurally capable of producing?
Not simply Did it complete the script? But it cover the objectives, Respect the participant? Produce trustworthy evidence? And not simply what can AI do now? But what research judgement must surround it. Automation doesn't remove the researcher's responsibility it moves their judgement upstream into the goal, context, boundaries and guidance.
It also creates a downstream responsibility review what happened, including how participants compensated. Some of this depends on researchers, some on what platforms expose and allow us to configure. A robust method requires us to test it, document failures and be clear about the standards platforms must maintain.
The interview may be automated. The judgement and responsibility for the quality of the research remains ours.
Technologies & Tools
- Claude Code
Concepts & Methods
- AI moderation
- Surveys
- Depth interviews
- Population prevalence
- Human-moderated interviews
- Open-text responses
- Qualitative comparison
- Asynchronous participation
- Moderation guidance
- Open-ended questions
- Neutral questioning
- Session conduct
- Research coverage
- Leading questions
- Structured objectives
- Depth controls
- Conversational repair
Organisations & Products
- Askable
- InsightStream













