Written and maintained by CASRAI Editorial Board
Last updated
Demand characteristics are cues in an experimental or survey setting — the instructions, the setting, the measures used, the order of tasks, or something a researcher says or does — that let a participant guess what the study is actually testing, and then, consciously or not, adjust their behavior to fit (or occasionally to defy) that guess. The term comes from Martin Orne’s 1962 paper On the Social Psychology of the Psychological Experiment, published in American Psychologist (17(11), 776–783), after a preliminary version presented in 1959. Orne’s core claim was uncomfortable for experimental psychology at the time: a participant is not a passive source of data. They are an active interpreter of the situation, and once they think they know what a study is “about,” that interpretation becomes part of what produces the result — not noise around it.
This matters for design, not just theory. A significant finding driven partly by participants performing the hypothesis they inferred is not a false result in the sense of being fabricated, but it is not a clean measurement of the effect under study either. It is a confounding variable hiding inside compliance. Because the confound tracks the hypothesis itself, it tends to inflate effect sizes in the direction the researcher expected — which is exactly the direction that gets published, replicated in citation networks, and taught in the next generation of methods courses, making it a quietly persistent threat to internal validity.
Where demand characteristics come from
Orne identified several distinct channels through which a study’s purpose leaks to participants, and later methods writing has kept largely to this structure:
- Prior information. Participant pools are rarely blank slates. Course-credit pools, professional panels, and repeat-participant registries all generate rumor: what a study “really” measures, what the trick question is, what the researcher is hoping to see. This is a known problem for university subject pools and for online panels alike.
- The setting itself. A room with a one-way mirror, a consent form describing “a study of moral judgment,” or a survey platform’s category label can telegraph the topic before a single item is presented.
- The order and structure of the procedure. Which measures come first, how items are grouped, and how many trials belong to each condition all give a attentive participant something to infer from, independent of the content of any single item.
- Explicit or implicit communication from the researcher. A leading tone of voice, a raised eyebrow at an “unexpected” answer, or over-detailed instructions can all signal expectation, even when the written protocol is neutral.
How participants respond: four roles
Not every participant who forms a guess responds the same way. Weber and Cook’s 1972 review of the demand-characteristics literature organized the range of responses into four roles that are still the standard shorthand in methods teaching:
- The good-participant role. The participant tries to confirm what they believe the hypothesis to be, out of a desire to be helpful or to see the study “work.” This is the role Orne’s original work worried about most, because it inflates the very effect the researcher is looking for.
- The negativistic (or “screw-you”) role. The participant deliberately behaves in the opposite direction, sometimes out of resentment at being studied, deceived, or bored — sabotaging the data in the other direction.
- The apprehensive role. The participant answers in whatever way seems least likely to make them look bad, prioritizing self-presentation over accuracy — the mechanism this role describes overlaps heavily with social desirability bias, though the two are not identical: demand characteristics require a guess about the specific hypothesis, while social desirability operates even when no hypothesis has been inferred.
- The faithful role. The participant tries to follow instructions exactly as given and to ignore any inferred hypothesis. Weber and Cook flagged this as the hardest role to actually occupy, since noticing a demand characteristic and then successfully behaving as though you had not noticed it is a genuinely difficult cognitive task.
Because these four roles pull in different directions — two toward confirming the hypothesis, one against it, one trying to cancel the whole effect out — the net bias in a given dataset is not automatically “everyone tried to help the researcher.” It depends on the population, the topic’s sensitivity, and how transparent the manipulation was, which is part of why demand characteristics are hard to reason about from theory alone and are better handled through design and detection.
Demand characteristics vs. the neighboring threats
Several related terms get used interchangeably in casual writing even though they describe different mechanisms. CASRAI’s observer effect guide covers the full family in more depth; the distinctions that matter most for demand characteristics specifically are:
- Demand characteristics vs. the Hawthorne effect. The Hawthorne effect is behavior changing simply because a participant knows they are being watched or studied, independent of any specific guess about the hypothesis. Demand characteristics require the extra step of inferring what the study is testing and then reacting to that specific inference.
- Demand characteristics vs. observer/experimenter bias. Demand characteristics describe the participant reacting to the study; observer bias describes the researcher’s own expectations distorting how data is collected, coded, or scored. A single study can have both problems at once, from opposite ends of the interaction.
- Demand characteristics vs. response bias generally. Response bias is the broader umbrella for any systematic, direction-consistent distortion in self-report data. Demand characteristics are one specific mechanism that produces response bias, tied specifically to hypothesis-guessing rather than, say, a poorly worded item or acquiescent responding.
Detecting demand characteristics after the fact
Because demand characteristics operate through what a participant believes the study is about, the most direct detection method is simply asking them. This is typically done through a post-experimental inquiry or funnel debriefing: a structured set of questions, administered after data collection but before full debriefing, that moves from general (“What did you think this study was about?”) to specific (“Did anything about the instructions or the setting suggest what we were testing?”) without tipping the participant off to the real hypothesis mid-interview. Orne’s original methodological program treated this inquiry as a required companion to any study where demand characteristics were a live concern, not an optional add-on.
A manipulation check — a direct item confirming that an experimental manipulation was perceived the way the researcher intended — serves a related but distinct purpose. A manipulation check confirms the independent variable landed; a post-experimental inquiry into demand characteristics asks whether the participant’s response was contaminated by guessing the point of the study in the first place. Well-designed studies with a plausible demand-characteristics risk typically run both, then report what fraction of participants correctly guessed the hypothesis and whether excluding or statistically adjusting for hypothesis-aware participants changes the result.
Designing against demand characteristics
No single technique eliminates demand characteristics; each of the standard mitigations trades some risk reduction against a real cost, usually to feasibility, ecological validity, or research ethics.
- Deception about the specific hypothesis (not about participation itself), paired with full debriefing afterward, remains the most direct control — if a participant cannot correctly infer the hypothesis, they cannot selectively perform it. This raises real ethical obligations around debriefing and, in some fields, requires specific IRB justification for withholding the study’s true purpose.
- Single- and double-blind designs reduce demand characteristics by keeping the participant (single-blind) or both the participant and the person administering the study (double-blind) unaware of condition assignment, cutting off one of the main channels — explicit or implicit cueing from the researcher — described above.
- Unobtrusive or implicit measures (reaction-time tasks, behavioral traces, physiological measures) are harder for a participant to consciously manipulate toward a guessed hypothesis than a direct self-report item is, though they raise their own interpretive and consent questions.
- Between-subjects experimental design limits how much of the manipulation any one participant is exposed to, which limits how much material they have available to infer a hypothesis from, compared with a within-subjects design where the same participant sees every condition.
- Standardizing experimenter contact — scripted instructions, recorded or computer-delivered materials, minimizing improvised interaction — closes off the explicit/implicit communication channel without requiring full deception.
These design controls sit alongside, rather than replace, a properly powered and randomized control group design: demand characteristics are a validity threat that operates on top of whatever the basic design already controls for, not a substitute for basic experimental structure.
Reporting demand characteristics honestly
When a study cannot fully rule out demand characteristics — which describes most research involving human participants who are aware they are being studied — the defensible move is to name the risk explicitly in the limitations section rather than omit it or assert the measure was unaffected. A specific, well-scoped limitation (“38% of participants correctly guessed the study’s purpose in post-experimental inquiry, which may have inflated the observed effect in the predicted direction”) is more useful to a reader, and more defensible to a peer reviewer, than either silence or a generic disclaimer that demand characteristics “cannot be entirely ruled out.” Reporting the actual post-experimental inquiry results, where they were collected, lets a reader judge the size and direction of the residual risk for themselves.
Frequently Asked Questions
Are demand characteristics the same thing as experimenter bias?
No. Demand characteristics describe the participant reacting to cues about the hypothesis; experimenter (observer) bias describes the researcher’s own expectations distorting data collection, coding, or scoring. They can co-occur in the same study but involve different people and different mechanisms.
Can demand characteristics ever help a study rather than hurt it?
Not in the sense of improving validity. Even when a “good-participant” response happens to push results toward what turns out to be a true effect, the study cannot distinguish that from participants performing a false hypothesis — the confound is present either way, and the result is not evidence for the effect independent of the confound.
Does a larger sample size fix demand characteristics?
No. Demand characteristics are a systematic, direction-consistent bias, not random noise. A larger sample reduces the standard error around a biased estimate; it does not reduce the bias itself, the same limitation that applies to response bias generally.
Is deception always necessary to control for demand characteristics?
No. Blinding, unobtrusive measurement, standardized administration, and between-subjects design all reduce the risk without concealing the study’s basic purpose from participants, and each carries a smaller ethical burden than deception paired with debriefing.
How do I report demand characteristics if I didn’t run a post-experimental inquiry?
State plainly that hypothesis-guessing was not assessed and that the study therefore cannot rule out this specific threat to internal validity, rather than implying the absence of a formal check means the risk was absent.








