What reviewers check in clinical psychology papers

Clinical psychology reviewing has its own set of reflexes, and they are not the ones you learn writing for general psychology journals. A reviewer for a clinical journal reads your method looking for a specific list: how the sample was diagnosed, what the comparison condition actually was, who delivered the treatment and whether their effects were modeled, whether anything happened to participants that you have not reported, and whether the change you found means anything to a person rather than to a significance test.

These checks exist because the field has spent two decades absorbing uncomfortable findings about its own literature: effect sizes that shrink dramatically when the control condition is active rather than a waitlist, treatment differences that track who developed the treatment, harms that go unreported in trials of psychological interventions, and samples described as clinical that were undergraduates above a questionnaire cut-off.

This page goes through the checks in the order a reviewer applies them, from how your sample was defined to what you claim in the discussion, and explains what evidence satisfies each one. It is written for treatment studies, symptom and mechanism research, and psychopathology work using clinical samples.

How the sample was diagnosed, and whether it is a clinical sample at all

This is the first check and it disqualifies more manuscripts than any other. A group of participants scoring above a cut-off on a self-report questionnaire is not a group of people with a disorder, and describing them as such is the error reviewers react to hardest. A PHQ-9 score of 10 or above identifies people at elevated risk of depression with a false positive rate that is substantial in general population samples; it is a screening result, not a diagnosis.

What satisfies the check is a structured or semi-structured diagnostic interview administered by trained assessors: the SCID-5, the MINI, the ADIS for anxiety disorders, the K-SADS for children and adolescents. Report which instrument and version, who administered it and what their training and supervision were, and the interrater reliability with the number of interviews that were double-coded and the coefficient obtained. Reliability stated without the number of cases it was computed on is not informative and reviewers ask for the number.

If you did not diagnose, the fix is linguistic and it is genuine rather than cosmetic. Call the sample what it is: participants with elevated depressive symptoms, an analogue sample, a community sample selected on symptom severity. Then be careful in the discussion not to slide back into disorder language, which is where the slippage usually occurs. Analogue samples are legitimate for many research questions, and papers using them are published constantly. What is not accepted is the analogue sample described as clinical.

Two associated reporting requirements. Comorbidity, because comorbid presentations are the rule rather than the exception in clinical samples and their distribution shapes what your results mean; report it, and say whether comorbid conditions were exclusion criteria. And medication status, including whether medication was stable, whether changes during the study were tracked, and how you handled them analytically. A psychotherapy trial that does not report concurrent pharmacotherapy will be asked about it.

What the comparison condition was, and what it licenses you to claim

The choice of comparator determines the meaning of your effect size, and reviewers read the effect size through it. This is now well enough established that reporting a large effect against a waitlist without acknowledging the comparator is treated as a naive claim rather than a strong result.

The pattern in the literature is consistent. Comparisons against waitlist controls produce the largest effects, and there is good evidence that waitlists function as more than a neutral baseline, potentially suppressing spontaneous improvement because participants know help is coming. Comparisons against treatment as usual produce smaller effects that vary with what usual care contains at your site, which is why usual care must be described in detail rather than named. Comparisons against an active, structurally matched control produce the smallest effects, and in adult psychotherapy for common mental disorders, differences between bona fide active treatments are typically modest.

What to do about this in your own paper: name the comparator precisely, describe what it consisted of including contact time, and interpret your effect size against studies using the same class of comparator rather than against the field average. If your control is a waitlist, say in the discussion that the estimate is not comparable to active-comparator trials and is expected to be larger for that reason. Volunteering this is far better than having it raised.

Reviewers also examine structural equivalence between conditions: equal contact time, equal therapist attention, equal credibility and expectancy. Where you measured treatment credibility and expectancy, report it by condition, since a difference there is an alternative explanation for a difference in outcome. Where you did not, expect a limitation to be requested.

Therapist effects and researcher allegiance

Two design features specific to psychotherapy research are checked by any reviewer with trials experience, and both are frequently absent from submitted manuscripts.

Therapist effects come first. Patients treated by the same therapist are not independent observations: a meaningful share of outcome variance in psychotherapy trials is attributable to the therapist. Analyzing patients as independent when they are nested within a small number of therapists underestimates standard errors and overstates the precision of the treatment effect. The expected treatment is to model therapist as a random effect, or at minimum to report how many therapists delivered each condition, how many patients each treated, and the intraclass correlation for therapist. Where therapists delivered both conditions, that is a design strength and should be stated, since it controls therapist effects across arms.

Allegiance is the second and it is more delicate. Effect sizes in comparative psychotherapy trials correlate with the theoretical allegiance of the research team, and reviewers are alert to trials in which the developer of the treatment designed the study, trained the therapists, supervised delivery and rated the outcomes. This is not an accusation of misconduct, it is a recognized source of bias with plausible mechanisms including differential competence in delivering the two treatments, and it is best handled by disclosure and design. State the team’s relationship to the treatments, and describe the safeguards: independent outcome assessment, blinded raters with a check on whether blinding held, treatment manuals for both conditions, supervision by an expert in each approach, and fidelity ratings for both arms rather than only the treatment of interest.

Fidelity is worth its own sentence. Reporting that sessions were recorded and rated, giving the proportion rated and the resulting adherence and competence scores by condition, closes a check that otherwise stays open. Fidelity monitoring described in the methods with no results reported is a gap that reviewers specifically look for.

Attrition, missing data and the analysis population

Dropout in psychotherapy trials is substantial and frequently differential between conditions, which makes the analysis population a live question rather than a technical footnote.

Reviewers expect an intention to treat analysis as primary, including all randomized participants in the condition they were assigned to, with per-protocol or completer analyses reported as secondary if at all. A paper whose primary analysis is on completers will be asked to change it, because attrition in psychological treatments is often related to how the treatment is going, which makes completer analyses systematically optimistic.

Missing data handling should be stated explicitly, with the method named and the assumption acknowledged: multiple imputation or full information maximum likelihood assume data are missing at random conditional on the variables in the model, and that assumption is questionable precisely when people drop out because they are deteriorating. This is where a sensitivity analysis earns its place, shifting imputed values for dropouts in the unfavorable direction and reporting how large the shift has to be before the conclusion changes.

The reporting requirements around it are equally checked: a CONSORT flow diagram with numbers at each stage, reasons for dropout by condition, and a comparison of completers with non-completers on baseline variables. Differential dropout between arms should be tested and discussed rather than reported and left.

Clinical significance: the check that catches statistically clean papers

A result can be statistically significant, correctly analyzed and clinically uninformative, and clinical psychology reviewers are trained to notice the gap. A mean reduction of 1.8 points on the BDI-II is a real difference and is smaller than the measurement error of the instrument for an individual person.

The tools that answer this check are standard and reviewers expect at least one of them.

Deterioration and adverse events

Harm reporting in trials of psychological interventions has historically been poor, and this is one of the fastest-moving expectations in the field. Reviewers increasingly require, rather than suggest, that a treatment study report whether adverse events were monitored, how they were defined and collected, what occurred, and the proportion of participants who deteriorated reliably in each condition.

Silence is no longer read as an absence of harm; it is read as an absence of monitoring. If you collected nothing systematic, say so as a limitation rather than omitting the topic. If your population includes suicidality, expect a specific question about your risk protocol: what triggered an escalation, who assessed it, and what happened to participants who met the threshold. This applies to online and unguided interventions with particular force, since the monitoring problem there is structural.

Mechanism claims and the measurement of change

Papers claiming to identify how a treatment works face an additional layer of scrutiny, because the design requirements for a mediation claim are stringent and rarely met.

The requirements reviewers apply: temporal precedence, meaning the proposed mediator must be measured before the change in outcome rather than at the same post-treatment assessment; ideally repeated measurement of both mediator and outcome across treatment so that change in one can be shown to precede change in the other, which session-by-session or weekly assessment makes possible; and evidence that the treatment actually moved the mediator, which is often assumed and not demonstrated. A mediation model in which mediator and outcome are both measured at post-treatment cannot establish any of this, and saying so before the reviewer does improves how the paper reads.

Measurement of change raises its own checks. If you compare symptom scores across time, you are assuming the instrument measures the same construct in the same metric at each occasion, which longitudinal measurement invariance tests and response shift can violate, particularly after an intervention that changes how people think about their symptoms. Reviewers in this area increasingly ask for it. Floor and ceiling effects matter too: a sample selected for high symptom severity regresses toward the mean, and an uncontrolled pre-post design cannot separate that from treatment effect, which is one of the main reasons uncontrolled outcome studies attract heavy criticism.

Sample description and generalizability

Clinical journals have become notably stricter about how samples are described, and thin demographic reporting now draws comments that it did not five years ago.

What is expected: race and ethnicity reported with the categories used and how they were collected, socioeconomic indicators, education, and the recruitment source, since a community-advertised sample, a clinic caseload and an online recruited sample differ systematically in severity, comorbidity and treatment history. Report treatment history explicitly, because a sample of treatment-naive participants and a sample of people who have already failed two courses of therapy are different populations and yield different effect sizes.

Eligibility criteria receive similar attention. Exclusions for comorbid substance use, personality disorder, psychosis, active suicidality or medication changes are common and each one narrows the population your results describe, sometimes to the point where the sample bears little resemblance to routine practice. List them, and then say in the discussion what they mean for generalization to the clinical settings you are addressing. The gap between trial samples and clinic populations is a recognized issue in this field and reviewers respect authors who name it.

Finally, the framing check. Reviewers read the abstract and the final paragraph of the discussion against everything above, looking for claims the design cannot support: efficacy language from an uncontrolled study, recommendations for practice from a small analogue sample, mechanism claims from a post-treatment mediation model. Aligning those two pieces of text with your actual design is the cheapest revision available and the one most often left undone.

Frequently asked questions

Can I call my sample clinical if I used a questionnaire cut-off?

No. A cut-off on a screening instrument identifies elevated symptoms with a substantial false positive rate, not a diagnosis. Describe the sample accurately as an analogue or community sample selected on symptom severity, and keep disorder language out of the discussion as well as the method. Analogue samples are publishable for many research questions; the objection is to the mislabeling, not to the design.

Do I need a structured diagnostic interview?

If you want to claim a diagnosed sample, yes, and you should report which instrument and version, who administered it, their training and supervision, and the interrater reliability together with how many interviews were double-coded. If a structured interview was not feasible, report what was used, whether it was a clinician diagnosis from records or a validated screener, and describe the sample in terms that match.

Why do reviewers object to waitlist control groups?

Because comparisons against waitlists produce systematically larger effects than comparisons against active or usual-care conditions, and there is evidence that waiting for a promised treatment can itself suppress improvement. A waitlist-controlled effect size is therefore not comparable to an active-comparator one. Waitlist designs remain acceptable in many circumstances; what is not acceptable is reporting the resulting effect size without acknowledging what the comparator implies.

What are therapist effects and do I have to model them?

Patients treated by the same therapist have correlated outcomes, so a meaningful portion of variance sits at the therapist level. Treating patients as independent when they are nested within therapists underestimates standard errors and overstates precision. The expected treatment is to include therapist as a random effect; at minimum, report the number of therapists per condition, the number of patients each treated, and the intraclass correlation. Therapists delivering both conditions is a design strength worth stating.

What counts as clinical significance?

A statement about whether the change matters to a person rather than to a test. The accepted tools are the reliable change index with the associated criterion for movement into the functional range, a priori response and remission rates using referenced thresholds, comparison against a published minimal important difference for your measure, and number needed to treat. Reporting at least one of these alongside the group means is now close to standard in clinical journals.

Do I have to report adverse events in a psychotherapy study?

Increasingly yes, and silence is read as absence of monitoring rather than absence of harm. Report whether adverse events were systematically collected, how they were defined, what occurred, and the proportion of participants showing reliable deterioration in each condition. If nothing systematic was collected, state that as a limitation. Where suicidality is present in the population, expect a specific question about the risk protocol and what happened to participants who crossed the threshold.

Can I claim a mechanism from a post-treatment mediation analysis?

Not credibly. A mediator measured at the same assessment as the outcome establishes no temporal precedence, and cross-sectional mediation estimates are biased relative to the longitudinal quantities they stand for. What supports a mechanism claim is repeated measurement of mediator and outcome during treatment, showing change in the mediator preceding change in the outcome, together with evidence that the treatment moved the mediator. Without that, present the analysis as consistent with the hypothesized process and say so explicitly.

How does it compare to the other tools?

You may also find this useful