What reviewers check in nursing research papers

Nursing journals review differently from general psychology journals, and authors moving between the two are frequently caught out. The statistical bar is not necessarily higher, but the procedural bar is: reporting guidelines are enforced rather than suggested, the distinction between research and quality improvement is taken seriously, instruments are expected to have validity evidence in your population rather than in the population where they were developed, and a paper without usable implications for practice will be rejected however clean the analysis.

The reviewers themselves are usually a mix. A clinical academic who will read your intervention description asking whether a ward could actually deliver it, and a methodologist who will check whether you accounted for patients being nested within units. Satisfying one and not the other is the standard reason a nursing manuscript gets major revisions.

This page sets out what each of them looks for: the guideline the journal will hold you to, the clustering problem that is nearly universal in ward-based data and nearly always ignored, what a translated scale needs before it can be used, how much intervention detail counts as enough, the qualitative reporting expectations, and why the implications section is a genuine gate rather than a formality.

The reporting guideline is checked first, and it is checked item by item

Most nursing journals of any standing endorse the EQUATOR network and require the completed checklist for your design as a submission file. Reviewers are frequently sent the checklist alongside the manuscript and asked to verify it, which means missing items are found rather than overlooked. This is the single largest procedural difference from psychology journals, where guidelines are named in the instructions and rarely enforced.

Match the guideline to the design, and note that authors get this wrong surprisingly often: a quasi-experimental ward-level intervention submitted with a CONSORT checklist, or a service evaluation submitted as though it were a trial.

Research or quality improvement: the distinction reviewers police

A great deal of nursing scholarship originates in service improvement: a new handover protocol on one unit, a pressure ulcer bundle, a change to discharge documentation. This work is publishable and valuable, and it is not automatically research. The distinction matters for three reasons reviewers will raise.

It determines the reporting guideline, SQUIRE rather than CONSORT or STROBE, and SQUIRE asks for things a trial report does not: the local problem and its context, the rationale for the intervention linking it to a theory of change, the iterative cycles of testing, and reflection on what would transfer elsewhere. It determines the ethics pathway, since many jurisdictions treat service evaluation as exempt from research ethics review while requiring institutional approval, and a manuscript claiming exemption without saying under what local rule invites a query. And it determines what can be claimed: an uncontrolled before-and-after change on one ward supports a claim about that ward, and the temptation to describe it as evidence of effectiveness is what draws the sharpest comments.

State early and explicitly which one your paper is, name the approval or exemption and its source, and use the matching guideline. Ambiguity here is a common cause of an otherwise avoidable rejection.

Clustering: the statistical issue that is nearly universal here

Nursing data is almost always nested. Patients within wards, wards within hospitals, nurses within units, students within cohorts, residents within care homes. People in the same unit share a manager, a staffing ratio, a skill mix, a patient population and a culture, so their outcomes are correlated in a way that ordinary regression assumes away.

The consequence is not cosmetic. Treating clustered observations as independent underestimates standard errors for cluster-level predictors, sometimes severely, which inflates the Type I error rate and turns noise into significant findings. A study reporting that staffing model predicts burnout across 12 units with 340 nurses, analyzed as 340 independent observations, has an effective sample size much closer to 12 for that predictor than to 340.

What reviewers expect to see: acknowledgement of the clustered structure, an intraclass correlation coefficient reported for the primary outcome so readers can judge how much clustering there is, and an analysis that accounts for it. Multilevel models are the standard answer where you have enough clusters, roughly 30 or more, and cluster-robust standard errors or generalized estimating equations are reasonable alternatives with fewer. With very few clusters, small-sample corrections such as Kenward-Roger become necessary and should be named.

If your design is a cluster randomized trial, the requirements are firmer still: the sample size calculation must include the design effect based on an assumed intraclass correlation with its source, and reviewers do check whether the assumed value is justified or invented. Failure to account for clustering in the power calculation and in the analysis is one of the recognized recurring errors in this literature.

Instruments: alpha from 1998 is not validity evidence for your sample

Nursing research uses a large number of instruments developed elsewhere, often in another language and another health system. Reviewers in this field are attentive to what happens in that transfer, more so than in many psychology journals.

The expectations for a translated instrument are specific. A forward and back translation process with a description of who performed it and how discrepancies were resolved, an expert panel review for content and cultural relevance, pretesting with members of the target population, and evidence that the adapted version measures what the original did. Citing a translation published elsewhere is fine, but cite the validation study, not just the translation.

The reliability point is smaller but appears in almost every review. Report the internal consistency computed in your own sample, not only the value from the original development paper, and report it for each subscale you use rather than for the total only. Where the scale is used to compare groups, expect a question about whether it functions equivalently across them, which is a measurement invariance question and worth reading about separately if your central claim is a group difference.

Two further checks that catch authors out. Permission and licensing: several widely used instruments in this field require registration or payment, and journals increasingly ask for a statement. And modification: if you dropped items, changed the response format or shortened the scale, the psychometric evidence from the original no longer applies to what you administered, and the reviewer will say so. A modified instrument needs its own factor structure and reliability evidence reported from your data.

Intervention description: could another unit deliver this?

This is where the clinical reviewer earns their place, and it is the most common substantive weakness in nursing intervention papers. A description reading “nurses in the intervention group received an educational programme on delirium prevention” is not a description. It cannot be replicated, evaluated or implemented, and the whole point of publishing it is that someone else should be able to do those things.

The TIDieR items are the checklist to write against, and they map onto the questions a reviewer will ask in almost the same order.

Sampling, response rates and outcome choice

Workforce surveys of nurses have a specific credibility problem that reviewers understand well: the people most affected by workload and burnout are the least likely to have time to complete a survey about workload and burnout. This makes non-response a substantive threat rather than a formality, and low response rates in this literature are common enough that reviewers expect them to be addressed rather than reported and passed over.

What helps: a clearly defined denominator, so the response rate means something; a comparison of respondents with the eligible workforce on any variables available from staffing records, such as grade, unit and shift pattern; and an explicit statement of the likely direction of non-response bias. Reporting that night-shift and bank staff are underrepresented, and reasoning about what that does to your estimate, is worth more than any statistical adjustment.

Outcome choice is the other recurring question. Nursing journals are interested in outcomes that are sensitive to nursing care, and reviewers will ask why you chose the one you did and whether it can plausibly respond to the intervention within the study period. Knowledge and self-reported confidence are acceptable as proximal outcomes but weak as endpoints; a paper whose only outcome is a knowledge test after an education session will be asked what happened to practice. Where patient outcomes are available, from incident reporting systems, clinical records or routinely collected quality indicators, using them substantially strengthens the paper, and reviewers know which data exist in most health systems.

A related check concerns timing. Outcomes measured immediately after an education intervention and never again support a claim about immediate effects only. If you have follow-up, report attrition at each point; if you do not, say so plainly rather than describing the immediate result as sustained.

Qualitative work: reflexivity and the insider problem

Qualitative research is central to nursing scholarship and reviewed with corresponding seriousness. Beyond the COREQ items, three things attract attention consistently.

The first is researcher positioning. COREQ asks who conducted the interviews, their credentials, occupation, gender, experience and training, and what relationship existed with participants before the study. This is not bureaucratic. A ward sister interviewing nurses she supervises, or a clinician interviewing her own patients, sits inside a power relationship that shapes what participants say, and reviewers expect it named and discussed rather than omitted. Insider status is not disqualifying and is often an asset; concealing it is the problem.

The second is analytic transparency. Which method, by name and with the reference for the specific version you followed, since thematic analysis in particular covers several distinct approaches with different epistemological commitments. How coding proceeded, who was involved, and how disagreements were handled. Note that agreement statistics such as kappa are contested in interpretive traditions and expected in others, so follow the convention of the method you named rather than adding a coefficient to look rigorous.

The third is the relationship between data and claims. Reviewers look for enough verbatim material to judge the interpretation, attributed to distinct participants rather than drawn repeatedly from the two most articulate ones, and for themes that are analytic rather than topic summaries. A theme called “communication” is a category; a theme making a claim about what participants experienced is a finding.

Implications for practice: a real gate, not a formality

Many nursing journals require a short structured statement of what the paper contributes and what it means for practice, sometimes as a separate box with a strict word limit. It is read closely, by editors as well as reviewers, and vague versions are sent back.

What fails: “these findings highlight the importance of communication in nursing practice”, “further research is needed”, “managers should consider the wellbeing of staff”. These could be attached to any paper in the field, which is precisely the objection.

What works is an implication at a level someone can act on, with the constraint attached. Who does what differently, at what point in a care pathway, and what your data does and does not establish about it. “Units using a paper-based handover recorded a median of 2.4 more omitted items per handover than those using the structured electronic template; where an electronic system already exists, adding the structured template requires configuration rather than new procurement, and our data support that change for medical-surgical units of this size, though not for critical care, which was not represented in the sample.” That is specific enough to be wrong, which is what makes it useful.

The same discipline applies to the contribution statement. Say what was not known before this study and what is known now, in one sentence each. Reviewers use these two sentences as a shortcut to judge whether the paper adds anything, so writing them carelessly is expensive.

Frequently asked questions

Which reporting guideline does my nursing study need?

Match it to the design: CONSORT for randomized trials, with the cluster extension if you randomized wards or clinics, STROBE for observational studies, PRISMA 2020 for systematic reviews, COREQ or SRQR for qualitative work, SQUIRE 2.0 for quality improvement, and CHERRIES for online surveys. Add TIDieR whenever an intervention is described. Most nursing journals require the completed checklist as a submission file and reviewers verify it item by item.

Is my project research or quality improvement?

Broadly, research aims to produce generalizable knowledge and quality improvement aims to change a local system, though the boundary is genuinely blurred and jurisdictions define it differently. The practical consequences are concrete: SQUIRE 2.0 rather than CONSORT or STROBE, a different ethics pathway, and a narrower set of claims. State which one your paper is, name the approval or exemption and the rule it rests on, and use the matching guideline.

Do I have to use a multilevel model if my data comes from several wards?

You have to account for the clustering somehow, and a multilevel model is the usual way when you have roughly 30 or more clusters. With fewer, cluster-robust standard errors or generalized estimating equations are defensible, and with very few clusters small-sample corrections should be named. At minimum, report the intraclass correlation for your primary outcome so readers can see how much clustering exists. Ignoring it inflates significance for exactly the unit-level predictors nursing papers usually care about.

Can I use a scale that was validated in another country and language?

Yes, with evidence about the transfer. Reviewers expect a described forward and back translation, expert review for cultural relevance, pretesting in the target population, and psychometric evidence in your sample rather than only from the original development study. Cite the validation of the adapted version, not just its translation, and if you modified items or the response format, report the factor structure and reliability from your own data, because the original evidence no longer applies.

My survey response rate was low. Will the paper be rejected?

Not automatically; low response rates are common in nurse workforce research and reviewers know why. What matters is that you define the denominator clearly, compare respondents with the eligible workforce on whatever variables staffing records make available, and state the likely direction of the resulting bias. A reasoned account of who is missing and what that does to your estimate is far more persuasive than a high rate reported without any non-response analysis.

How much detail does my intervention description need?

Enough that another unit could deliver it. Write against the TIDieR items: what was delivered, by whom with what training, how, where, how often and for how long, what was tailored or changed during the study, how fidelity was assessed and what it showed, and what the comparison group actually received. “Usual care” needs describing too, because it varies between sites and determines what your effect size means.

What makes an implications for practice section acceptable?

Specificity and a stated boundary. Name who would do what differently, at what point in the care pathway, and say what your data supports and what it does not. Statements that could be appended to any paper in the field, such as the importance of communication or the need for further research, are the ones sent back. Aim for an implication concrete enough that a reader could disagree with it.

How does it compare to the other tools?

You may also find this useful