You have the manuscript finished, the chat window open, and a reasonable question: can this thing review my paper before I send it to a journal? The two answers you normally get are both useless. One says yes, machines review papers now, paste it in and go. The other says no, it invents everything. Neither helps you decide what to do this afternoon.
The useful answer is that a general-purpose chat model does some parts of a review genuinely well, and other parts so badly that acting on them can cost you a paper. What it does well are the tasks that only require reading your own text carefully. What it does badly are the tasks that need knowledge it does not have and a disposition it was not built with.
I review for Q1 psychology journals, so this is written from the side of the desk that writes the reports. I also build a tool that does part of this work, linked further down the page, so weigh what follows against your own experience rather than taking it on trust. Below: what to trust, what to throw away, the exact instructions to give it if you are going to use it today, and the part no prompt fixes.
Yes for language, internal consistency and over-claiming. No for judgement. A chat model is very good at holding your whole manuscript in view at once and telling you where the text contradicts itself, where a claim outruns the design that produced it, and where a sentence takes three readings to parse. It is bad at telling you whether the work is worth publishing, where it belongs, and whether the literature you lean on exists.
Here is the same rule in a form you can apply while you work. Trust it on anything it can verify inside the document you gave it. Distrust it on anything that needs information from outside that document. Whether your abstract reports the same sample size as your method section is inside. Whether your target journal publishes this kind of study, and whether the paper you just cited is real, are outside.
That rule is not a slogan, it falls out of how the thing works. The model produces text that plausibly continues the text it was given. With your manuscript in the prompt, the plausible continuation stays anchored to your actual words, which is why it is reliable there. Ask about something absent from the prompt and the plausible continuation is a well-formed invention, and a well-formed invention about a reference looks exactly like a reference.
Four things come out consistently well. All four are things a careful reader with infinite patience would also catch, which is precisely why a machine is good at them.
This is the strongest use, and in psychology it is the one that saves manuscripts. Ask it to flag every sentence that states or implies a causal relationship and to name the design that would be needed to support it. You will find “the intervention improved wellbeing” sitting on a single-group pre-post design with no control, “predicts” used for a regression estimated on data collected at one time point, and a discussion that quietly upgrades an association in the results into a mechanism in the conclusion. It catches this because the mismatch is inside your own document: the method says one thing, the discussion says something stronger. That drift is what authors stop seeing after the fifth revision and what a reviewer opens the report with.
Absence has a shape. A results section of this kind usually contains effect sizes with confidence intervals, a justification for the sample size, a statement on how missing data were handled, exclusion criteria applied before rather than after looking at outcomes, and evidence that assumptions were checked. When one of those is not in your text the model can tell, because it has read enough papers of this shape to know the slot exists and is empty. Treat the output as a checklist rather than as criticism: some items will not apply to your design. But an omission it names is one a reviewer can name too.
Give it the manuscript with the tables and ask for one job: list every number that appears in more than one place and does not match. This is where the boring, embarrassing errors live. The sample size in the abstract that stopped matching the method after you dropped four participants. A p value written as .04 in the text and .004 in the table. A coefficient called significant in the discussion that is not significant in the table it points to. One caveat you must respect: it can also misread a number and flag a mismatch that is not there, so treat every flag as a place to look, never as a finding.
If you write from a lab where English is not the working language, this is probably the highest-value item on the list. Not because your grammar is bad, but because the editor screening your submission has minutes and will not spend them reconstructing a forty-word sentence with three nested clauses. Ask it to shorten sentences, cut the hedging to one layer and remove phrases that carry no information. Then read the result against your original, because fluent rewriting flattens precision: a carefully qualified claim about a subgroup becomes a confident claim about everyone, and “associated with” drifts to “leads to”.
None of these are bugs waiting for the next version. Each follows from how the tool is built and what it is optimised for, which is why the mechanism is more useful to you than the warning: if you know why it fails, you can predict where it will fail on your paper.
A citation is one of the most predictable shapes in academic writing. Two or three plausible surnames, a year, a title assembled from the vocabulary of your topic, a journal that publishes that topic, a volume and a page range. A model that generates likely text is extremely good at producing that shape, and producing the shape is not the same as retrieving a real record. Unless the tool is genuinely looking the reference up and showing you where it found it, assume the reference was written rather than found.
What makes this dangerous is that the invented ones are not obviously wrong. They name authors who really do work in that area, in journals that really do publish that kind of study, with titles you would believe. That is exactly the profile of a citation you skim past. So the rule has no exceptions: nothing enters your manuscript unless you have found the actual paper and read enough of it to know it says what you claim. Search the exact title in Crossref, PubMed or your library catalogue. If nothing comes back, treat the citation as non-existent until you have the actual record in front of you: a paper you cannot locate is one you cannot cite either.
Conversational assistants are shaped to be helpful and pleasant to talk to. For most tasks that is what you want. It is a disaster in a reviewer, because the entire value of a review comes from someone being willing to tell you what you did not want to hear. When you ask “is my discussion well argued?”, the question signals the answer you are hoping for, and the agreeable continuation is a yes with three compliments and one soft suggestion.
You can push against this by assigning a hostile role, banning praise and asking for grounds for rejection instead of feedback. It helps a lot and you should do it, but the pull is structural, so calibrate: if the output made you uncomfortable nowhere, it reviewed nothing. Watch for the second-order version too. Push back on a criticism and it will often withdraw it, because agreeing is the agreeable move again. That is not reconsideration, and it is never a reason to leave your paragraph as it was.
Many papers that die never reach a reviewer. An editor decides in a few minutes that the work does not belong here, and that decision runs on knowledge the model does not have: what this journal has published in the last two years, which debates its readers are having, how a scope statement is interpreted in practice rather than on the page.
Ask whether your paper fits a named journal and you get a fluent, confident answer assembled from the journal’s name and general reputation. It reads like analysis with nothing behind it, and believing it costs you a submission cycle you did not need to spend. Do this part yourself: read the last four to six issues and look for two papers comparable to yours in design and topic. If you cannot find them, it is not your journal. Then check your own reference list, because a journal you cite five or six times is usually where your conversation is already happening.
Paste a manuscript, ask for a review, and you get: consider adding more recent references, the discussion could be strengthened, ensure the limitations are acknowledged. Every one of those is true of every manuscript ever written, which is why they are worth nothing. The mechanism is the same as before. A vague request has one most-likely answer, and the most likely answer is the average review, which is generic by construction.
The test takes five seconds. Read the feedback and ask whether it would still make sense pasted onto somebody else’s paper, in another field, on another topic. If it would, delete it. Real criticism is not portable, because it points at your sentence, your table, your design decision.
Nobody has a secret prompt. What the instructions below do is remove the three conditions that make the output useless: no standard to measure against, no obligation to point at your text, and permission to be nice. Give it one job at a time, in its own message. A single request to review my paper is the least productive thing you can send.
The over-claiming pass. Paste your results and discussion: “List every sentence that states or implies causality, prediction or effectiveness. For each one, name the design feature in my method needed to support it, and say whether my method has it.” Then fix the language or the claim, but check the list before you argue with it.
The consistency pass. Paste abstract, method, results and tables together and ask for one thing: “List every number, sample size, degrees of freedom, p value or coefficient that appears in more than one place with different values, or that the text describes differently from the table. Give the location of each mismatch and comment on nothing else.” Then verify each flag with your eyes.
The desk-reject rehearsal. Paste only the title, abstract and first paragraph of the method, which is roughly what an editor reads: “You are an editor with forty submissions in the queue and five minutes for this one. Give me the three reasons you would reject it without sending it out for review, in the order you would think of them. Do not tell me anything you liked.”
Those three passes cost you an evening and are much better than nothing. What you end up with is a competent read of the surface of your manuscript by something that never gets bored. What you do not end up with is a review, and the difference bites hardest exactly when the stakes are highest.
Three limits survive every instruction on this page, which is why a good prompt is a floor and not a ceiling.
The first is the standard of your field. A model with no field measures your paper against the average of everything it has read across all disciplines, and that average sits well below what a Q1 psychology journal expects. Do not count on it volunteering that Cronbach’s alpha alone is thin evidence of reliability and that a reviewer is likely to ask for omega, or stopping you comparing latent means between groups without testing measurement invariance first, or objecting that a mediation estimated on cross-sectional data cannot be described in causal terms however carefully the sentence is built. Ask about any of those and you get a sensible answer, which is the trap: it only checks what you already knew to ask about.
The second is verification. A generator cannot audit itself, because checking a reference means going and looking, not thinking harder. Anything it tells you about the literature outside your document remains your job to confirm, one item at a time.
The third is stability. Run the same prompt twice and you can get two different sets of criticisms with no indication of which run was better. That is tolerable while drafting and corrosive when you are deciding whether the paper is ready, because what you really want to know is what it failed to notice, and silence is the one output you cannot interpret.
The gap is the same in all three cases. It can tell you what is in your manuscript. It cannot tell you what a reviewer in your field will refuse to accept. Closing that needs either a person who reviews in your area, or a tool with that standard written into it, and, wherever the literature is involved, references checked against real bibliographic databases instead of generated. In my own tool that reference check belongs to the paid report; the free one reads your manuscript, not your bibliography.
Two checks, both worth five minutes, both easy to skip when the deadline is Friday.
First, what happens to the text you upload. That answer lives in the provider’s terms of use and privacy policy rather than on its marketing page, and it can differ between free and paid versions of the same product, so read the current terms yourself rather than a summary somebody wrote last year. If your manuscript contains clinical data, identifiable material, or anything covered by an ethics approval or a data-sharing agreement, this check is not optional and may not be yours alone to make.
Second, what your target journal asks of you. Journals set their own rules on the use of AI tools by authors, and those rules differ between titles and change over time, so the only version worth trusting is the one in that journal’s instructions for authors today. Find it before you submit and do what it says, including any disclosure. If you are on the other side, holding a manuscript sent to you to review, that is a separate question with a separate answer: what you may do with someone else’s unpublished work is governed by the instructions the journal sent you.
It does part of the job well. It catches over-claiming and causal language your design does not support, notices standard reporting elements that are missing, finds numbers that disagree between your text and your tables, and improves your English. It cannot judge whether the work is publishable, whether it fits a particular journal, or whether the literature you cite exists. Use it on what is inside your document and verify anything that comes from outside it.
One that supplies a role, a standard and a ban on praise, and asks for one job at a time. Tell it to review as a reviewer in your field whose task is to find grounds for rejection, paste the journal guidelines and the reporting guideline for your design, require every criticism to quote the exact sentence it refers to, and ask for the five most serious problems ranked by likelihood of rejection. The quotation requirement does most of the work: if it cannot quote your text, the criticism was generic or invented.
Because a conversational assistant is shaped to be helpful and agreeable, and your question usually signals the answer you are hoping for. A reviewer starts from the opposite stance, looking for reasons to say no. Use this as a test rather than a complaint: if the feedback made you uncomfortable nowhere, it reviewed nothing. And when you push back on a criticism and it immediately withdraws it, that is agreeableness, not reconsideration.
Assume yes unless it is actually looking them up and showing you where. A citation has a highly predictable shape, so a model that produces likely text produces convincing citations without retrieving anything. The invented ones name real authors in the right area, in journals that publish that topic, with believable titles. Search the exact title in Crossref or PubMed before any citation enters your manuscript. If nothing comes back, treat it as non-existent until you have the actual record in front of you.
It can check what you reported: inconsistent numbers, degrees of freedom that do not match your stated sample size, missing effect sizes or intervals, a test described in the text differently from the table. It cannot verify that the analysis was appropriate without your data, and it will accept a badly specified model described in fluent prose. Treat every flag as a place to look, and recompute anything that matters.
That depends on the provider and on what is in your manuscript, and the answer is in the terms of use and privacy policy rather than on the marketing page. Check what happens to the text you submit and whether it can be used for training, and check the current version, since terms change. If the manuscript contains clinical or identifiable data, that check is mandatory and may involve your ethics approval or your institution.
Journals set their own rules on this and they are not identical, so read the instructions for authors of the journal you are targeting and follow what it asks, including any disclosure statement. Do that check close to submission rather than from memory of what some journal required a year ago. Using a tool to improve your own writing and using one to generate content you cannot vouch for are different things, and the second is where authors get into trouble.
That is a different question from using it on your own paper, and you should not answer it by analogy. A manuscript sent to you for review is someone else’s unpublished work, handed over in confidence, and what you may do with it is set by the journal that invited you. Read the instructions the journal sent to reviewers before you paste any part of it into any tool.