Let me start with what works against me: ChatGPT is free, you already have it open, and with a decent prompt it produces useful observations about your manuscript. Any comparison that pretends otherwise is lying to you, and you can disprove it in two minutes.
Here is the second thing working against me: my tool is a language model too. I am not selling different magic. What changes is not the technology but what it is asked to do, which standard it is measured against, and what gets verified before you are shown the result.
This page compares the two properly: what you get when you ask a general-purpose model to play reviewer, where that breaks down in psychology and health sciences, and the cases where paying for anything would be a waste because the chat is enough.
With a decent prompt, a general-purpose model catches real things: that the introduction never states the gap you are filling, that your stated aim and your research question do not match, that the discussion restates results instead of interpreting them, that a conclusion is not supported by any of your tables. Those are legitimate objections and you get them for free.
It also does something I do not: it converses. You can ask why it said what it said, ask it to rewrite a paragraph, propose three ways to phrase a limitation, translate your response to reviewers. No specialised tool has that flexibility, mine included.
If your manuscript is at an early stage, if you are still deciding what story you are telling, if what you need is an interlocutor rather than a verdict, the chat is the right tool and spending money on anything else is premature.
The problem is not that the model is bad. It is that a reviewer with no field evaluates against the average of everything it has read, and the average across all disciplines sits well below what a psychology journal demands. When you ask it to review, it answers with what is true of any paper, because that is all it can assert without knowing your area.
In practice that means it misses precisely the objections that sink manuscripts in this field. It will not tell you that reporting Cronbach’s alpha alone as evidence of reliability is no longer enough and that a reviewer will ask for McDonald’s omega. Or that comparing latent scores between men and women without testing measurement invariance invalidates the comparison. Or that a mediation estimated on cross-sectional data cannot be described in causal terms, however elegantly the sentence is written. Or that effect sizes with confidence intervals are missing, which is the first thing a trained reviewer looks for.
There are three further problems, formal but heavy. The first is consistency: two runs of the same prompt can return different criticisms, and you have no way of knowing which run was the good one. The second is agreeableness: a conversational assistant is optimised to be helpful, and that pushes it towards validating what you show it, especially when your question already signals that you expect a yes. The third is that it can invent references with impeccable formatting, a well-documented risk in general-purpose models, and one fake citation in a response to reviewers does damage that is hard to undo.
Saying it again because it matters: underneath there is a language model, exactly as in the chat. The difference is four layers built on top, layers you would otherwise have to build yourself.
The first is the standard. The role, the report sections, the severity scale and the things that must be checked in a psychology or health-sciences manuscript are written by someone who reviews for these journals. None of it depends on you guessing the right prompt.
The second is the structure of the process. The Pro report is not a single pass: there are two independent reviewers, an editor who arbitrates between them, and a desk-reject screen beforehand. It is a deliberate imitation of the real journal process, so that a criticism never rests on one reading.
The third is verification. Every reference in the manuscript is checked one by one, up to sixty, against Crossref, OpenAlex, PubMed, Semantic Scholar and DOAJ, and the report is checked before delivery so it does not go out with invented citations. That layer exists precisely because the hallucination risk is real and cannot be fixed by asking the model not to hallucinate.
The fourth is the shape of the output: a report with an overall grade and grades per section, every criticism carrying a verbatim quote and its fix, a reanalysis plan with the code written in the software you actually used — the one your analysis section names: R, SPSS, jamovi, JASP, Mplus, Stata — and a shortlist of Q1 and Q2 journals where the work fits. That is a document you can work from, not a chat thread you have to reread.
I have no interest in you misusing a free tool so mine looks better. If you are going to ask a general-purpose model to review your paper, these four things improve the result a great deal:
It does not pay if you are on an early draft, if your field is not psychology or health sciences, or if what you need is somebody to think out loud with. For that the chat is a better tool than mine, and it is free.
It pays when you are about to submit and the cost of being wrong stops being measured in euros and starts being measured in months. A lost review round at a Q1 journal is three to six months; a desk rejection is two weeks plus the resubmission. Against that, spending the price of a coffee on a full report is not really a financial decision.
And there is a middle path, which is the one I recommend: the free diagnostic report, no account, no card. Upload it, compare it against what the chat gave you with your best prompt, and keep whichever told you something you did not know. If the chat wins, you have saved money and I have work to do.
It can give you a useful critical read, especially of structure and argument, and with a good prompt it catches real problems. What it cannot do is apply the specific standard of a psychology or health-sciences journal, because it evaluates against the average of all disciplines. It works for the draft; for the submission it falls short on the field-specific objections.
It is, and saying otherwise would be dishonest. Underneath there is a language model, exactly as in the chat. What changes is four layers built on top: a standard written by a reviewer for Q1 psychology journals, a process with two reviewers and an arbitrating editor, verification of every reference against five databases, and a structured report instead of a conversation.
That is a question you should ask of any tool, mine included, and the answer lives in each provider’s terms of use and privacy policy rather than on its marketing page. Check what happens to the text you upload and whether it is used for training. If your manuscript contains clinical data or sensitive material, that check is not optional.
Because a conversational assistant is optimised to be helpful and agreeable, and that pushes it towards validating what you show it, especially when your question already implies you expect a yes. A real reviewer starts from the opposite position: looking for reasons to reject. If you use the chat, explicitly tell it to say nothing positive and to quote your text in every criticism.
One that supplies role, standard and a ban on praise: act as a reviewer for a specific journal in your field, with the author guidelines and the relevant reporting guideline pasted into the message, with the explicit brief of finding grounds for rejection, and with an obligation to quote the exact sentence from your manuscript in every objection. If it cannot quote it, it invented it.
By checking them one by one in Crossref, PubMed or OpenAlex before they enter your manuscript. There is no shortcut: asking the model not to invent guarantees nothing. That check is exactly what the Pro report automates, verifying up to sixty references against five databases.