AI tools for peer review: what to check before you trust one with your manuscript

I review manuscripts for psychology journals, and I also build one of the tools this page is about, so read what follows with that in mind. What a product page will not give you is the shape of the problem: there are only three ways to get a critical read of your manuscript before you submit it, and each fails in a way built into what it is, not into how well it was made.

This is not a feature table. Feature tables about AI products go stale fast, which is why you will not find anybody else’s prices, limits or feature counts on this page. What does not date is the structure: what a general purpose model can and cannot know, what you inherit when a tool applies criteria somebody else wrote, and what you are buying when you pay a person.

Below: the three approaches and the limit baked into each, eight questions to ask before you upload anything, a twenty minute test using a paper whose flaws you already know, and, labelled as such, where my own tool sits.

Three approaches, not twenty products

Strip the branding away and every tool offering to review your paper is one of three things: a general purpose model that you instruct, a specialised tool applying criteria somebody else wrote, or a human who reads it and takes money, or a favour, for an opinion. They blur in practice, since a specialised tool has a model underneath, but they fail differently and predictably. Working out which one you are looking at tells you where to point your scepticism.

Approach one: a general purpose assistant that you instruct

A chat assistant of the kind you probably have open in another tab, told to act as a reviewer and criticise your paper. Cheapest route, and by far the most flexible. Its structural limit is having no criterion of its own: it answers with roughly the average of everything it has read, and that average spans every discipline there is, most of them more forgiving than a psychology journal about measurement. The standard has to come from you, in the prompt, so the tool is only as demanding as what you already know reviewers demand.

Approach two: a tool specialised in review

Somebody fixed the criteria in advance, so you get a structured report with the same sections, the same severity scale and the same checks every time. Output becomes comparable, which lets you run version three against version one and see whether you actually fixed anything, and the awkward questions get asked whether or not you knew to ask them. The cost is inheriting somebody else’s judgement in full, usually without seeing it written down: if they believe the important thing is clean English, you get a clean English report and a fatal design flaw goes unmentioned. So the first question is not what it costs, it is who wrote the criteria and whether that person is named.

Approach three: a human read

A colleague who owes you a favour, a statistician hired for a few hours, or a pre-submission editing service. This differs in kind: a person who can be asked why, who changes their mind when you push back, and who carries responsibility for what they told you. A human read also does the thing no automated tool does reliably, which is to tell you that the paper you wrote is not the paper you should have written from this dataset. The costs are money and time: these are professional services, billed at professional rates, and the turnaround is days or weeks rather than minutes. Ask for the price and the delivery date in writing before you commit to either.

How each one fails, mechanically

Failure modes beat feature lists: they follow from how the thing works, so you can anticipate them instead of meeting them in a rejection letter.

Invented references, and why asking nicely does not fix it

A language model produces the most plausible continuation of a text, and a citation is one of the most predictable strings in academic writing: surname, initial, year, a title of familiar words, a journal, a volume, a code shaped like a DOI. It can generate one that is perfectly formed and refers to nothing, with no internal marker separating a remembered reference from a constructed one. Telling it not to invent changes little, since it cannot check itself. The only fix is a query against a real index, so ask whether a tool looks citations up and where. With a general model that check is yours, in Crossref, PubMed or OpenAlex.

Agreement dressed up as assessment

The second failure is subtler because the output looks like a review: praise, three mild suggestions about expanding the discussion, a closing line about a valuable contribution. It reads as validation and it is mostly politeness. Count the objections that would change something if you acted on them. Not "consider strengthening the methods section", but "comparing latent means across groups needs a test of measurement invariance first". A report with none of the second kind flattered you.

It does not know your journal

Fit is one of the most common reasons manuscripts die before anyone reviews them, and it is the judgement automated tools are worst at. Whether your paper belongs in a journal depends on what it published in the last two years, what its editors are tired of receiving, and what is sitting in its queue now. None of that reaches a model unless you put it there, so a tool that hands you target journals without saying what evidence it used is guessing at the thing that matters most. Do that part yourself: open the last four to six issues and look for two papers comparable to yours.

The variance nobody mentions

Run the same manuscript through the same general model twice and you get two different reviews. Some of the difference is phrasing, some of it is a serious objection appearing in one run and not the other, which makes a single clean report weak evidence. A fixed rubric narrows the variance and reading the manuscript twice narrows it further, but nothing removes it. Treat one pass as one opinion.

Eight questions to ask before you upload anything

These work on all three categories, including the human one. If a tool makes any of them hard to answer, that difficulty is itself the answer.

1. Who set the standard, and is that person named?

Every review tool encodes somebody’s opinion of what matters in a manuscript, and the useful question is whose. Look for a named person with a traceable record in your field: someone who reviews for these journals, publishes in them, or edits for them. "Built with leading researchers" is not a name. Check the about page, then that name in Google Scholar or ORCID.

2. Does it quote your text, or talk about it in the abstract?

The fastest test of whether a review is real. A criticism that quotes the sentence it objects to can be checked, argued with and acted on; "the methodology could be strengthened" can be written without reading anything. Take the first output you get and count the criticisms containing a verbatim sentence of yours. Zero means you are holding a description of papers in general.

3. Does it verify the references, or repeat them?

Given how invented citations arise, this is not a refinement: a tool that recommends literature without checking it exists hands you a liability carrying your name, not its own. Ask whether references are checked against an index and which one, then verify by hand. Take three citations it produced and search them in Crossref or PubMed. One miss and you stop using it for citations.

4. Does it tell you what to fix, or only that something is wrong?

Diagnosis without treatment is where most automated review stops. "Your sample size is not justified" is true and you already suspected it; what you need is which justification your design calls for, what to compute, and what sentence goes into the manuscript. Read one full criticism and ask whether you could act on it tomorrow morning without consulting anybody. If not, it is a list of worries, not a plan.

5. What happens to your manuscript afterwards?

You are uploading unpublished work, sometimes with clinical material behind it. What matters is whether the file is stored, for how long, whether it trains anything, and whether you can have it deleted. This applies to every category, mine included. Read the privacy policy rather than the marketing page and search it for "train", "retain" and "delete"; if those words never appear, the silence is information.

6. Does it rewrite your paper, and what does that do to your authorship?

There is a real line between a tool that tells you a paragraph does not work and one that hands you a replacement paragraph. Cross it often and you are submitting sentences you did not write and cannot defend when a reviewer asks about them. Whatever your journal’s rules, and you need to read those rather than take my word for anything, the logic holds: responsibility for the text sits with people who can answer for it, and a tool cannot answer for anything. Look in the instructions for authors, usually under a heading on generative AI or research integrity.

7. Is the criticism specific to your field, or true of any paper?

Psychology and health sciences have objections that barely exist elsewhere and that sink manuscripts here: what counts as adequate reliability evidence, measurement invariance before comparing latent scores, causal language on a cross sectional design, effect sizes with intervals instead of a naked p value. A general read will not raise them because they are not general. Run a manuscript you know well and ask whether anything came back that a reader outside your field could not have said. If not, you bought a writing tutor.

8. Can you tell whether it read the whole thing?

Long manuscripts can get truncated, and a review of your introduction presented as a review of your paper buys false confidence about the part that gets you rejected. The fatal objections live in method and analysis, in the middle of the file. Plant the test yourself: put a checkable oddity deep in the results, a reported degrees of freedom that cannot match your sample size, and see whether it comes back.

The twenty minute test that settles it

Every claim on this page, mine included, is checkable in one sitting. The protocol needs a manuscript whose flaws you already know, and the best candidate is a paper of yours rejected after review, because you hold the reports and therefore the answer key. A paper you refereed yourself works almost as well.

Feed it in and compare. You are measuring how many real objections it recovered, how many it invented, and whether it raised anything the reviewers missed. Do not expect an exact match; two referees on one manuscript do not reproduce each other either. What you want is overlap on the serious design and analysis points, plus a few specifics nobody mentioned. Ten criticisms with zero overlap means it is describing the genre.

Where my own tool sits, declared plainly

I am not going to perform neutrality at the end of a page like this. The AI Paper Reviewer on this site is a category two tool: fixed criteria, structured output, a language model underneath. Here is how it answers its own eight questions.

Who set the standard: I did, and my name is on it. Víctor Ciudad-Fernández, PhD in psychology from the University of Valencia, reviewing for Q1 journals in the field. The criteria are the ones I apply when I review, which also means they are one person’s criteria in one field, not a universal measure of quality.

Every criticism in the paid report carries a verbatim sentence from your manuscript, because that is the only way you can tell whether an objection is about your paper or about papers. In the paid report, and only there, your references are checked one at a time, up to sixty, against Crossref, OpenAlex, PubMed, Semantic Scholar and DOAJ rather than trusted; the free report does not check them, so treat any literature it mentions as unverified until you look it up. Every criticism you are shown arrives with what to do about it, and where the fix is statistical the paid report adds the code written in the software your manuscript uses — R, SPSS, jamovi, JASP, Mplus or Stata — not in whatever language the model prefers. At neither tier does it rewrite your manuscript, and that is deliberate: the words stay yours, and so does the defence.

What it does not do: it does not know your journal’s queue or what its editor is tired of reading, it reviews any empirical discipline but its criteria were built and tested in psychology and health sciences, so treat it as strongest there, and it is a language model, so run it twice if a finding surprises you. The diagnostic report is free, with no account and no card: a score overall and by section, a verdict, one major and one minor criticism with their fixes plus another major and another minor that open when you leave your email, and an editorial recommendation. That is the diagnosis, not the whole review. The paid tier is a one-off charge, not a subscription. If a general purpose assistant with a good prompt tells you more than my report does, use the assistant and keep your money.

Frequently asked questions

Which AI tool is best for peer review?

That question has no stable answer, because these products change faster than any comparison can track. The answerable version is which category fits your problem: a general purpose assistant early in the draft, when you want an interlocutor; a specialised review tool for a near final manuscript in a field it was built for; a human read when what you need to decide is whether this is the right paper.

Is it safe to upload an unpublished manuscript to an AI tool?

It depends on the terms of the specific tool, and the answer lives in its privacy policy rather than on its front page. Check whether files are stored, for how long, whether they train anything, and whether you can request deletion. Ask the same of mine. If your data is clinical, what your ethics approval allows binds you whatever the provider permits.

Do I have to declare that I used an AI tool on my paper?

Check the journal rather than generalising from a page like this one. The information sits in the instructions for authors, usually under a heading about generative AI or research integrity, and often again in the submission form. The logic is stable even when the wording differs: responsibility for the text sits with people who can answer for it.

Can I use an AI tool on a manuscript I was asked to referee?

That is a different question from using one on your own work, and it is mainly about confidentiality: a manuscript sent to you for review is not yours to upload anywhere. The rules come from the journal that invited you: check the invitation itself and the journal’s guidelines for reviewers. If you cannot find them there, ask the editor before you upload anything.

Are paid review tools better than free ones?

Price does not predict quality here. What money can buy is criteria written by someone accountable, verification of references against real indexes, a process that reads the manuscript more than once, and output structured enough to work from. A paid tool with none of those is worse than a free assistant driven by someone who knows what to ask.

Can an AI tool replace a human reviewer?

No, and the reason is accountability rather than performance. A reviewer signs a judgement, can be argued with, and can be wrong in a way that has consequences. What these tools replace is the waiting: instead of learning months from now that your measurement model was never going to convince anyone, you find out this afternoon.

How does it compare to the other tools?

You may also find this useful