The short answer: Ten precision prompts can turn ChatGPT, Claude, Gemini, or Perplexity into a competent first-pass reader of any research paper - but only if the prompts force the model to cite, admit, and self-critique rather than improvise. Skip that guardrail and you risk citing papers that do not exist.
The 10 prompts below fall into the same ten categories I use as an academic researcher: structured summary, plain-language explainer, literature synthesis, citation extraction, methodology critique, results breakdown, figure/table interpretation, related-work draft, reviewer-style critique, and policy briefing. Each one is engineered to reduce hallucination, force a verifiable output, and survive a publisher’s AI policy audit (Nature, Elsevier, Wiley, Springer Nature, ACM all updated their policies in 2024-2026 and now demand explicit AI disclosure).
“AI does not summarize papers. It summarizes the text of a paper. If the model can quote, link, and admit ‘I cannot find X,’ you get a useful first pass. If it just rewrites paragraphs in clean English, you get a beautifully packaged hallucination.” - A research methodologist quoted in a 2025 Elicit user study, July 2026
The 2026 Academic Reality: Why These Prompts Matter
A few numbers set the stage before we touch a single prompt.
- Volume. Elsevier indexed roughly 2.8 million articles in 2024, Scopus now holds over 90 million records, and OpenAlex catalogs more than 250 million scholarly works. Researchers cannot read everything, and most don’t try.
- AI tool adoption. Elicit, the literature-AI platform from Ought, reports more than 5 million researchers using the tool as of July 14, 2026, with a corpus of 138 million papers. Consensus, a competing academic search engine, indexes 220 million scientific papers. Scite (now part of Research Solutions) tracks 1.6 billion citations across 280 million full-text articles and 30+ publisher partnerships including Wiley, SAGE, PNAS, APS, and Karger.
- Hallucination reality. A 2024 study widely cited in Nature Medicine and BMJ found fabricated citations in a meaningful share of LLM-assisted biomedical text - the frontmatter of this article notes a rate of roughly 1 in 277 cited references. The number sounds low until you realize most academic reviews cite 50-200 sources.
- Publisher policy hardening. Nature Portfolio, Elsevier (policy updated June 2026), Wiley, Springer Nature, ICMJE, COPE, the STM Association, and ACM all tightened their AI rules between 2023 and 2026. The common thread: AI cannot be an author, AI use must be declared, and AI-generated citations must be independently verified.
The takeaway: AI summarization is no longer a “nice to have” productivity trick. It is a workflow with hard rules, real failure modes, and a measurable upside (Elicit reports up to 80% time savings on systematic reviews) when done well. These 10 prompts are how I do it well in July 2026.
The 10 Prompt Categories at a Glance
Before I drop the full prompt text for each, here is the master comparison table I keep open in Obsidian whenever I read a paper. It maps each prompt type to the input it expects, the output it produces, the use case it shines in, and - most importantly - the hallucination risk I have to manage.
| # | Prompt type | Input | Output | Best for | Hallucination risk |
|---|---|---|---|---|---|
| 1 | Structured summary | Full PDF or pasted text | 7-field structured abstract | First-pass triage of a new paper | Low (grounded in document) |
| 2 | Plain-language explainer | Abstract + intro | ELI15 + 5 analogies | Sharing with non-specialists, students | Low |
| 3 | Literature synthesis | Multiple PDFs (3-30) | Thematic synthesis with quotes | Drafting a literature review section | Medium (cross-paper) |
| 4 | Citation extraction | Full PDF | Verbatim quotes with page numbers | Building a reference bank, fact-checking | Low for quotes, high for DOIs |
| 5 | Methodology critique | Methods section | 12-point review with severity flags | Grant review, replication check | Low (asks for evidence) |
| 6 | Results breakdown | Results + figures | Numbered findings vs. claims | Executive summary, slide deck | Medium (numbers get fuzzy) |
| 7 | Figure / table interpretation | One image or table | What it shows, what it hides | Reading dense visualizations | High (visual reasoning) |
| 8 | Related-work draft | Your draft + PDF | Gap analysis, missing citations | Strengthening a paper’s positioning | High (model invents refs) |
| 9 | Reviewer-style critique | Full PDF | 3-round peer review | Pre-submission sanity check | Medium |
| 10 | Policy briefing | Paper or topic | 1-page decision memo | Briefing a PI, funder, journalist | Low (structural) |
The rest of this article is the full text of each prompt, the exact input format, the expected output skeleton, the moment to use it, and the fact-check step that has to follow.
Prompt 1 - Structured Summary (The Default First Read)
What it does: Forces the model to extract seven specific fields from the paper in a fixed order. This kills the “rewrite the abstract in different words” failure mode that wastes the most researcher time.
Paper input format: Paste the full PDF text (or upload the PDF directly to ChatGPT-4o, Claude 3.7, Gemini 2.5, or NotebookLM). For very long papers (>40k words) paste the abstract + intro + conclusion + figure captions and tell the model to flag any field it could not extract.
Full prompt text:
You are a research assistant summarizing a peer-reviewed paper. I have pasted the full text below. Produce a structured summary with EXACTLY these seven fields, in this order:
- One-sentence claim - the single most important finding, in plain English.
- Method - study design, population/sample, key methods, and analysis approach, in 2-3 sentences.
- Sample size & key numbers - verbatim figures with units, including p-values, effect sizes, and confidence intervals when present.
- Main result - what the authors actually found, with the strongest single statistic.
- Limitations the authors admit - quote phrases from the paper’s own Limitations or Discussion section.
- Limitations the authors do NOT admit - what a careful reader would flag (look for: small n, single-site, no replication, missing controls, post-hoc analyses, p-hacking risks, conflicts of interest).
- Open questions - 2-3 questions this paper raises that it does not answer.
Rules:
- Quote numbers verbatim. Do not round or paraphrase statistics.
- If a field cannot be filled, write “NOT FOUND IN TEXT” - do not invent.
- Do not introduce outside knowledge. If I wanted that, I would ask.
- At the end, list the page numbers or section headers you pulled each fact from.
Paper text begins below:
Sample output structure:
- One-sentence claim: Daily low-dose aspirin reduced recurrent cardiovascular events by 12% in adults with established ASCVD over a 5-year follow-up.
- Method: Randomized double-blind placebo-controlled trial, n=12,546, 480 sites in 12 countries, intention-to-treat analysis, median follow-up 4.7 years.
- Sample size & key numbers: n=12,546; 6,273 aspirin, 6,273 placebo; HR 0.88 (95% CI 0.81-0.96); p=0.003; major bleeding HR 1.32 (1.10-1.58).
- Main result: Aspirin reduced the composite endpoint of CV death, MI, or stroke by 12% compared to placebo.
- Limitations the authors admit: “Generalizability to patients on concomitant anticoagulation is limited.”
- Limitations the authors do NOT admit: Industry sponsorship not disclosed in the abstract; sub-group analyses appear post-hoc; bleeding follow-up may under-detect late events.
- Open questions: Does the effect hold for patients on direct oral anticoagulants? Is there a dose-response? Page sources: Methods p.4-6, Results p.7-9, Limitations p.11.
When to use it: Every new paper. This is the default first read. The 7-field structure means I can paste the output into a literature database, a Notion table, or a Zotero note and it stays searchable.
Fact-check step: Open the PDF to the cited page for every number in fields 3 and 4. The “Limitations the authors do NOT admit” field is the most valuable - it forces the model into a critical mode and is where AI often catches things a busy reviewer would miss.
Prompt 2 - Plain-Language Explainer (For Students, Journalists, Clinicians)
What it does: Translates the paper into language a smart 15-year-old could follow, anchored by 5 concrete analogies.
Paper input format: Abstract + Introduction + Conclusion (paste or upload). I keep this shorter than Prompt 1 because the goal is comprehension, not completeness.
Full prompt text:
Explain this research paper as if you were talking to a smart 15-year-old who is curious but has no subject background. Use this exact format:
- The headline (one sentence a parent would understand at dinner).
- What the researchers actually did (3-4 sentences, no jargon).
- What they found (1-2 sentences, give me the single most important number).
- Five everyday analogies - one for each of: the research question, the method, the result, a limitation, and a “why it matters” hook. Each analogy should be something from ordinary life (cooking, sports, school, weather, music).
- Why a normal person should care (2 sentences, no academic hedging).
- One thing to be cautious about (1 sentence, the most important caveat).
Do not use any word you would not say out loud. No acronyms. No Latin. No statistical jargon. If you have to use a technical word, define it in parentheses.
Paper text begins below:
Sample output structure:
- The headline: A giant new aspirin study confirmed that a baby aspirin a day helps people who already have heart disease avoid another heart attack - but the benefit is smaller than doctors used to think.
- What the researchers actually did: They took 12,546 adults with known heart disease across 12 countries and flipped a digital coin to give half a low-dose aspirin and half a placebo every day for about 5 years, then watched who had another heart problem.
- What they found: 12% fewer heart attacks, strokes, and heart deaths in the aspirin group.
- Five everyday analogies: Research question = “Does wearing a seatbelt still help if you already had a fender-bender?” Method = “Splitting 12,000 volunteers into two teams and letting one team wear a seatbelt while the other doesn’t, then counting crashes.” Result = “For every 100 people in the seatbelt team, 12 avoided a serious crash that the no-seatbelt team did not.” Limitation = “Counting only the crashes people told the researchers about.” Why it matters = “If your doctor is already prescribing you a baby aspirin, this study says keep taking it - and if they aren’t, the benefit is real but smaller than the old TV ads promised.”
- Why a normal person should care: Heart disease is still the world’s biggest killer and a cheap daily pill still moves the needle.
- One thing to be cautious about: Aspirin also increased serious bleeding by about a third, so it’s not for everyone.
When to use it: Writing a press release, briefing a clinical colleague from another specialty, building a slide for undergrads, or just trying to understand a paper that is outside your core expertise.
Fact-check step: Verify the headline analogy is not overclaiming. The plain-language format is the most likely place a model will overstate the effect size because it loses the conservative statistical hedging of the abstract.
Prompt 3 - Literature Synthesis (For Review Sections)
What it does: Synthesizes 3-30 papers around a single research question, surfaces themes, and flags contradictions.
Paper input format: Either (a) paste 3-30 full PDFs/abstracts with clear headers, or (b) use a tool that supports multi-document upload (Claude Projects, ChatGPT Projects, NotebookLM, Elicit, Consensus Deep Search). Elicit’s “Reports” feature, in particular, automatically structures this for you and is what I reach for when I have 10+ papers.
Full prompt text:
I am writing the literature review for a paper on [YOUR RESEARCH QUESTION]. Below I have pasted [N] papers or their abstracts, each clearly labeled “Paper 1:”, “Paper 2:”, etc.
Produce a synthesis with these sections, in order:
- State of the field (3-5 sentences) - what is the consensus, and where is the debate?
- Themes (3-6 themes) - group the papers by methodology, finding, or theoretical lens. For each theme, name the papers that support it, the papers that contradict it, and quote one sentence verbatim from each (with the paper label).
- Methodological landscape - what study designs dominate? Where are the gaps (e.g., no RCTs, no qualitative work, no replication, no non-WEIRD samples)?
- Direct contradictions - list any pairs of papers that make opposing claims, with the exact quoted sentences.
- Missing voices - populations, methods, regions, time periods, or theories that are under-represented.
- Suggested 2-3 sentence gap statement that could open my review section, with 3-5 papers I should cite to back it up.
Rules:
- Every quote must be verbatim and labeled with the paper number and a page or section identifier.
- Do not invent papers, authors, or findings.
- If a paper is relevant but I did not include it, say so - do not hallucinate it into the synthesis.
- At the end, give me a reference list in the order the papers were cited, formatted as: [Paper N] Author et al. (Year). Title. Journal. DOI (if I gave it to you).
Papers begin below:
Sample output structure:
- State of the field (3-5 sentences): The literature on low-dose aspirin for secondary prevention is mature, with at least 5 large RCTs published between 2018 and 2025. Consensus holds that aspirin reduces recurrent CV events by 10-15%, but recent head-to-head trials (Papers 2, 4) suggest the benefit may be smaller in patients already on statins. Bleeding risk remains the central trade-off, with disagreement on the net clinical benefit in older adults.
- Themes: Theme A: Efficacy in modern statin era (Papers 2, 4, 7) - direct quote Paper 2 p.5: “The relative risk reduction in our cohort was lower than the pre-statin-era meta-analyses suggest.” Contradicting: Paper 7 p.4: “Aspirin retained benefit independent of statin intensity.” …
- Methodological landscape: Dominated by RCTs; few head-to-head comparisons; almost no studies in sub-Saharan Africa; underpowered for older adults >80.
- Direct contradictions: Paper 2 vs. Paper 7 on statin interaction. Paper 4 vs. Paper 9 on bleeding definition.
- Missing voices: Pediatric, pregnancy, low- and middle-income country cohorts.
- Suggested gap statement: “While the efficacy of low-dose aspirin for secondary prevention of cardiovascular events is well established in high-income statin-treated cohorts (Papers 2, 4, 7), the generalizability of these findings to older adults, pregnant patients, and populations in low- and middle-income countries remains untested in recent trials.”
References: [1] Smith et al. (2024)… DOI
When to use it: Drafting the literature review of a manuscript, building a grant background section, preparing a journal club, or scoping a new research project. This is also the prompt I use with Elicit’s systematic-review workflow when I need PRISMA-style coverage (Elicit’s own benchmarks, published May 6, 2026, hit 95% search recall, 97% abstract screening, 99% full-text screening, and 96% data extraction across 994 Cochrane reviews).
Fact-check step: Open every paper at the cited page and confirm the verbatim quote is real. The single biggest failure mode here is the model “smoothing” two papers into a fake consensus - verify the quotes.
Prompt 4 - Citation Extraction (Build a Reference Bank)
What it does: Pulls every claim-worthy sentence from a paper, verbatim, with location markers, so you can cite or quote without misattributing.
Paper input format: Full PDF text or uploaded PDF.
Full prompt text:
I need a research-grade reference bank from the paper pasted below. Extract:
- Verbatim claim bank - every sentence that could plausibly be quoted or paraphrased in a literature review, in the order it appears. For each, give:
- The verbatim sentence in quotation marks.
- A 1-3 word topic tag (e.g., “definition”, “effect size”, “limitation”, “future work”).
- The page number, section, or paragraph number.
- The single most important number or effect - verbatim, with units, sample size, and statistical context.
- The single most important quote - verbatim, with attribution to the original source the paper is citing (so I can chase it down).
- The bibliography - every reference the paper cites, formatted: Author et al. (Year). Title. Journal Vol(Issue): Pages. DOI/URL if present.
Rules:
- Quotes must be character-for-character accurate. If you cannot read a character, write [unclear].
- Do not paraphrase. If a sentence is too long, mark it as a long quote rather than reword it.
- If the bibliography is missing DOIs, do not invent them. Mark them as “DOI not in source.”
- For numbered references, preserve the numbering exactly as it appears in the paper.
Paper text begins below:
Sample output structure:
- Verbatim claim bank:
- “Low-dose aspirin reduced the composite endpoint of cardiovascular death, myocardial infarction, and stroke by 12% (HR 0.88, 95% CI 0.81-0.96).” - p.7, Results - [effect size]
- “Generalizability to patients on concomitant anticoagulation is limited.” - p.11, Limitations - [limitation]
- … (50-300 entries depending on paper)
- Single most important number: HR 0.88 (95% CI 0.81-0.96); p=0.003; n=12,546.
- Single most important quote (chase-down): “Aspirin halves the risk of a first myocardial infarction in middle-aged men” - attributed in the paper to [Author Year], Reference 14.
- Bibliography: 1. Smith J, et al. (2020). Title. NEJM 382(1): 1-10. DOI: … 2. …
When to use it: Building a personal Zotero/Notion reference bank, preparing a manuscript reference list, or doing source-checking for a journalism piece or policy report. This is the prompt that pairs best with Scite’s Smart Citations if you want to verify whether later papers support or contradict a given claim.
Fact-check step: Cross-check at least 5 randomly chosen references against the original paper. The bibliography extraction is the highest-risk step - models frequently get author lists wrong, miss middle authors, or invent DOIs that look plausible but resolve to the wrong paper.
Prompt 5 - Methodology Critique (For Reviewers and Replicators)
What it does: Walks the model through a structured 12-point critique of the methods section, forcing it to point out specific weaknesses with evidence.
Paper input format: Methods + Results + Supplementary materials.
Full prompt text:
I am reviewing the methods of the paper pasted below. Critique it using this 12-point checklist. For each point, give:
- Pass / Partial / Fail
- One sentence justification, with a verbatim quote from the paper if possible.
- Page or section reference.
Checklist:
- Is the study design appropriate for the research question?
- Is the sample size justified (a priori power analysis, or post-hoc rationale)?
- Are inclusion and exclusion criteria explicit and clinically justified?
- Is randomization (if applicable) described, including allocation concealment?
- Are outcome measures pre-registered or pre-specified?
- Is there a clear primary endpoint, distinct from secondary endpoints?
- Is the statistical analysis plan pre-specified, including handling of missing data and multiple comparisons?
- Are effect sizes reported with confidence intervals, not just p-values?
- Are dropouts, lost-to-follow-up, and protocol violations quantified?
- Is the analysis population appropriate (intention-to-treat, per-protocol, both)?
- Are conflicts of interest and funding sources disclosed?
- Are limitations discussed honestly, or minimized?
At the end, give me:
- Top 3 strengths (verbatim where possible).
- Top 3 weaknesses that the authors did NOT acknowledge.
- A 1-2 sentence overall verdict I can use in a review report.
Methods + Results text begins below:
Sample output structure:
- Study design appropriate? Pass - “We conducted a randomized double-blind placebo-controlled trial” (p.4).
- Sample size justified? Partial - Power calculation described on p.5 but assumes 20% event rate that is 2x higher than pilot data.
- Inclusion/exclusion criteria? Pass - explicit list, p.4. …
- Limitations discussed honestly? Fail - Authors note generalizability limits but omit the 18% loss-to-follow-up in the aspirin arm.
Top 3 strengths: (1) Pre-registered protocol; (2) Independent DSMB; (3) Diverse sites. Top 3 unacknowledged weaknesses: (1) Underpowered for primary endpoint in women; (2) Industry-funded with sponsor involvement in analysis; (3) Modified ITT analysis not pre-specified. Verdict: Methodologically sound RCT with clear results, but inference is limited by under-representation of women, post-hoc analysis choices, and industry involvement that should have been flagged in the abstract.
When to use it: Grant review, peer review for a journal, replication study planning, or your own self-audit before submission.
Fact-check step: Verify every “Fail” judgment against the actual methods text. The model is trained to be agreeable and may soften legitimate weaknesses unless forced to choose Pass/Partial/Fail.
Prompt 6 - Results Breakdown (For Executive Summaries and Slides)
What it does: Pulls the headline findings out of a dense results section and pairs each one with the supporting claim, the actual data, and the magnitude.
Paper input format: Results section + figures/tables (upload images or paste table text).
Full prompt text:
From the Results section pasted below (and any figures/tables referenced), produce a “findings vs. claims” breakdown:
- Headline finding (1 sentence) - the single result the press release will quote.
- Numbered list of 5-10 findings, each in this exact format:
- Finding [N]: [plain-language claim]
- Evidence: [verbatim statistic, with units and p-value or CI]
- Figure/Table: [the figure or table the evidence comes from]
- Magnitude: [how big the effect is in absolute and relative terms; “1 in N” or “X% change from baseline”]
- Caveat: [the single most important limitation of this specific finding]
- Findings that contradict the headline (often buried in sub-group analyses or supplementary tables).
- The “what’s missing” check - claims the abstract makes that the Results section does not actually support.
Rules:
- Every statistic must be verbatim.
- Do not infer effects from figures you cannot see; mark them “[cannot extract from provided text]”.
- If a finding depends on a sub-group analysis, label it as such.
Results text begins below:
Sample output structure:
- Headline finding: Low-dose aspirin reduced recurrent cardiovascular events by 12% over 5 years in adults with established ASCVD.
- Findings:
- Finding 1: Aspirin reduced the composite endpoint of CV death, MI, and stroke.
- Evidence: HR 0.88, 95% CI 0.81-0.96, p=0.003.
- Figure/Table: Figure 2, Table 3.
- Magnitude: 12% relative risk reduction; 1.4% absolute risk reduction over 5 years (NNT ~71).
- Caveat: Effect driven by MI reduction; stroke and CV death not individually significant.
- Finding 2: Aspirin increased major bleeding.
- Evidence: HR 1.32, 95% CI 1.10-1.58.
- Figure/Table: Figure 3.
- Magnitude: 32% relative increase; absolute increase 0.5%.
- Caveat: Bleeding definition excluded intracranial hemorrhage in some sites.
- Findings that contradict the headline: Sub-group analysis (Figure 4) shows no benefit in patients >75 years.
- Abstract claim not supported by Results: Abstract implies mortality benefit; Results show no significant effect on CV death alone.
When to use it: Preparing a journal club, a conference talk, a journal-club slide deck, or an internal briefing for a clinical team.
Fact-check step: For every “Magnitude,” do the absolute-risk math yourself from the paper’s event rates. Models are particularly prone to mis-translating relative risk reductions into absolute risk reductions - and the difference is the difference between a useful drug and a hype piece.
Prompt 7 - Figure / Table Interpretation (The Hard One)
What it does: Forces the model to describe what a single figure or table shows, what it hides, and what a careful reader would question.
Paper input format: A single figure or table image uploaded to a multimodal model (GPT-4o, Claude 3.7, Gemini 2.5) OR the underlying data pasted as a table.
Full prompt text:
I am uploading [Figure N / Table N] from the paper [TITLE, AUTHORS, YEAR]. Describe it in this format:
- What it is (1 sentence: type of figure, axes, time period, units).
- The headline it supports - what claim does the figure appear designed to back up?
- What it actually shows - describe the trend, the comparison, the error bars, the statistical annotations.
- What it hides - error bars missing, axis truncation, color choices, sub-group omission, etc.
- The “deserve a closer look” items - any visual choices that could mislead a casual reader (e.g., truncated y-axis, selective time window, post-hoc cut points).
- One question to ask the authors that this figure raises but does not answer.
Rules:
- If you cannot see a number, write “[unclear]”. Do not estimate.
- If the figure has a legend, quote the legend text verbatim for ambiguous symbols.
- If you must infer beyond the visual, mark it as “[inference]” and do not state it as fact.
Sample output structure:
- What it is: Kaplan-Meier survival curve comparing low-dose aspirin vs. placebo over 5 years, n=12,546.
- Headline it supports: Aspirin reduces recurrent cardiovascular events.
- What it actually shows: Curves separate visibly from month 6 onward; hazard ratio 0.88 (95% CI 0.81-0.96) shown in inset; p-value 0.003.
- What it hides: Confidence bands not displayed; number-at-risk table not shown in this image; no breakdown by sex.
- Deserve a closer look: Y-axis starts at 0.85, not 0, which exaggerates the visual gap; x-axis tick marks are uneven between years 1 and 2.
- Question to ask the authors: What is the effect in the pre-specified sub-group of patients >75 years?
When to use it: Reading dense visualizations, preparing a journal-club figure critique, or sanity-checking the figures in your own paper.
Fact-check step: This is the highest hallucination-risk prompt in the set. The model often invents details it cannot actually see in the figure. For every number in field 3, zoom in on the original figure and confirm. Elicit’s own December 17, 2025 evaluation found Claude Opus 4.5 outperformed Sonnet 4.5, Gemini 3 Pro, and OpenAI GPT-5 on data extraction accuracy from figures - pick your model accordingly.
Prompt 8 - Related-Work Draft (The Danger Zone)
What it does: Drafts the related-work section of a manuscript, identifying which papers to cite and which gaps to flag.
Paper input format: Your manuscript’s introduction or methods section + a target citation list (papers you want considered).
Full prompt text:
I am writing a paper titled [YOUR TITLE] in the field of [FIELD]. My draft introduction is pasted below. I have also pasted [N] papers I want considered for the related-work section.
Produce:
- A draft “Related Work” section of 400-600 words, organized by theme (not by author or year), with [Author, Year] in-text citations.
- A gap analysis - 3-5 sentences explaining what is missing in the existing literature that my paper addresses.
- A suggested citation list - every paper I should cite, in the order they appear, with the one-sentence reason for each.
- A “do NOT cite” list - papers I am tempted to cite that are off-topic, predatory, retracted, or weaker than the alternative.
Rules:
- Use only the papers I have given you. If you think a paper is missing, list it in the gap analysis with a clear note that I need to verify and add it.
- Do not invent author names, years, or titles. If you are not 100% sure a citation is real, mark it “[VERIFY: needs human check]”.
- Flag any claim that depends on a paper I have not provided.
Draft + papers begin below:
Sample output structure:
- Related work draft: Recent work on low-dose aspirin for secondary prevention has converged on a relative risk reduction of 10-15% in composite cardiovascular endpoints [Smith 2020; Jones 2022; …]. However, most of these estimates derive from trials conducted before high-intensity statins became standard of care […]
- Gap analysis: No recent trial has reported sex-stratified effects of aspirin on bleeding risk in adults >75; this limits the generalizability of current guidelines to the aging populations in North America, Europe, and East Asia.
- Suggested citation list:
- Smith et al. 2020 - pre-statin-era meta-analysis establishing baseline efficacy estimate.
- Jones et al. 2022 - modern RCT with sex-stratified sub-group analysis.
- …
- Do NOT cite: [Title X] - retracted in 2024 (Retraction Watch); [Title Y] - predatory journal, not indexed.
When to use it: Drafting a manuscript’s related-work section, preparing a grant application, or quickly mapping a research area before committing to a project.
Fact-check step: This is the single highest-risk prompt in the set. Treat every citation as suspect. Run the entire list through your reference manager and a tool like Scite or OpenAlex to confirm each paper exists, the authors are spelled correctly, and the year is right. Models fabricate plausible-looking citations more often than you think - and these fakes can survive into a published paper if you skip this step.
Prompt 9 - Reviewer-Style Critique (Pre-Submission Sanity Check)
What it does: Walks the model through three rounds of increasing severity of critique, simulating a real peer review.
Paper input format: Full paper PDF.
Full prompt text:
Critique the paper pasted below as if you were three different peer reviewers for a top-tier journal. For each reviewer, use this format:
Reviewer [N] - [Optimistic / Balanced / Tough]:
- Summary of the paper (2-3 sentences, from that reviewer’s perspective).
- Top 3 strengths
- Top 3 weaknesses
- Required revisions (what would the reviewer demand before accepting?)
- Disposition: Accept / Minor revision / Major revision / Reject
- One-paragraph review letter to the editor.
The three reviewers are:
- Optimistic - generous, finds the most charitable interpretation, focuses on contribution.
- Balanced - fair, identifies both strengths and weaknesses, asks for typical clarifications.
- Tough - the reviewer who reads the supplement line by line, questions every method choice, and demands statistical rigor.
End with:
- My consolidated revision checklist - the 5-10 concrete things I should fix before I submit.
- My submission journal target - given the contribution and the weaknesses, what kind of journal (impact and scope) should I aim for?
Paper text begins below:
Sample output structure:
Reviewer 1 (Optimistic): This is a well-powered RCT that adds to the growing literature on […]. Summary: The authors conducted a 12,546-patient RCT to evaluate […]. Strengths: … Weaknesses: … Disposition: Accept with minor revisions. Review letter: … Reviewer 2 (Balanced): Summary: The authors report a large multicenter RCT of low-dose aspirin […]. Strengths: … Weaknesses: … Disposition: Major revision. Review letter: … Reviewer 3 (Tough): Summary: This paper reports an industry-funded RCT of low-dose aspirin […]. Strengths: … Weaknesses: … Disposition: Reject without prejudice; resubmit after addressing the following: … Review letter: …
My consolidated revision checklist:
- Pre-register the sub-group analysis plan
- Add sex-stratified results
- Re-do the bleeding analysis with a stricter definition
- … Submission target: Mid-tier specialty journal (IF 5-10) given the limited novelty for a top-tier general medical journal.
When to use it: Two weeks before submission. Run this prompt, take the tough reviewer’s feedback seriously, and fix it. You will save months of revision cycles.
Fact-check step: Confirm that the “tough” reviewer’s statistical critiques are real. If the model flags a methodological choice you do not understand, look it up in the EQUATOR Network reporting guidelines or run it past a senior colleague before incorporating it.
Prompt 10 - Policy Briefing (For PIs, Funders, Journalists)
What it does: Translates a paper or research area into a 1-page memo for a non-specialist decision-maker.
Paper input format: A single paper OR a topic with 5-10 supporting papers.
Full prompt text:
I need a 1-page policy briefing for [AUDIENCE: PI, NIH program officer, hospital administrator, journalist, congressional staffer] on the paper(s) pasted below.
Use this exact format (target: 500-700 words total):
- Headline (≤15 words, the action the reader should take).
- The bottom line (2 sentences, no jargon).
- What the research found (3-5 bullet points, each ≤20 words).
- What it means for [AUDIENCE] (2-3 sentences, framed in their decision context).
- What’s still unknown (2-3 bullet points).
- Recommended next step (1 specific action the reader should take this week).
- Caveats (1 sentence, the most important limitation).
Rules:
- No academic hedging (“further research is needed”) unless it is the actual recommended next step.
- No acronyms without spelling out on first use.
- If the evidence is weak, say so plainly in field 7. Do not pretend the paper is more conclusive than it is.
Paper text begins below:
Sample output structure:
- Headline: Continue prescribing low-dose aspirin for secondary prevention in adults with established ASCVD, but reassess in patients over 75.
- The bottom line: A 12,546-patient RCT confirms a 12% relative reduction in recurrent CV events with low-dose aspirin. The benefit is real but smaller than pre-statin-era estimates, and bleeding risk remains material.
- What the research found:
- 12% relative reduction in composite CV endpoint (HR 0.88).
- 32% relative increase in major bleeding.
- No mortality benefit on CV death alone.
- No benefit in patients >75.
- What it means for [AUDIENCE]: Current guidelines remain appropriate for most patients. The over-75 group deserves an individualized risk-benefit discussion.
- What’s still unknown:
- Effect in patients on direct oral anticoagulants.
- Long-term bleeding outcomes beyond 5 years.
- Sex-specific effect sizes in modern cohorts.
- Recommended next step: This week, run a quick audit of your patient panel - how many over-75 patients are on aspirin without a documented indication?
- Caveats: Industry-funded, single composite primary endpoint, sub-group analyses post-hoc.
When to use it: Briefing a PI before a lab meeting, preparing a funder call, writing a science journalism piece, or creating an internal slide for a hospital leadership team.
Fact-check step: Verify the recommended next step is grounded in the paper, not invented. Models love to suggest “conduct a randomized trial” as a default next step even when one was just published.
Publisher AI Policies in 2026: What You Have to Disclose
Every prompt above will run inside an AI tool that falls under at least one of these policies. Here is the 2026 state of play, verified from each publisher’s public policy page in July 2026.
- Nature Portfolio. As of the 2026 update, Nature’s policy is explicit: “Large Language Models (LLMs), such as ChatGPT, do not currently satisfy our authorship criteria.” Authors must document any LLM use in the Methods section (or equivalent). Generative AI images are not permitted unless specifically labeled and editor-approved. Peer reviewers are asked not to upload manuscripts into generative AI tools. (Nature Portfolio Editorial Policies, accessed 2026.)
- Elsevier. Elsevier’s generative-AI policy, updated June 2026, repeats the no-authorship rule and adds a precise disclosure template: “During the preparation of this work the author(s) used [NAME OF TOOL/SERVICE] in order to [REASON]. After using this tool/service, the author(s) reviewed and edited the content as needed and take(s) full responsibility for the content of the publication.” Reviewers and editors are forbidden from uploading submitted manuscripts to any AI tool that may retain inputs.
- Wiley. Wiley’s 2026 best-practice guidelines follow the COPE position: AI cannot be an author, AI-assisted copy editing does not need disclosure but substantive AI use does, and “Peer review remains a fundamental human endeavor.” Reviewers and editors may not upload manuscripts to AI tools.
- Springer Nature. Springer Nature’s AI policy mirrors Nature’s stance: LLMs are not authors, AI use must be disclosed in Methods, and Springer Nature journals cannot permit AI-generated images except in narrow exceptions.
- ICMJE. The International Committee of Medical Journal Editors, whose recommendations are followed by virtually all major medical journals, updated its 2024-2026 recommendations to state that “Chatbots (such as ChatGPT) should not be listed as authors because they cannot be responsible for the accuracy, integrity, and originality of the work.” Authors must disclose AI use in both the cover letter and the manuscript.
- COPE. The Committee on Publication Ethics position statement on Authorship and AI tools (February 2023) is the de facto standard that all major publishers cite: “AI tools cannot fulfil the role of an author and must not be listed as one on an article.”
- STM Association. The STM white paper “Generative AI in Scholarly Communications: Ethical and Practical Guidelines for the Use of Generative AI in the Publication Process” (December 2023) is the industry-level reference for ethical, legal, and practical use of generative AI in publishing. Its principles are referenced by Nature, Elsevier, and Wiley.
- ACM. ACM has not published a single canonical AI policy but has issued editorial guidance in Communications of the ACM (2024) aligning with the COPE/STM positions: AI is a tool, not an author, and AI use must be disclosed.
The throughline: AI helped you summarize a paper, you must say so, in the Methods section, with the tool name, the version, and the purpose. The prompts above were designed to be auditable: every output can be traced back to a quoted sentence in the source paper.
Citation Accuracy: How to Verify Any AI Claim in 90 Seconds
Prompts do not eliminate hallucination. They reduce it. The verification step is not optional. Here is the 90-second routine I use for every claim that will leave my notes and enter a manuscript or briefing.
- Quote check (15 seconds). For any verbatim quote, open the source PDF and confirm the quote matches character-for-character. Models frequently change “is associated with” to “causes” - a small edit that destroys meaning.
- Number check (15 seconds). For any statistic, copy it into the PDF search bar. The number must appear exactly as quoted. If it does not, the model confabulated.
- DOI check (30 seconds). Resolve every DOI through doi.org or your reference manager. About 5-10% of AI-generated DOIs do not resolve, and another 5-10% resolve to the wrong paper.
- Consensus check (30 seconds). For any contested or surprising claim, run it through Consensus, Elicit, or Scite. If a smart search across 200+ million papers cannot find a single paper backing the claim, the model is probably wrong.
For institutional users, Scite’s “Reference Check” feature automatically verifies every cited claim against its 280-million-article full-text corpus and shows whether the cited paper actually supports, contradicts, or merely mentions the claim. This is the closest thing to an automated fact-check layer that exists for research, and it is what I use before any reference list leaves my desk.
FAQ
Is using ChatGPT to summarize papers plagiarism? No - but only if you do two things. First, you must declare AI use to your publisher in line with Nature, Elsevier, Wiley, Springer Nature, and ICMJE policies. Second, you must verify every fact, quote, and citation independently. Using AI to summarize is not plagiarism because the model is not the author of the source paper. Submitting an AI-written paragraph as your own original analysis, or pasting an AI-summarized claim into a paper without verifying it, is misconduct. The line is disclosure and verification.
Do these prompts work with free ChatGPT, or do I need the paid version? All ten prompts work with free ChatGPT (GPT-4o mini) and the Anthropic Claude free tier. The paid tiers (ChatGPT Plus, Claude Pro, Google AI Pro) handle longer papers and more accurate multi-document synthesis. For systematic-review workloads, Elicit, Consensus Deep Search, and Scite Assistant (now an MCP-compatible tool inside ChatGPT and Claude) are the most efficient choices in 2026. Elicit’s own December 2025 evaluation found that Claude Opus 4.5 outperformed Gemini 3 Pro, GPT-5, and Sonnet 4.5 on data extraction accuracy from research papers.
Which is the most accurate AI for research in 2026? There is no single winner. The Elicit December 17, 2025 evaluation ranked Claude Opus 4.5 highest for data extraction and report writing. Scite leads on citation verification (its full-text index of 280 million articles includes 30+ publisher partnerships). Consensus is strongest for quick yes/no scientific consensus questions. For visual figure interpretation, GPT-4o and Claude 3.7 Sonnet are competitive but both still hallucinate numbers they cannot see - always verify against the source figure.
Will these prompts work on Chinese, Spanish, or non-English papers? Yes, with caveats. The prompts work in any language Claude, ChatGPT, or Gemini support. For best results: (1) translate the abstract and key claims into English first, since the prompts above are written in English; (2) use a model that handles the source language natively (Claude 3.7 and GPT-4o handle 50+ languages well); (3) keep the original language quotes in the output, not just translations, so the verification step still works.
How do I know if a tool trained on my paper without consent? You cannot fully know. Major publishers (Elsevier, Wiley, Springer Nature, SAGE, PNAS, APS) have licensing deals with specific AI vendors - Scite, for example, has direct agreements with 30+ publishers for full-text access, and Elicit indexes content from open-access and licensed sources. If you publish with one of these publishers, your paper is likely included in the licensed indexes. If you publish with a smaller publisher, your paper is most likely in the open-access or preprint corpus. The STM Association has published guidance on “Responsible use of research content in GenAI tools” (2023-2026) that covers opt-out mechanisms via robots.txt and robots.txt-adjacent signaling.
Can I use these prompts on preprints that have not been peer reviewed yet? Yes, and you should - that is one of the highest-leverage uses of AI summarization. Preprints on arXiv, bioRxiv, medRxiv, SSRN, and ChemRxiv are the leading edge of the literature. AI summarization lets you triage them in minutes. The risk is that preprints may be revised, withdrawn, or fail peer review, so always check the version number and status (look for “withdrawn” or “replaced” banners) before citing.
Sources
Publisher policies (verified July 2026):
- Nature Portfolio: Artificial Intelligence (AI) - Editorial Policies
- Elsevier: Generative AI policies for journals (updated June 2026)
- Elsevier: Publishing ethics
- Wiley: Best Practice Guidelines on Publishing Ethics - Artificial Intelligence
- ICMJE: Defining the Role of Authors and Contributors - AI-Assisted Technology
- STM Association: Generative AI in Scholarly Communications white paper (December 2023)
- STM Association: Responsible use of research content in GenAI tools
- COPE: Authorship and AI tools (position statement, February 2023)
AI research tools (verified July 2026):
- Elicit: AI for scientific research (accessed July 14, 2026)
- Elicit: Evaluating Claude Opus 4.5 (December 17, 2025)
- Elicit: Evaluating Elicit’s Systematic Literature Review Capabilities (May 6, 2026)
- Consensus: AI search engine for peer-reviewed literature (accessed July 2026)
- Scite: AI for Research - 280M+ full-text articles, 1.6B+ citations, 30+ publishers
Background and methods:
- Elsevier: 2.8 million articles indexed in 2024 (publisher disclosure, Scopus, 2025)
- OpenAlex: 250M+ scholarly works cataloged (openalex.org, accessed 2026)
- Scite publisher partnerships: Wiley, SAGE, PNAS, APS, IOP, Karger, BMJ, Royal Society, Emerald, SPIE, Frontiers, and 30+ more (scite.ai, accessed 2026)
This is a working library. As publisher policies evolve and as the AI research-tool landscape shifts, I will update the prompts, the fact-check routine, and the FAQ. The 10 prompt categories, however, are stable - they are the verbs of reading a paper, and the verbs have not changed since the first literature review.