Systematic reviews are among the most trustworthy artifacts in medical research: a structured, reproducible synthesis of everything known about a question, built specifically so that a clinician, a guideline panel, or a regulator can rely on it. But, they’re slow. A single review can take a year or more to complete, and by the time it’s published, new studies have often already outpaced it. That gap between how fast evidence accumulates and how fast reviews can responsibly be produced is exactly where AI has moved in over the past two years.

An interesting piece in The New Republic walked through what that looks like in practice as the state-of-the-art, and the picture is mixed. Some tools on the market let a user type a research question into a chatbot style interface and get back a search, a set of included studies, and a synthesized answer, all generated in one pass with little visibility into how any of it was decided. Others build AI into the structured, staged process reviewers have used for decades, search, screening, extraction, appraisal, synthesis, with the model doing one bounded task at a time. Both get called “AI for systematic reviews.” They are not the same thing, and the difference matters for whether the resulting review can actually be trusted.


What Is At Stake?


The concerns aren’t hypothetical. Reproducibility is one of the foundational requirements of a systematic review, and generative AI tools that summarize across studies often can’t guarantee it: the same question, asked twice, can surface different studies and reach different conclusions. Traceability is another. When a model generates a claim or a summary, a reviewer needs to be able to trace that output back to the specific line in the specific study that supports it. Many tools simply can’t show their work. And there’s a subtler risk that’s easy to miss: a slick AI search can create the impression of comprehensiveness even when the model only had access to a fraction of the literature, often just what’s freely available online, leaving paywalled and low-resource-setting research invisible to the review.

These are the exact failure modes the field’s own methods bodies have been racing to get ahead of. In late 2025, Cochrane, the Campbell Collaboration, JBI, and the Collaboration for Environmental Evidence, working through the International Collaboration for Automation in Systematic Reviews, released RAISE: a set of actionable standards for responsible AI use in evidence synthesis. RAISE doesn’t treat all AI use as equally acceptable. Some applications, like using AI to help draft a first-pass search strategy, are treated as exploratory and supplementary. Others, like AI-assisted screening or data extraction, are considered usable only once validated within the specific review, either through full human verification or a documented check against expert benchmarks. And some, like having a language model synthesize conclusions across studies on its own, are flagged as not acceptable at all, at least not yet.


Building to the Standard, Not Around It


This is the landscape Nested Knowledge was built for, and it shapes how every AI feature on the platform works.

Every output stays traceable to its source. Whether it’s Smart Search proposing a Boolean string, Smart Screener applying inclusion criteria, or Adaptive Smart Tags pulling a data point from a full text, the result points directly back to the passage in the source study that generated it. That traceability is the direct answer to the black box problem now under scrutiny across the field.

Nothing finalizes without a human. Screening, tagging, and critical appraisal are all designed around a review-and-curate model. A reviewer sets the criteria, checks the output, and makes the final call. That’s the human-in-the-loop standard RAISE lays out for every one of these stages, not an afterthought layered on top of automation.

We show reviewers exactly where the line is, tool by tool. Not every stage of a review is equally ready for AI, and pretending otherwise is part of what got the field into this position. That’s why we’ve mapped every Nested Knowledge tool against RAISE’s own acceptability tiers in our RAISE Compliance Guide, so a team can see plainly where a tool’s output can be relied on directly, where it needs in-review validation first, and where the heavy lifting, like quantitative and qualitative synthesis, still has to be done by a person.

Search stays a supplement, not a replacement. The false-comprehensiveness problem is one of the most consequential risks on this list, because it’s invisible until something important gets missed. In line with RAISE’s own recommendation, AI-assisted search in Nested Knowledge is built to extend traditional search methods, not stand in for them.


The Distinction That Matters


The gap between “AI wrote the review” and “AI supported a reviewer who is accountable for the review” is the whole story right now. One produces something that loosely resembles a systematic review. The other produces an SLR that actually behaves like one: reproducible, traceable, and answerable to a human being who can defend every decision in it. That’s the distinction the field is converging on through frameworks like RAISE, and it’s the one we’ve built Nested Knowledge around from the start.

A blog about systematic literature reviews?

Yep, you read that right. We started making software for conducting systematic reviews because we like doing systematic reviews. And we bet you do too.

If you do, check out this featured post and come back often! We post all the time about best practices, new software features, and upcoming collaborations (that you can join!).

Better yet, subscribe to our blog, and get each new post straight to your inbox.