How much can LLMs help with evidence reviews?
This guest blog by Dr. Savita Bailur (MERL Tech Initiative and School of International and Public Affairs, Columbia University) and Dr. Caryl Feldacker (Gates Foundation and Department of Global Health, University of Washington) was originally posted by Panoply on August 11, 2026. With thanks to Alex Tyers-Chowdhury for facilitating this research, and to Ambarish Haldipur for writing the GitHub code in Prompt 3.
For many researchers, the evidence review stage of a project can feel like being an archaeologist – digging through journals and reports, getting through paywalls, downloading PDFs, and reading abstracts just to see if a paper is relevant. This piece is a summary of an experiment we conducted together to see how LLMs may (or may not) speed up literature searches. We set out to see whether LLMs could support, or even replace, part of a manual literature search on AI interventions in low-income women’s health, specifically around reproductive, maternal, newborn, and child health (RMNCH) programs. This piece captures the high level findings – we have more details including the prompts here.
A manual literature search can look something like this:

The process can be time consuming and piecemeal. Tools such as ChatGPT, Claude, Elicit, Consensus, and Semantic Scholar promise cutting down both searching and reading time, offering synopses and generally scaling and speeding research at an unprecedented level. An LLM-assisted single paper search instead may look like this:

Going one step further (what we call Prompt 2 below) researchers can extend the same LLM prompt for an entire scan of the web. However, as anyone who has attempted any of the above LLM assisted searches also knows, there is always a sense of missing something. The above searches may result in superficial resources. There are also trusted sources in one’s field, including peer-reviewed journals, or authors known to work in a particular area, and it’s possible these get omitted from the results. The chosen field of research may also be complex to define if it is interdisciplinary, for example, assessing the impact of AI interventions in low-income women’s health … the topic of this experiment.
The starting point: mixed feelings about AI in the development sector
Interest in AI impact among foundations and nonprofits has grown fast, sometimes driven more by fear of falling behind than by clear evidence of value. A 2025 survey of over 200 foundations and 450 nonprofits found that most organizations felt uncertain about AI’s value and worried about bias and hallucination, yet still wanted to experiment with it. Despite a run in “AI for good” playbooks, such as The Generative AI Evals Playbook (Center for Global Development, Agency Fund, and ID Insight), J-Pal’s AI Evidence Playbook and Dalberg’s People Centered Playbook there is still a need to understand the impact pathways of AI (there is also often a need to define AI). As noted by Emeka Nwankwo, all the above constitute a disparate set of guidance. While there are many commentaries, or positive opinion pieces on the value of AI, it is more challenging to find rigorous studies of impact, and researchers often spend valuable time trying to sift through publications to look for the nuggets of evidence.
We tested three AI-assisted approaches
Against this disparate evidence, we started with a manual review of 14 papers on AI for impact, using criteria defined together with the Gates Foundation (see Appendix A for criteria).
We then tested three AI-assisted approaches.
In Prompt 1, we asked Claude to analyze a single paper against the same criteria used in the manual review. We initially tested six products – ChatGPT, Claude, DeepSeek, Elicit, Gemini, and NotebookLM. We did not track the model versions for each of these though we accept this would also impact on results. We narrowed down to Claude for Prompt 1 for practical reasons (Claude handled the prompts with more detailed outputs and we had a paid subscription), Claude and Elicit for Prompt 2 to compare results at scale, and Gemini for Prompt 3 (again the practical reason was to stay within the Google ecosystem, as this was an experiment to integrate with Google Scholar[1]). [link to Prompt 1]

In Prompt 2, we scaled this up into a broader search across the web, using both Claude and Elicit [link to Prompt].
In Prompt 3, recognizing that many researchers set up Google Scholar email alerts, we automated the process further by setting up a script to access Google Scholar emails in one’s inbox [link to Github].

[1] Towards the end of this exercise, Anthropic and Gates announced a $200 partnership in grant funding, Claude usage credits, and technical support for programs in global health, life sciences, education, and economic mobility but this exercise pre-dated that.
However, as Prompt 3 requires more automation, including handing over Gmail access and a credit card for access to the Google cloud ecosystem, we anticipate this may be less popular among researchers. The automation may also break if the alert format is changed.
Takeaways for researchers using AI for evidence reviews
We document the findings in the more detailed article which follows. However, it was clear that Prompt 1 was the most useful because a) we had identified a paper that we had sourced and considered useful and b) we had designed an extremely detailed, well-caveated prompt. Even here, however, there are limitations. An LLM tends to describe what’s present in a paper rather than notice what’s missing, which is something an experienced researcher can pick up on. Second, results are inconsistent: the same model, the same prompt, run minutes apart, can produce different conclusions. Third, and perhaps most concerning, using an LLM to analyze a paper on AI tends to inflate the value of AI itself and is tech-positivist (applying AI will solve everything).
We started this piece with the comparison of an academic researcher to an archaeologist. Both involve life-long learning, passion and curiosity. At the end of this experiment, Caryl wrote a reflective note: “Conducting a lit review is a little tedious, but it also helps us actually understand the field deeply. It is a pleasure to gain insights and expertise by actually deeply reading and mulling over the strengths and weaknesses of multiple papers, from multiple places, with often contradictory findings or approaches. While AI is moving at the pace of light, we mortal humans cannot. We should not use this set of insights on how better to use AI to cede the power of knowledge acquisition that we seek. We should not value these tools as any more than short-cuts. We all know that some of the pleasure is in the pain of completing a tough journey or the drive itself. A short cut is just that — a cheat that deprives us of the true worth of an activity, learning or experiential. I say this all with mistakes, and with humility, since I am not AI ”. It is this passion that still drives us as researchers, and while using AI is something we are getting to grips with as a new technical skill, we need to acknowledge the limitations, and not outsource our passion, curiosity or our critical thinking skills.
Appendix A: GE-DC Criteria
These were the parameters of the AI-enabled search (also captured in each prompt – see Appendix B & C for the prompts). The idea is that these criteria could be adapted according to the scope of the researcher:
These were the parameters of the AI-enabled search (also captured in each prompt – see Appendix B & C for the prompts). The idea is that these criteria could be adapted according to the scope of the researcher:
- Gender should be prioritized but it is not the sole focus (although should be noted when it is absent)
- A connectivity segment lens (unconnected, under-connected, meaningfully connected, AI ready) should be prioritized – with an assessment of who was excluded from the intervention or how inclusion was quantified but it is not the sole focus (although should be noted when it is absent)
- Unconnected: women who are aware of internet but have not used it in the last 3 months
- Under-connected: women who have used internet at least once in the last 3 months but want to use it more
- Meaningfully connected: Individuals with own phone or decision-making power for a shared reliable (ideally, smart) phone and consistent/reasonably affordable internet; and
- AI ready: Individuals who a) have access to or ownership of smart phones of sufficient quality to hold and run AI applications, b) live in geographies with strong, fast, reliable and affordable data infrastructure to implement scalable AI models, and, when/where possible, c) those who live in geographies with other enabling environment factors (e.g., GENAI in local language, compute capacity, digital governance, data access, etc.).
- Literature should be published between 2021-2026, with a bias towards more recent literature and reports, 2025+
- Key sources will include, but not be limited to academic journals and conferences such as ACM Dev, AEA, CHI, FACCT, HCI, IT for Development, ICT4D, Nature (in alphabetical order)
- Key institutional publications will include, but not be limited to Agency Fund, Caribou, Center for Global Development, DAIR, ID Insight, Gates Foundation, GitLab Foundation, GSMA, Humane Intelligence, MERL Tech, UNESCO, UN Women, Yale Inclusion Economics, World Bank (in alphabetical order)
- Evaluations can include, but are not limited to RCTs
- Geographic scope should prioritize Gates countries (Rwanda, Kenya, India, Pakistan, Ethiopia, Tanzania, Nigeria, etc) but other countries can be included
- In health, RMNCH likely to be priority (FP, education, ag/dev, fintech)
- Key beneficiaries when considering impact – intermediary, end user, other?
- Confidence in methodological rigour (high, medium, low)
Leave a Reply
You might also like
-
Digital and AI-enabled Social and Behavior Change: Snippets from the 2026 SBCC Summit
-
Building Practice around Responsible AI Principles in the Social Sector
-
Learning about the environmental impacts of data centers in Brazil with Rhavena Madeira and André Fernandes
-
Meta just dropped a bomb on chatbot builders. Here’s how it impacts the development and humanitarian sectors


I would be interested if you tried a literature search LLM like Consensus?