AI Evals: How do we approach gender-intentional chatbots?
Written by Savita Bailur, Alex Fulcher, Najihah Ahmad Fikri, Maria Luiza Alves da Costa, Yasmina Male and Ana Montañez

Image Credit: Girl Effect
Where do girls turn to if their period is late, if they had sex without protection, or if there’s something they might be too ashamed to ask someone? Maybe they don’t know who to ask – a parent might shame them, a friend might gossip, a clinic might be too far away and they don’t know how to get there. Google, Claude (or other LLM), or an influencer may be the answer if they know who to turn to and what they need to ask.
Another path may be information provided by governments or NGOs. AI-enabled chatbots have emerged as a promising mechanism for providing SRH (sexual and reproductive health) information to girls in an accessible, non-judgemental way. These chatbots have potential if built with the right guardrails, language, and relevant direction to resources. They can act as a private space to ask questions girls might not feel comfortable posing to a parent or peer and come at a lower cost than traditional health counselling. The issue is that we often have no evaluation framework to understand what’s working and what’s not.
“AI evals” are a hot topic right now, but in SRH (sexual and reproductive health) they need to be designed thoughtfully. Earlier this year, Girl Effect and Columbia SIPA partnered to design and test an evaluation framework for model performance of SRH chatbots. We tested this scorecard with Girl Effect’s own product, BigSis, as well as four other chatbots (those will remain in an internal report).
We aligned our evaluation to the “Level 1” evaluation layer of the Agency Fund/Centre for Global Development/ID Insight GenAI Evals Playbook.

Source: GenAI Evals Playbook
The idea is that this is a public good to be improved upon, and to that end, we’ve hosted a first version of the framework and a tentative scorecard on MTI’s Playground site. We will also share reflections from each of the team members in related blogs.
Methods
In order to draft this framework, we reviewed SRH and chatbot literature, spoke to twelve experts across SRH, AI, UX design, and digital development, and built a framework and scorecard of more than 30 weighted criteria organized around three questions:
- Is it built responsibly (Product Design)?
- Does it actually work for the user (User Experience)?
- Is it set up to last and to prove its impact (Forward-Looking)?

To do so, the team tested each of the selected chatbots with three tiers of questions ranging from general SRH information (Tier 1) to sensitive or moderate-risk scenarios (Tier 2) to high-risk and emergency situations (Tier 3).
Some of the concerns the team raised and will write about in subsequent outputs included:
- Chatbots performed worse when prompts were more sophisticated, rather than a simple input prompt. Vaguer prompts, particularly those in Tiers 2 and 3, like “I think I’m pregnant and I’m scared” was consistently harder for chatbots to interpret and respond to appropriately than more explicit requests. When crisis-related prompts lacked specificity, responses were handled inconsistently, and some chatbots required multiple follow-up exchanges before addressing the core concern. This is particularly consequential given that real users in distress are unlikely to phrase their situations with clinical precision.
- All the chatbots reviewed began with a clarification of consent. But what does consent actually mean? And what are the guardrails of privacy?
- What does “de-implementation” of such a chatbot service look like? What if funding runs out or the number is changed? De-implementation planning was the weakest area across all the five products we reviewed. Almost none of the chatbots had a plan for what happens to users and their data if the service shuts down. This is a concern not just for privacy and data security, but for the users themselves: many of these chatbots run on WhatsApp, so if funding ends the number simply goes unreachable, with no advance notice and nowhere to redirect people. This is of particular concern as many chatbots are hosted on WhatsApp.
- How generalizable is such a framework and scorecard to other chatbot topics? How many of the criteria are specific to SRH, and what is generalizable?
Over the next few posts, we’re going to unpack all the above points and what we learned during the process. And if you’d like to check out the framework and scorecard as a public good, please check it out here or contact savita @ merltech.org for more information.
You might also like
-
Nov 5th: Digital backlash, techno-authoritarianism and gender justice
-
Closer to Home: Radical Theories for Africa’s Digital Realities- Data Tamasha Africa Event Recap
-
Event recap: Launching the new book “Aid and ID: Making People Matter” with editor Margie Cheesman and book contributors
-
Exploring AI in Evaluation and Evaluation of AI at the 2026 European Evaluation Society (EES) Conference
