Skip to content
Blog

Observations from CIES 2026: What the education field is actually wrestling with

Marc Shotland 9 April 2026

IDinsight Chief Technical and Learning Officer Marc Shotland (second from right) at the CIES Annual Conference 2026 along with fellow panelists | L-R: Alex Their (Lapis Group), Gouri Gupta (Central Square Foundation), Natalia Kucirkova (International Centre for EdTech Impact), Marc Shotland (IDinsight), and Ally Mackintosh (Fab.AI)

I recently spent five days at the Comparative and International Education Society (CIES) annual conference in San Francisco (along with my colleagues, Doug Johnson, Neha Raykar, Jeff McManus, and Valentina Brailavskaya). I attended sessions, had a lot of coffee, co-presented at a pair of panels, and tested my Hindi on an AI-powered oral reading assessment (more on that shortly). The field is genuinely wrestling with five themes that came up again and again, across sessions, hallway conversations, scheduled meetings, and the occasional networking event.

1. AI is already in LMIC classrooms. We’re playing catch-up.

There is a version of the AI-in-education conversation that goes: high-income countries will adopt AI tools, and low-income countries will be left behind due to access gaps, digital literacy, and infrastructure. That framing is rapidly becoming outdated.

Teachers and students in low- and middle-income countries are already using AI—often without any system-level awareness, let alone governance. Central Square Foundation’s BaSE 2025 report, a large-scale survey of 12,500 households and 2,500 teachers across ten Indian states, finds that 64% of children are using EdTech, and 35% of EdTech-using children already use GenAI for learning, with 69% of those using it daily. Among children aware of GenAI, 68% were introduced to it by peers. This is not theoretical future adoption; it is happening now, largely peer-driven and outside any formal system.

The question is not whether AI will reach these classrooms. It is whether the systems around those classrooms will be ready to shape how it is used. Sessions across CIES (including the five Fab AI panels, the EdTech Hub panels I attended, plus side events) returned repeatedly to this urgency: the field needs to move from only “what are the ways we can leverage AI in LMIC education?” and “what are the potential risks?” to “how do we understand and shape what is already underway?”

2. Equity is a quality issue, not just an access issue.

The equity conversation at CIES was more sophisticated than I usually hear, and it connects directly to the adoption reality above.

Yes, access remains a real barrier. The BaSE data shows meaningful rural-urban gaps, and the digital divide is real. But some of the benchmarking work by my colleagues at Fab.AI show that there’s a real quality disparity. Low-resourced and low-connectivity contexts often drive adoption toward older, cheaper models or offline-first tools. And language constraints add further barriers. Many of those models/languages score significantly lower on benchmarks that measure their content and pedagogical knowledge. So even as AI adoption spreads, there is a genuine risk of a two-tier system: higher-quality AI in higher-resourced contexts, lower-quality AI in precisely the places that need it most.

Claudia Ramly making the case for Fab.AI and IDinsight’s Quality Assurance Facility.

Related to Equity, Natalia Kucirkova presented a 5Es framework—Efficacy, Effectiveness, Ethics, Environment, Equity—that explicitly places equity in the frame for quality standards. That is the right instinct. Equity needs to be a design constraint on tools themselves, not just an aspiration for distribution.

3. On implementation: we’re measuring the wrong things, or not measuring at all.

The What Works Hub for Global Education (led by Noam Angrist and Kate Sturla) facilitated a workshop on implementation measurement. I was super excited, given my recent interest in implementation science. I’ll be honest: when they started walking through the framework’s basic structure—there are administrative systems, schools, teachers, and students—I nearly rolled my eyes. It felt too obvious to be useful or novel.

But then it got interesting. The first was a distinction between fidelity to design vs fidelity to best-practice, which can diverge in important ways. A program can be implemented exactly as designed and still fall far short of what the evidence says works. Conversely, a frontline implementer might deviate from the written protocol in ways that actually bring them closer to best practice. It’s an interesting new way to think about what program fidelity is and how it can lead to impact.

The second was their finding that quantity (dosage) is a stronger predictor of impact than quality of delivery. My instinct was to push back on this: dosage probably matters more conditional on quality being above some threshold. If you are starting from very low quality, marginal improvements in quality likely matter more than additional dosage of something that is not working well. The finding should not be read as a license to ignore quality. Context matters.

The headline statistic across all of this was stark: only 12 or 13% (I forget which, even though they repeated it a dozen times) of education program evaluations include any implementation data at all. That means when something does not work, or when it does, we almost never know how it was actually delivered, whether fidelity held, or what broke down. The field is flying blind on the mechanisms. 

4. Scaling with government is genuinely hard, and the frameworks are getting sharper.

Multiple sessions grappled with the gap between what works in a pilot and what survives contact with a government system at scale. This is not a new observation, but the quality of thinking about it is improving.

It was particularly nice to see my old colleague Ashleigh Morrell from TaRL Africa present (albeit remotely) on their experiences navigating this gap: the genuine complexity of sustaining fidelity to core principles while adapting to the realities of different government systems across the continent.

On the frameworks side, IPA’s new Scaling Ingredients Framework offers a useful structured lens. It breaks the scaling process into six ingredients across four levels—purpose, model, implementer, and ecosystem—covering evidence of cost-effectiveness, implementation feasibility, implementation capacity, business or funding model, favorable ecosystem, and relevant problem or need. What I appreciate about this framing is that it treats scale-up as a diagnostic exercise, not just a logistics problem: asking which ingredients are present and which need investment, rather than assuming that a proven model will automatically find its footing in new contexts.

5. The middle tier is underleveraged, especially in program design.

Multiple sessions at CIES were dedicated to the “middle tier” of education systems—teacher coaches, cluster, block, and district offices that sit between national policy and individual schools. I found this genuinely exciting, not because the observation is new, but because the field seems to be taking it more seriously and with more rigor. The middle tier was a key ingredient of success in Pratham’s TaRL scale-up in Haryana.

Dhir Jhingran from Language and Learning Foundation (LLF) chaired a session that featured presentations from LLF, VVOB, and ideas42, with Ben Piper, Director of Global Education Programs at the Gates Foundation, as discussant. Ben pressed hard on the incentive question: what actually motivates the middle tier to do coaching? What does the supervisor’s supervisor’s incentive structure look like? It is the right question, and it usually does not get asked.

Some of global education’s preeminent thought-leaders wrestling with the same issues

I offered an example from Pratham’s work: when Academic Block Resource Coordinators were treated as mentors—as sources of expertise—rather than as postal workers delivering information downstream, they engaged differently. That is tapping into intrinsic motivation rather than compliance.

Shveta Lall, of Language and Learning Foundation, made a point that stuck with me: when the middle tier get to witness moments of genuine expertise and success, when they see something work, that becomes a powerful source of motivation that external incentives often cannot match. It made me think about Daniel Pink’s Autonomy-Mastery-Purpose framework. The implication for program designers: you are not just designing a supervision structure. You are designing conditions for people to experience mastery. And those moments of mastery are what sustain engagement when central attention and external support move on.

CIES 2026 left me cautiously encouraged. The field is asking better questions than it was five years ago: about implementation, about scale, about what equity actually requires. The challenge now is generating the evidence and frameworks to answer them.

A personal note: I did test Pratham’s PadhAI tool, an AI-powered oral assessment delivering the ASER reading test in Hindi. I scored perfectly on the paragraph-level task, despite making at least one small error that the tool missed—maybe blame my terrible accent. We mercifully ended the assessment there, and agreed not to find out how I’d do on comprehension.