Authors: Brihi Joshi, Keyu He, Sahana Ramnath, Sadra Sabouri, Kaitlyn Zhou, Souti Chattopadhyay, Swabha Swayamdipta, Xiang Ren Venue: Association for Computational Linguistics (ACL), 2025 — accepted

Summary

A good explanation is not one thing. Explaining the same thing to a fifth grader and to a graduate student are different tasks, and ELI-Why asks whether a language model can tell them apart. The benchmark holds 13.4K “Why” questions and is built to test whether models adapt an explanation to the educational background of the person asking.

Two human studies measured how well GPT-4 does this. It landed on the intended grade level only about half the time, and its explanations were rated 20% less suitable for learners’ needs than explanations curated by laypeople. Generating a fluent answer, in other words, is not the same skill as pitching it at the right reader — by this measure current models are not good teachers yet.

Paper · Code · Data

Thread: AI for Education — the research thread. Research Agenda — covers the “users from different knowledge backgrounds” domain.