Abstract
Background: Multimodal generative artificial intelligence (GenAI) can process written language, images, typography, layout, and interface-like elements, creating opportunities for literature education but also new epistemic risks. Vision-capable models may combine accurate recognition, plausible inference, and invented visual detail in a single fluent interpretation. The central challenge is not access to multimodal information, but disciplined separation of observation, inference, interpretation, and human judgment.
Objective: This perspective paper develops Multimodal Evidence-Bounded Interpretation (MEBI), a human-centered framework designed to preserve close reading, interpretive plurality, evidential accountability, and student authorship when multimodal GenAI is used with literary texts.
Methods: The study uses an evidence-informed conceptual synthesis and a worked multimodal case. Recent peer-reviewed research on GenAI-supported learning, student agency, literary analysis, multimodal AI, and visual hallucination is integrated with established scholarship on multimodality and graphic narrative and with human-centered AI guidance. The framework is demonstrated through selected scan locations from Nora Dåsnes’s Czech graphic narrative Na kočičí svědomí (2024).
Results: MEBI comprises five design propositions—human-first encounter, modal traceability, observation–inference separation, preserved interpretive plurality, and human adjudication and authorship—and a five-phase sequence: Encounter, Describe, Hypothesize, Audit, and Adjudicate & Author. Five evidence tags track verbal, image, sequential-spatial, typographic-graphic, and digital-interface evidence. A worked case, classroom sequence, evidence ledger, and assessment heuristic operationalize the framework.
Conclusion: MEBI positions AI as a generator of descriptions, hypotheses, counter-readings, and questions rather than as interpretive authority. It offers a researchable design for multimodal literature education while requiring empirical validation across genres, languages, age groups, and educational contexts.
References
Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö., & Mariman, R. (2025). Generative AI without guardrails can harm learning: Evidence from high school mathematics. Proceedings of the National Academy of Sciences, 122(26), e2422633122. https://doi.org/10.1073/pnas.2422633122
Bellot, A. R., Gutiérrez-Colón Plana, M., & Baran, K. A. (2025). Redefining literature education: The role of ChatGPT in undergraduate courses. International Journal of Artificial Intelligence in Education, 35, 3185–3201. https://doi.org/10.1007/s40593-025-00497-3
Chan, C. K. Y., & Wong, H. Y. H. (2025). Assessing the ability of generative AI in English literary analysis through Bloom’s Taxonomy. Discover Computing, 28, 127. https://doi.org/10.1007/s10791-025-09572-8
Cohn, N. (2013). The visual language of comics: Introduction to the structure and cognition of sequential images. Bloomsbury Academic.
Darvishi, A., Khosravi, H., Sadiq, S., Gašević, D., & Siemens, G. (2024). Impact of AI assistance on student agency. Computers & Education, 210, 104967. https://doi.org/10.1016/j.compedu.2023.104967
Dåsnes, N. (2024). Na kočičí svědomí (J. Jindřišková, Trans.). Albatros. (Original work published 2020 as Ti kniver i hjertet, Aschehoug.)
Deng, R., Jiang, M., Yu, X., Lu, Y., & Liu, S. (2025). Does ChatGPT enhance student learning? A systematic review and meta-analysis of experimental studies. Computers & Education, 227, 105224. https://doi.org/10.1016/j.compedu.2024.105224
European Commission. (n.d.). AI literacy – Questions & answers. Shaping Europe’s digital future. Retrieved August 9, 2026, from https://digital-strategy.ec.europa.eu/en/faqs/ai-literacy-questions-answers
European Parliament & Council of the European Union. (2024). Regulation (EU) 2024/1689 of the European Parliament and of the Council of 13 June 2024 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act). Official Journal of the European Union. https://eur-lex.europa.eu/eli/reg/2024/1689/oj/eng
Fan, Y., Tang, L., Le, H., Shen, K., Tan, S., Zhao, Y., Shen, Y., Li, X., & Gašević, D. (2025). Beware of metacognitive laziness: Effects of generative artificial intelligence on learning motivation, processes, and performance. British Journal of Educational Technology, 56(2), 489–530. https://doi.org/10.1111/bjet.13544
Guan, T., Liu, F., Wu, X., Xian, R., Li, Z., Liu, X., Wang, X., Chen, L., Huang, F., Yacoob, Y., Manocha, D., & Zhou, T. (2024).
HallusionBench: An advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 14375–14385). https://doi.org/10.1109/CVPR52733.2024.01363
Huang, W., Liu, H., Guo, M., & Gong, N. Z. (2024). Visual hallucinations of multi-modal large language models. In Findings of the Association for Computational Linguistics: ACL 2024 (pp. 9614–9631). https://doi.org/10.18653/v1/2024.findings-acl.573
Istrate, O. (2026). Generative AI and multimodal pedagogy: Implications for teaching, learning, and assessment. Frontiers in Education, 11, 1828953. https://doi.org/10.3389/feduc.2026.1828953
Kress, G. (2010). Multimodality: A social semiotic approach to contemporary communication. Routledge.
Mašát, M., Kopecký, K., Krejčí, V., & Voráč, D. (2026). Anchored to the text, owned by the student: A policy & practice review for generative AI in literature education. Frontiers in Education, 11, 1805617. https://doi.org/10.3389/feduc.2026.1805617
Miao, F., & Cukurova, M. (2024). AI competency framework for teachers. UNESCO. https://unesdoc.unesco.org/ark:/48223/pf0000391104
Miao, F., & Holmes, W. (2023). Guidance for generative AI in education and research. UNESCO. https://unesdoc.unesco.org/ark:/48223/pf0000386693
Mikkonen, K. (2017). The narratology of comic art. Routledge.
Painter, C., Martin, J. R., & Unsworth, L. (2013). Reading visual narratives: Image analysis of children’s picture books. Equinox.
Sanz-Tejeda, A., Domínguez-Oller, J. C., Baldaquí-Escandell, J. M., Gómez-Díaz, R., & García-Rodríguez, A. (2026). The impact of generative AI on academic reading and writing: A synthesis of recent evidence (2023–2025). Frontiers in Education, 10, 1711718. https://doi.org/10.3389/feduc.2025.1711718
Shneiderman, B. (2022). Human-centered AI. Oxford University Press.
Torres Vico, M., & Fernández-Herrero, J. (2026). Impact of large multimodal language models (MLMs) with screen-vision in higher education: A quasi-experimental study. Technology, Knowledge and Learning. Advance online publication. https://doi.org/10.1007/s10758-026-10006-7
Yan, L., Greiff, S., Lodge, J. M., & Gašević, D. (2025). Distinguishing performance gains from learning when using generative AI. Nature Reviews Psychology, 4, 435–436. https://doi.org/10.1038/s44159-025-00467-5
Zhai, C., Wibowo, S., & Li, L. D. (2024). The effects of over-reliance on AI dialogue systems on students’ cognitive abilities: A systematic review. Smart Learning Environments, 11, 28. https://doi.org/10.1186/s40561-024-00316-7

