The moments at which a learner's understanding reorganises itself sit at the heart of teaching practice, yet they are among the hardest things to observe from the outside. This paper proposes reading the internal states of a dialogue model not as a measure of the model, but as an instrument for measuring the person in front of it, inverting surprisal theory from comprehension to production. As a first step, we built a preliminary system coupling a conference speakerphone and a PTZ camera to a cloud speech dialogue API, and characterised it as an observation instrument. The subject was the author alone (self-experimentation). Both the spoken dialogue loop and image-grounded dialogue were established: the noise floor measured ?63.4 dBFS against speech at ?22 to ?28 dBFS, and the model read back a string printed on the subject's shirt, information available only from that frame. Against this, we observed responses describing a visual input in detail when none had been sent, even though the recogniser had correctly registered its absence. A control turns under near-identical conditions asked appropriately for repetition instead, and nothing in the returned output distinguishes the two. The obstruction proved consistent across every stage of the observation chain: the microphone's signal processing exposes no parameters and cannot be disabled, and the API returns neither log probabilities for the learner's speech, nor hidden representations, nor attention weights. In the visual case, we could adjudicate the confabulation against the frame we had sent; a learner's inner state affords no such original. For an instrument, the inability to trace why a response was produced is disqualifying. What this research requires, then, is not better performance but ownership of the instrument.
Breazeal, C. (2003). Emotion and sociable humanoid robots. International Journal of Human-Computer Studies, 59(1–2), 119–155.
Chi, M. T. H. (1997). Quantifying qualitative analyses of verbal data: A practical guide. The Journal of the Learning Sciences, 6(3), 271–315.
Ericsson, K. A., & Simon, H. A. (1993). Protocol analysis: Verbal reports as data (Rev. ed.). MIT Press.
Frank, S. L., Otten, L. J., Galli, G., & Vigliocco, G. (2015). The ERP response to the amount of information conveyed by words in sentences. Brain and Language, 140, 1–11.
Hale, J. (2001). A probabilistic Earley parser as a psycholinguistic model. In Proceedings of the Second Meeting of the North American Chapter of the Association for Computational Linguistics (pp. 1–8).
Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., Ishii, E., Bang, Y., Madotto, A., & Fung, P. (2023). Survey of hallucination in natural language generation. ACM Computing Surveys, 55(12), 1–38.
Knowles, M. S. (1980). The modern practice of adult education: From pedagogy to andragogy (Rev. ed.). Cambridge Adult Education.
Kozima, H., Michalowski, M. P., & Nakagawa, C. (2009). Keepon: A playful robot for research, therapy, and entertainment. International Journal of Social Robotics, 1(1), 3–18.
Levy, R. (2008). Expectation-based syntactic comprehension. Cognition, 106(3), 1126–1177.
Maynez, J., Narayan, S., Bohnet, B., & McDonald, R. (2020). On faithfulness and factuality in abstractive summarisation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (pp. 1906–1919).
Pennebaker, J. W. (1997). Writing about emotional experiences as a therapeutic process. Psychological Science, 8(3), 162–166.
Reeves, B., & Nass, C. (1996). The media equation: How people treat computers, television, and new media like real people and places. Cambridge University Press.
Schön, D. A. (1983). The reflective practitioner: How professionals think in action. Basic Books.
Smith, N. J., & Levy, R. (2013). The effect of word predictability on reading time is logarithmic. Cognition, 128(3), 302–319.
Tomasello, M. (1995). Joint attention as social cognition. In C. Moore & P. J. Dunham (Eds.), Joint attention: Its origins and role in development (pp. 103–130). Lawrence Erlbaum Associates.
Ohshima, N. (2026). A Voice-and-Vision Dialogue System for Observing Emergent Cognition in Learners: Design Rationale and What the Cloud cannot Provide. International Journal of Academic Research in Business and Social Sciences, 16(9), 1088–1100.
Copyright: © 2026 The Author(s)
Published by Knowledge Words Publications (www.kwpublications.com)
This article is published under the Creative Commons Attribution (CC BY 4.0) license. Anyone may reproduce, distribute, translate and create derivative works of this article (for both commercial and non-commercial purposes), subject to full attribution to the original publication and authors. The full terms of this license may be seen at: http://creativecommons.org/licences/by/4.0/legalcode