Keeping track of how students behave in class is hard to do consistently, and right now it mostly comes down to a teacher's own judgment, formed by watching the room. Visual Question Answering (VQA) datasets rarely help here: most are English, built around everyday photos, or simply not designed with a classroom in mind. This paper presents Classroom VQA, a Vietnamese dataset built specifically for classroom learning-behavior understanding, comprising 3,000 classroom images drawn from Places 365 and 9,000 question–answer pairs produced and checked through a six-stage annotation pipeline. The questions span four categories – confirmation, object and attribute, counting, and action and state – chosen to cover different kinds of visual reasoning a classroom-monitoring system might need. A PhoBERT–EVA02 baseline is also proposed, using bidirectional Co-Cross-Attention so that Text-to-Image and Image-to-Text signals can inform one another. Tested across four EVA02 backbone sizes, this baseline reaches 70.78–71.44% accuracy, with under one percentage point separating the smallest and largest models. Classroom VQA is compared against six existing VQA datasets, and its question distribution, category-level error patterns, ethical considerations, and limitations are examined in detail. Together, the dataset and baseline give Vietnamese classroom VQA a concrete starting point and, more broadly, a foundation for education-oriented multimodal research.
Research Article | Open Access | Download Full Text
Volume 5 | Issue 3 | Year 2026 | Article Id: DST-V5I3P102 DOI: https://doi.org/10.59232/DST-V5I3P102
Classroom VQA: A Vietnamese Visual Question Answering Dataset and a PhoBERT–EVA02 Co-Cross-Attention Baseline for Classroom Learning-Behavior Monitoring
Huy Tran, Linh Luong
| Received | Revised | Accepted | Published |
|---|---|---|---|
| 25 May 2026 | 28 Jun 2026 | 20 Jul 2026 | 29 Aug 2026 |
Citation
Huy Tran, Linh Luong. “Classroom VQA: A Vietnamese Visual Question Answering Dataset and a PhoBERT–EVA02 Co-Cross-Attention Baseline for Classroom Learning-Behavior Monitoring.” DS Journal of Digital Science and Technology, vol. 5, no. 3, pp. 11-26, 2026.
Abstract
Keywords
Visual Question Answering dataset, Benchmark construction, Classroom behavior monitoring, Cross-Attention, EVA02, PhoBERT, Vietnamese NLP.
References
[1] William M. Baum, “What Counts as Behavior? The Molar Multiscale View,” The Behavior Analyst, vol. 36, no. 2, pp. 283-293, 2017.
[CrossRef] [Google Scholar] [Publisher Link]
[2] Qingtang Liu, Xinyu Jiang, and Ruyi Jiang, “Classroom Behavior Recognition using Computer Vision: A Systematic Review,” Sensors, vol. 25, no. 2, pp. 1-22, 2025.
[CrossRef] [Google Scholar] [Publisher Link]
[3] Stanislaw Antol et al., “VQA: Visual Question Answering,” Proceedings of the IEEE International Conference on Computer Vision (ICCV), pp. 2425-2433, 2015.
[Google Scholar] [Publisher Link]
[4] Yash Goyal et al., “Making the v in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 6904-6913, 2013.
[Google Scholar] [Publisher Link]
[5] Mengye Ren, Ryan Kiros, and Richard Zemel, “Exploring Models and Data for Image Question Answering,” Advances in Neural Information Processing Systems, vol. 28, pp. 1-9, 2015.
[Google Scholar] [Publisher Link]
[6] Justin Johnson et al., “CLEVR: A Diagnostic Dataset for Compositional Language and Elementary Visual Reasoning,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 2901-2910, 2017.
[Google Scholar] [Publisher Link]
[7] Jason J. Lau et al., “A Dataset of Clinically Generated Visual Questions and Answers about Radiology Images,” Scientific Data, vol. 5, no. 1, pp. 1-10, 2018.
[CrossRef] [Google Scholar] [Publisher Link]
[8] Khanh Quoc Tran et al., “ViVQA: Vietnamese Visual Question Answering,” Proceedings of the 35th Pacific Asia Conference on Language, Information and Computation, pp. 1-9, 2021.
[9] Dat Quoc Nguyen, and Anh Tuan Nguyen, “PhoBERT: Pre-Trained Language Models for Vietnamese,” Findings of the Association for Computational Linguistics: EMNLP 2020, pp. 1037-1042, 2020.
[CrossRef] [Google Scholar] [Publisher Link]
[10] Yuxin Fang et al., “EVA: Exploring the Limits of Masked Visual Representation Learning at Scale,” 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, BC, Canada, pp. 19358-19369, 2023.
[CrossRef] [Google Scholar] [Publisher Link]
[11] Byeong Su Kim et al., “Visual Question Answering: A Survey of Methods, Datasets, Evaluation, and Challenges,” ACM Computing Surveys, vol. 57, no. 10, pp. 1-35, 2025.
[CrossRef] [Google Scholar] [Publisher Link]
[12] Mateusz Malinowski, Marcus Rohrbach, and Mario Fritz, “Ask your Neurons: A Neural-based Approach to Answering Questions about Images,” Proceedings of the IEEE International Conference on Computer Vision (ICCV), pp. 1-9, 2015.
[Google Scholar] [Publisher Link]
[13] Haoyuan Gao et al., “Are you Talking to a Machine? Dataset and Methods for Multilingual Image Question,” Advances in Neural Information Processing Systems, vol. 28, 2015.
[Google Scholar] [Publisher Link]
[14] Jiayuan Mao et al., “The Neuro-Symbolic Concept Learner: Interpreting Scenes, Words, and Sentences from Natural Supervision,” arXiv preprint, pp. 1-28, 2019.
[CrossRef] [Google Scholar] [Publisher Link]
[15] Remi Cadene et al., “RUBi: Reducing Unimodal Biases for Visual Question Answering,” Advances in Neural Information Processing Systems, vol. 32, pp. 1-12, 2019.
[Google Scholar] [Publisher Link]
[16] Christopher Clark, Mark Yatskar, and Luke Zettlemoyer, “Don't Take the Easy Way Out: Ensemble based Methods for Avoiding Known Dataset Biases,” Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pp. 4069-4082, 2019.
[CrossRef] [Google Scholar] [Publisher Link]
[17] Mateusz Malinowski, and Mario Fritz, “A Multi-World Approach to Question Answering about Real-World Scenes based on Uncertain Input,” Advances in Neural Information Processing Systems, vol. 27, pp. 1-9, 2014.
[Google Scholar] [Publisher Link]
[18] Bolei Zhou et al., “Places: A 10 Million Image Database for Scene Recognition,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 40, no. 6, pp. 1452-1464, 2018.
[CrossRef] [Google Scholar] [Publisher Link]
[19] Kushal Kafle, and Christopher Kanan, “Visual Question Answering: Datasets, Algorithms, and Future Challenges,” Computer Vision and Image Understanding, vol. 163, pp. 3-20, 2017.
[CrossRef] [Google Scholar] [Publisher Link]
[20] Kiet Van Nguyen et al., “UIT-VSFC: Vietnamese Students’ Feedback Corpus for Sentiment Analysis,” 2018 10th International Conference on Knowledge and Systems Engineering (KSE), Ho Chi Minh City, Vietnam, pp. 19-24, 2018.