Computer-Vision-Informed Visual Explanation Cards for Autonomous-Driving Traffic-Sign Alerts: Localization, Classification, and Retrieved Evidence on GTSDB

Authors

  • Ruiyan Ma Software Engineering, UC Irvine, CA, USA
  • Long Zhang Transportation Systems Engineering, Southern Methodist University, TX, USA
  • Tiffany Song Human-Computer Interaction Design, Indiana University Bloomington, Bloomington, IN, USA

DOI:

https://doi.org/10.51903/ijgd.v3i2.3991

Keywords:

traffic-sign detection, autonomous driving, computer vision, visual explanation, graphic design

Abstract

This paper develops a computer-vision-informed visual explanation card for traffic-sign alerts in autonomous-driving and driver-assistance interfaces. The card organizes a detected sign crop, predicted class, confidence, default semantic display-priority tier, retrieved visual precedents, scene-location cue, and concise action prompt. The empirical study uses the complete German Traffic Sign Detection Benchmark (GTSDB), comprising 900 road scenes in the standard 600-scene development and 300-scene evaluation portions. The same scenes support localization, crop classification, retrieval, calibration, and end-to-end analysis. The first 600 scenes were divided at the scene level into training and validation subsets; the 300 evaluation scenes were held out until model choices, retrieval depth, fusion weight, and detector threshold had been fixed. A learned class-agnostic localizer filters color-connected-component proposals with a histogram-of-oriented-gradients and color classifier. Five crop classifiers and four nearest-neighbor settings were evaluated, with retrieval treated primarily as example-based explanation support. On the held-out scenes, the selected localizer achieved an AP at IoU 0.50 of 0.211, an AP averaged over IoU 0.50–0.95 of 0.098, and a recall of 0.260 at the validation-selected operating point. The selected crop classifier achieved 0.784 accuracy and 0.574 macro-F1 on ground-truth crops. With predicted crops, correct-class end-to-end coverage was 0.177, and correct-tier end-to-end coverage was 0.244. These results define the information that the proposed card can receive from the evaluated vision pipeline. They do not measure driver comprehension, glance behavior, response time, trust, usability, or deployment safety, which require separate human-centred evaluation.

References

Aamodt, A., & Plaza, E. (1994). Case-based reasoning: Foundational issues, methodological variations, and system approaches. AI Communications, 7(1), 39–59. https://doi.org/10.3233/AIC-1994-7104

Amershi, S., Weld, D., Vorvoreanu, M., Fourney, A., Nushi, B., Collisson, P., Suh, J., Iqbal, S. T., Bennett, P. N., Inkpen, K., Teevan, J., Kikin-Gil, R., & Horvitz, E. (2019). Guidelines for human-AI interaction. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems (pp. 1–13). Association for Computing Machinery. https://doi.org/10.1145/3290605.3300233

Ben-Bassat, T., & Shinar, D. (2006). Ergonomic guidelines for traffic sign design increase sign comprehension. Human Factors, 48(1), 182–195. https://doi.org/10.1518/001872006776412298

Breiman, L. (2001). Random forests. Machine Learning, 45(1), 5–32. https://doi.org/10.1023/A:1010933404324

Campbell, J. L., Brown, J. L., Graving, J. S., Richard, C. M., Lichty, M. G., Sanquist, T., Bacon, L. P., Woods, R., Li, H., Williams, D. N., & Morgan, J. F. (2016). Human factors design guidance for driver-vehicle interfaces (Report No. DOT HS 812 360). National Highway Traffic Safety Administration. https://www.nhtsa.gov/sites/nhtsa.gov/files/documents/812360_humanfactorsdesignguidance.pdf

Chen, C., Li, O., Tao, D., Barnett, A., Rudin, C., & Su, J. K. (2019). This looks like that: Deep learning for interpretable image recognition. Advances in Neural Information Processing Systems, 32, 8930–8941. https://proceedings.neurips.cc/paper/2019/hash/adf7ee2dcf142b0e11888e72b43fcb75-Abstract.html

Chen, Y., & Li, M. (2025). From hand-drawn sketches to interactive web prototypes: A reproducible vision-language approach with structural and visual consistency evaluation. Journal of Technology Informatics and Engineering, 4(2), 364–384. https://doi.org/10.51903/jtie.v4i2.490

Chen, Y., & Xu, H. (2026). Trust-calibrated multilingual RAG for humanitarian information platforms: Empirical evaluation on OMoS-QA for migration information access. International Journal of Graphic Design, 4(1), 141–164. https://doi.org/10.51903/ijgd.v4i1.3552

Cortes, C., & Vapnik, V. (1995). Support-vector networks. Machine Learning, 20(3), 273–297. https://doi.org/10.1007/BF00994018

Dalal, N., & Triggs, B. (2005). Histograms of oriented gradients for human detection. In 2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (Vol. 1, pp. 886–893). IEEE. https://doi.org/10.1109/CVPR.2005.177

Endsley, M. R. (1995). Toward a theory of situation awareness in dynamic systems. Human Factors, 37(1), 32–64. https://doi.org/10.1518/001872095779049543

Guo, C., Pleiss, G., Sun, Y., & Weinberger, K. Q. (2017). On calibration of modern neural networks. In D. Precup & Y. W. Teh (Eds.), Proceedings of the 34th International Conference on Machine Learning (Vol. 70, pp. 1321–1330). PMLR. https://proceedings.mlr.press/v70/guo17a.html

Houben, S., Stallkamp, J., Salmen, J., Schlipsing, M., & Igel, C. (2013). Detection of traffic signs in real-world images: The German Traffic Sign Detection Benchmark. In 2013 International Joint Conference on Neural Networks (pp. 1–8). IEEE. https://doi.org/10.1109/IJCNN.2013.6706807

International Organization for Standardization. (2021). Road vehicles—Ergonomic aspects of transport information and control systems (TICS)—Procedures for determining priority of on-board messages presented to drivers (ISO/TS 16951:2021). https://www.iso.org/standard/81103.html

Jin, J. (2025a). Evidence-chain reliable RAG: Hallucination detection, source attribution, and deterministic provenance explanations. Journal of Technology Informatics and Engineering, 4(2), 520–533. https://doi.org/10.51903/jtie.v4i2.535

Jin, J. (2025b). LLM-style evidence cards for scientific search interfaces: A UI/UX design framework for retrieval transparency, ranking trust, and visual evidence hierarchy. International Journal of Graphic Design, 3(2), 397–414. https://doi.org/10.51903/ijgd.v3i2.3698

Kim, B., Khanna, R., & Koyejo, O. O. (2016). Examples are not enough, learn to criticize! Criticism for interpretability. Advances in Neural Information Processing Systems, 29, 2280–2288. https://proceedings.neurips.cc/paper_files/paper/2016/hash/5680522b8e2bb01943234bce7bf84534-Abstract.html

Koo, J., Kwac, J., Ju, W., Steinert, M., Leifer, L., & Nass, C. (2015). Why did my car just do that? Explaining semi-autonomous driving actions to improve driver understanding, trust, and performance. International Journal on Interactive Design and Manufacturing, 9, 269–275. https://doi.org/10.1007/s12008-014-0227-2

Lee, J. D., & See, K. A. (2004). Trust in automation: Designing for appropriate reliance. Human Factors, 46(1), 50–80. https://doi.org/10.1518/hfes.46.1.50_30392

Li, C., Zhou, B., & Gao, K. (2025). Risk-calibrated patient-facing AI safety cards: A UI/UX benchmark for explainable medical AI response interfaces. International Journal of Graphic Design, 3(2), 381–394. https://doi.org/10.51903/ijgd.v3i2.3709

Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., & Zitnick, C. L. (2014). Microsoft COCO: Common objects in context. In D. Fleet, T. Pajdla, B. Schiele, & T. Tuytelaars (Eds.), Computer vision—ECCV 2014 (pp. 740–755). Springer. https://doi.org/10.1007/978-3-319-10602-1_48

Mendez, L., & Okafor, S. (2026). Adaptive Graphic Interaction Model: A Mixed-Method Framework for Future Factory Design. International Journal of Graphic Design, 4(1), 1–16. https://doi.org/10.51903/IJGD.V4I1.3194

Munzner, T. (2014). Visualization analysis and design. CRC Press.

Norman, D. A. (2013). The design of everyday things (Rev. and expanded ed.). Basic Books.

Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., Vanderplas, J., Passos, A., Cournapeau, D., Brucher, M., Perrot, M., & Duchesnay, É. (2011). Scikit-learn: Machine learning in Python. Journal of Machine Learning Research, 12, 2825–2830.

Ribeiro, M. T., Singh, S., & Guestrin, C. (2016). “Why should I trust you?”: Explaining the predictions of any classifier. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 1135–1144). Association for Computing Machinery. https://doi.org/10.1145/2939672.2939778

Ware, C. (2012). Information visualization: Perception for design (3rd ed.). Morgan Kaufmann.

Wickens, C. D., Lee, J. D., Liu, Y., & Gordon-Becker, S. E. (2004). An introduction to human factors engineering (2nd ed.). Pearson Prentice Hall.

Wogalter, M. S. (Ed.). (2006). Handbook of warnings. Lawrence Erlbaum Associates.

Xin, Q. (2025). Uncertainty-aware late fusion for 3D perception (confidence calibration + fusion rule learning). Journal of Technology Informatics and Engineering, 4(1), 215–238. https://doi.org/10.51903/jtie.v4i1.485

Xin, Q. (2026). LiDAR–camera object-level fusion for multi-target tracking using JPDA and EKF: A reproducible empirical study on a PandaSet-parameterised five-sequence dataset. Journal of Technology Informatics and Engineering, 5(1), 54–76. https://doi.org/10.51903/jtie.v5i1.486

Xu, H., Chen, Y., & Med, A. (2025). Automatic detection and explanation of dark patterns from interface microcopy: Empirical comparison of BERT-style encoders, RoBERTa-style encoders, and LLM-style decoders on the ec-darkpattern dataset. Journal of Technology Informatics and Engineering, 4(3), 590–612. https://doi.org/10.51903/jtie.v4i3.491

Zhang, Y., & Zhang, H. (2025a). A therapist-facing session copilot for live counseling support: Reasoning-guided retrieval and ranking from multi-turn counseling dialogues. Journal of Technology Informatics and Engineering, 4(2), 464–486. https://doi.org/10.51903/jtie.v4i2.547

Zhang, Y., & Zhang, H. (2025b). Visualizing the right counseling support: Evidence-linked recommendation cards for explainable mental health intake interfaces. International Journal of Graphic Design, 3(1), 214–229. https://doi.org/10.51903/ijgd.v3i1.3722

Zhou, B., Jin, J., & Zhao, D. (2025). Calibrated resume-job matching for trustworthy LLM-assisted recruiter screening: Pairwise matching, probability calibration, and selective refusal on two public recruitment datasets. Journal of Technology Informatics and Engineering, 4(3), 625–648. https://doi.org/10.51903/jtie.v4i3.529

Zhou, B., Li, C., & Liu, L. (2025). Risk-calibrated patient-facing AI safety cards: A UI/UX design framework for rubric-based medical risk communication. International Journal of Graphic Design, 3(2), 365–380. https://doi.org/10.51903/ijgd.v3i2.3696

Downloads

Published

2025-10-22

How to Cite

Computer-Vision-Informed Visual Explanation Cards for Autonomous-Driving Traffic-Sign Alerts: Localization, Classification, and Retrieved Evidence on GTSDB. (2025). International Journal of Graphic Design, 3(2), 455-473. https://doi.org/10.51903/ijgd.v3i2.3991