Evidence-Grounded Visual Explanation Design for E-commerce Recommendation Cards: Complementary and Substitute Relations, Image Availability, and Interface Hierarchy
DOI:
https://doi.org/10.51903/ijgd.v4i1.3992Keywords:
E-commerce recommendation, visual explanation, complementary products, substitute products, interface hierarchyAbstract
E-commerce recommendation cards must explain why an item is related while maintaining a compact, consistent visual hierarchy. This study presents an evidence-grounded card framework and evaluates its retrieval inputs with the final TREC Product Recommendation 2025 topics, a 48,616-product corpus, 7,633 NIST judgments for 47 topics, and the SQID supplementary image-URL table. Fixed lexical methods generated separate top-10 complement and substitute rankings over the full corpus before assessor judgments were applied. Under the official scoring procedure, the relation-aware lexical method achieved an average nDCG@10 of 0.1840, with a complement nDCG@10 of 0.0562 and a substitute nDCG@10 of 0.3119. Its average difference from full-text TF–IDF was not statistically reliable (difference = 0.0111, 95% CI [-0.0090, 0.0321]). The complement-minus-substitute gap was -0.2557 (95% CI [-0.3151, -0.1937]), confirming that complementary recommendation remains the weaker relation. Requiring a confirmed supplementary image URL increased confirmed URL coverage among displayed items to 100% but reduced average nDCG@10 to 0.0138. The findings support relation-qualified reasons, confirmed-image and text-led fallback states, and conservative language for complementary suggestions. The framework specifies evidence and layout behavior; shopper responses and different verbalization methods require direct comparative evaluation.
References
Al Ghossein, M., Chen, C.-W., & Tang, J. (2024). Shopping Queries Image Dataset (SQID): An image-enriched ESCI dataset for exploring multimodal learning in product search [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2405.15190
Bai, J., Wang, H., Wu, Q., & Zhang, B. (2026). Privacy-robust incrementality estimation in cookieless settings via uplift modeling: Reproducible evidence from the Hillstrom e-mail experiment. Journal of Technology Informatics and Engineering, 5(1), 17–38. https://doi.org/10.51903/jtie.v5i1.468
Bai, J., & Wu, Q. (2026). Privacy-safe marketing mix modeling and budget optimization under identifier loss: A controlled simulation study. International Journal of Electronics and Communications Systems, 6(1), 83–95. https://doi.org/10.24042/ijecs.v6i1.30533
Burke, R. (2002). Hybrid recommender systems: Survey and experiments. User Modeling and User-Adapted Interaction, 12(4), 331–370. https://doi.org/10.1023/A:1021240730564
Chang, X., Lu, Y., & Zhong, Z. S. (2026). Review-grounded explainable recommendation with faithfulness evaluation on Amazon reviews. Journal of Electrical Engineering and Computer Sciences, 11(1), 9–22. https://doi.org/10.54732/jeecs.v11i1.2
Chen, Y., & Xu, H. (2026). Trust-calibrated multilingual RAG for humanitarian information platforms: Empirical evaluation on OMoS-QA for migration information access. International Journal of Graphic Design, 4(1), 141–164. https://doi.org/10.51903/ijgd.v4i1.3552
Crossing Minds. (2026). Shopping Queries Image Dataset (SQID) [Data set]. Hugging Face. Retrieved July 23, 2026, from https://huggingface.co/datasets/crossingminds/shopping-queries-image-dataset
Go, E. M., & Mothelsang, K. (2024). Trends in Modern Typography Design: Visual Preferences on E-Commerce Platforms. International Journal of Graphic Design, 2(2), 248–263. https://doi.org/10.51903/IJGD.V2I2.2147
Järvelin, K., & Kekäläinen, J. (2002). Cumulative gain-based evaluation of IR techniques. ACM Transactions on Information Systems, 20(4), 422–446. https://doi.org/10.1145/582415.582418
Jin, J. (2025a). Calibrated resume-job matching for trustworthy LLM-assisted recruiter screening: Pairwise matching, probability calibration, and selective refusal on two public recruitment datasets. Journal of Technology Informatics and Engineering, 4(3), 625–648. https://doi.org/10.51903/jtie.v4i3.529
Jin, J. (2025b). Evidence-chain reliable RAG: Hallucination detection, source attribution, and deterministic provenance explanations. Journal of Technology Informatics and Engineering, 4(2), 520–533. https://doi.org/10.51903/jtie.v4i2.535
Jin, J. (2025c). LLM-style evidence cards for scientific search interfaces: A UI/UX design framework for retrieval transparency, ranking trust, and visual evidence hierarchy. International Journal of Graphic Design, 3(2), 397–414. https://doi.org/10.51903/ijgd.v3i2.3698
Koren, Y., Bell, R., & Volinsky, C. (2009). Matrix factorization techniques for recommender systems. Computer, 42(8), 30–37. https://doi.org/10.1109/MC.2009.263
Lu, Y., Zhou, H., & Zhang, Y. (2025). A constrained, data-driven budgeting framework integrating macro demand forecasting and marketing response modeling. Journal of Technology Informatics and Engineering, 4(3), 493–520. https://doi.org/10.51903/jtie.v4i3.466
Lundberg, S. M., & Lee, S.-I. (2017). A unified approach to interpreting model predictions. In I. Guyon, U. von Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, & R. Garnett (Eds.), Advances in neural information processing systems (Vol. 30, pp. 4765–4774). Curran Associates, Inc. https://proceedings.neurips.cc/paper_files/paper/2017/hash/8a20a8621978632d76c43dfd28b67767-Abstract.html
Manning, C. D., Raghavan, P., & Schütze, H. (2008). Introduction to information retrieval. Cambridge University Press.
Miller, T. (2019). Explanation in artificial intelligence: Insights from the social sciences. Artificial Intelligence, 267, 1–38. https://doi.org/10.1016/j.artint.2018.07.007
Mu, J., Lu, Y., & Hwang, E. (2026). Structured visual brief interfaces for advertising design: A UI/UX framework for turning creative intentions into designer-editable graphic design cards. International Journal of Graphic Design, 4(1), 192–208. https://doi.org/10.51903/ijgd.v4i1.3702
Mu, J., Ye, T., & Patel, P. (2025). Offline counterfactual evaluation for advertising and recommendation slot policies: A reproducible study on the Open Bandit Dataset (Small). Journal of Technology Informatics and Engineering, 4(3), 521–543. https://doi.org/10.51903/jtie.v4i3.500
National Institute of Standards and Technology. (2026, March 17). 2025 Product Search Track. https://trec.nist.gov/data/product2025.html
Nielsen, J. (1994). Enhancing the explanatory power of usability heuristics. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (CHI ’94) (pp. 152–158). Association for Computing Machinery. https://doi.org/10.1145/191666.191729
Norman, D. A. (2013). The design of everyday things (Rev. & expanded ed.). Basic Books.
Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., VanderPlas, J., Passos, A., Cournapeau, D., Brucher, M., Perrot, M., & Duchesnay, É. (2011). Scikit-learn: Machine learning in Python. Journal of Machine Learning Research, 12(85), 2825–2830. https://www.jmlr.org/papers/v12/pedregosa11a.html
Reddy, C. K., Màrquez, L., Valero, F., Rao, N., Zaragoza, H., Bandyopadhyay, S., Biswas, A., Xing, A., & Subbian, K. (2022). Shopping Queries Dataset: A large-scale ESCI benchmark for improving product search [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2206.06588
Ribeiro, M. T., Singh, S., & Guestrin, C. (2016). “Why should I trust you?”: Explaining the predictions of any classifier. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 1135–1144). Association for Computing Machinery. https://doi.org/10.1145/2939672.2939778
Ricci, F., Rokach, L., & Shapira, B. (Eds.). (2015). Recommender systems handbook (2nd ed.). Springer. https://doi.org/10.1007/978-1-4899-7637-6
Robertson, S., & Zaragoza, H. (2009). The probabilistic relevance framework: BM25 and beyond. Foundations and Trends in Information Retrieval, 3(4), 333–389. https://doi.org/10.1561/1500000019
Su, W., Chen, S., & Qian, E. (2026). Narrative-aware scientific claim verification agent with evidence ranking for ClimateCheck. Journal of Technology Informatics and Engineering, 5(1), 327–340. https://doi.org/10.51903/jtie.v5i1.549
Su, W., Chen, S., & Zhao, C. (2025). Budgeted multi-hop retrieval agent for compositional question answering: A retrieval-policy evaluation on the official MultiHop-RAG benchmark. Journal of Technology Informatics and Engineering, 4(3), 649–662. https://doi.org/10.51903/jtie.v4i3.543
Su, W., Rao, H., & Ma, E. (2026). Privacy and data-integrity risk cards for LLM agents: A UI/UX design framework for secure human oversight under prompt-injection attacks. International Journal of Graphic Design, 4(1), 186–191. https://doi.org/10.51903/ijgd.v4i1.3699
Sweller, J. (1988). Cognitive load during problem solving: Effects on learning. Cognitive Science, 12(2), 257–285. https://doi.org/10.1207/s15516709cog1202_4
Tintarev, N., & Masthoff, J. (2015). Explaining recommendations: Design and evaluation. In F. Ricci, L. Rokach, & B. Shapira (Eds.), Recommender systems handbook (2nd ed., pp. 353–382). Springer. https://doi.org/10.1007/978-1-4899-7637-6_10
TREC Product Search Track. (2025). Product recommendation 2025 [Data set]. Hugging Face. https://huggingface.co/datasets/trec-product-search/product-recommendation-2025
TREC Product Search Track Organizers. (2025). Recommendation task instructions. https://trec-product-search.github.io/recommendations
Tufte, E. R. (2001). The visual display of quantitative information (2nd ed.). Graphics Press.
Wang, H., Ren, Y., & Chang, X. (2025). Layout-aware progressive PDF rendering: AI prioritization of PDF slices to reduce time-to-functional-first-frame on FUNSD. Journal of Technology Informatics and Engineering, 4(2), 425–446. https://doi.org/10.51903/jtie.v4i2.523
Xin, Q. (2026a). Behavior retrieval plus response generation for interpretable conversational personalized recommendation. International Journal of Electrical, Energy and Power System Engineering, 9(2), 120–136. https://doi.org/10.31258/ijeepse.9.2.120-136
Xin, Q. (2026b). Self-supervised customer representation learning for segmentation and next-purchase prediction on UCI Online Retail. Journal of Information and Technology, 14(1), 20–37. https://doi.org/10.32664/j-intech.v14i01.2229
Xu, H., Chen, Y., & Med, A. (2025). Automatic detection and explanation of dark patterns from interface microcopy: Empirical comparison of BERT-style encoders, RoBERTa-style encoders, and LLM-style decoders on the ec-darkpattern dataset. Journal of Technology Informatics and Engineering, 4(3), 590–612. https://doi.org/10.51903/jtie.v4i3.491
Xu, K., Zhou, H., Zheng, H., Zhu, M., & Xin, Q. (2024). Intelligent classification and personalized recommendation of e-commerce products based on machine learning. Applied and Computational Engineering, 64(1), 147–153. https://doi.org/10.54254/2755-2721/64/20241365
Ye, T., Mu, J., & Hunter, J. (2026). Off-policy evaluation and conservative policy selection for slot-level dynamic bidding and ranking on the Open Bandit Dataset (Small). Journal of Technology Informatics and Engineering, 5(1), 178–199. https://doi.org/10.51903/jtie.v5i1.503
Zhang, B., Ren, Y., & Zou, J. (2025). LLM-style explainable e-commerce recommendation cards: A UI/UX design framework for trust-calibrated product recommendation. International Journal of Graphic Design, 3(2), 381–396. https://doi.org/10.51903/ijgd.v3i2.3697
Zhang, Y., & Chen, X. (2020). Explainable recommendation: A survey and new perspectives. Foundations and Trends in Information Retrieval, 14(1), 1–101. https://doi.org/10.1561/1500000066
Zhang, Y., & Zhang, H. (2025a). A therapist-facing session copilot for live counseling support: Reasoning-guided retrieval and ranking from multi-turn counseling dialogues. Journal of Technology Informatics and Engineering, 4(2), 464–486. https://doi.org/10.51903/jtie.v4i2.547
Zhang, Y., & Zhang, H. (2025b). Visualizing the right counseling support: Evidence-linked recommendation cards for explainable mental health intake interfaces. International Journal of Graphic Design, 3(1), 214–229. https://doi.org/10.51903/ijgd.v3i1.3722
Zhong, Z. S., Li, C., & Rao, H. (2026). Trajectory reliability prediction for generalist AI agents: Tool-use failure analysis and success forecasting on ZClawBench. Journal of Technology Informatics and Engineering, 5(1), 341–360. https://doi.org/10.51903/jtie.v5i1.539
Zhou, B., Li, C., & Liu, L. (2025). Risk-calibrated patient-facing AI safety cards: A UI/UX design framework for rubric-based medical risk communication. International Journal of Graphic Design, 3(2), 365–380. https://doi.org/10.51903/ijgd.v3i2.3696
Zhou, B., Wang, H., & Chang, X. (2025). Distilling VMAF into an edge-deployable quality predictor: A pilot shot-level proxy with LLM-ready quality tokens. Journal of Technology Informatics and Engineering, 4(2), 447–463. https://doi.org/10.51903/jtie.v4i2.522
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Hailin Zhou

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.









5.png)
