Beyond CTR and VCR: LLM-Assisted Design Evaluation with Eye-Tracking Validation for Graphic Advertising
DOI:
https://doi.org/10.51903/ijgd.v4i1.3994Keywords:
graphic design evaluation, advertising banner, LLM-assisted scoring, eye tracking, fixation-density map, visual hierarchyAbstract
Digital advertising design is often judged with delivery or response metrics such as click-through rate (CTR), video-completion rate (VCR), and viewability, although these measures do not describe whether a banner communicates through a clear visual hierarchy. This study evaluates whether large language model (LLM)-assisted scores and lightweight image features can support graphic advertising review while keeping design-quality prediction distinct from measured visual attention. GraphicDesignEvaluation provided 1,200 principle–image rating rows for alignment, overlap, and white space. Under grouped five-fold cross-validation by base design, the GPT-only Ridge model reached R2 = 0.408, RMSE = 1.566, MAE = 1.257, Pearson = 0.639, and Spearman = 0.634. Direct GPT–human agreement was uneven: overlap was strong (Spearman = 0.769), white space was moderate (0.637), and alignment was weaker (0.519). An external construct analysis used 1,000 advertisements from ADD1000 with aggregate human fixation-density maps and matched EAID ratings. A combined gradient and local-contrast proxy modestly exceeded a center prior (CC = 0.457, SIM = 0.562, KL divergence = 0.637), while lower fixation entropy was associated with higher aesthetic ratings (Spearman = -0.319). BannerRequest400 was used only to describe brief, CTA-language, logo, and format constraints because it contains no human design-quality or gaze labels. The findings support a preliminary, human-supervised design-review framework: LLM scores are useful screening evidence, especially for overlap, but they do not replace human creative judgment, eye tracking, or advertising-performance measures.
References
Bai, J., Wang, H., Wu, Q., & Zhang, B. (2026). Privacy-robust incrementality estimation in cookieless settings via uplift modeling: Reproducible evidence from the Hillstrom E-mail experiment. Journal of Technology Informatics and Engineering, 5(1), 17–38. https://doi.org/10.51903/jtie.v5i1.468
Bai, J., & Wu, Q. (2026). Privacy-safe marketing mix modeling and budget optimization under identifier loss: A controlled simulation study. International Journal of Electronics and Communications System, 6(1), 83–95. https://doi.org/10.24042/ijecs.v6i1.30533
Breiman, L. (2001). Random forests. Machine Learning, 45(1), 5–32. https://doi.org/10.1023/A:1010933404324
Chang, X., Lu, Y., & Zhong, Z. S. (2026). Review-grounded explainable recommendation with faithfulness evaluation on Amazon Reviews. Journal of Electrical Engineering and Computer Science, 11(1), 9–22. https://doi.org/10.54732/jeecs.v11i1.2
Chang, X., Ye, T., & Luo, S. (2023). LLM-as-reranker for personalized recommendation: Popularity bias mitigation and faithful natural-language explanations on MovieLens 100K. Journal of Advanced Computing Systems, 3(8), 61–78. https://doi.org/10.69987/JACS.2023.30806
Chen, T., & Guestrin, C. (2016). XGBoost: A scalable tree boosting system. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 785–794). https://doi.org/10.1145/2939672.2939785
Chen, Y., & Li, M. (2025). From hand-drawn sketches to interactive web prototypes: A reproducible vision-language approach with structural and visual consistency evaluation. Journal of Technology Informatics and Engineering, 4(2), 364–384. https://doi.org/10.51903/jtie.v4i2.490
Chen, Y., & Xu, H. (2026). Trust-calibrated multilingual RAG for humanitarian information platforms: Empirical evaluation on OMoS-QA for migration information access. International Journal of Graphic Design, 4(1), 141–164. https://doi.org/10.51903/ijgd.v4i1.3552
Haraguchi, D., Inoue, N., Shimoda, W., Mitani, H., Uchida, S., & Yamaguchi, K. (2024). Can GPTs evaluate graphic design based on design principles? In SIGGRAPH Asia 2024 Technical Communications (Article 5, pp. 1–4). Association for Computing Machinery. https://doi.org/10.1145/3681758.3698010
IAB Europe. (2023). Guide to attention in digital marketing. https://iabeurope.eu/wp-content/uploads/IAB-Europe-Guide-to-Attention-in-Digital-Marketing-25.05-1-1.pdf
Inoue, N., Kikuchi, K., Simo-Serra, E., Otani, M., & Yamaguchi, K. (2023). LayoutDM: Discrete diffusion model for controllable layout generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 10167–10176). https://doi.org/10.1109/CVPR52729.2023.00980
Interactive Advertising Bureau. (2024). Attention measurement explainer: Data signal approaches. https://www.iab.com/wp-content/uploads/2024/08/IAB_Attention_Measurement_Explainer_August_2024.pdf
Itti, L., Koch, C., & Niebur, E. (1998). A model of saliency-based visual attention for rapid scene analysis. IEEE Transactions on Pattern Analysis and Machine Intelligence, 20(11), 1254–1259. https://doi.org/10.1109/34.730558
Jia, P., Li, C., Yuan, Y., Liu, Z., Shen, Y., Chen, B., Chen, X., Zheng, Y., Chen, D., Li, J., Xie, X., Zhang, S., & Guo, B. (2023). COLE: A hierarchical generation framework for multi-layered and editable graphic design [Preprint]. arXiv. https://arxiv.org/abs/2311.16974
Jin, J. (2025a). Evidence-chain reliable RAG: Hallucination detection, source attribution, and deterministic provenance explanations. Journal of Technology Informatics and Engineering, 4(2), 520–533. https://doi.org/10.51903/jtie.v4i2.535
Jin, J. (2025b). LLM-style evidence cards for scientific search interfaces: A UI/UX design framework for retrieval transparency, ranking trust, and visual evidence hierarchy. International Journal of Graphic Design, 3(2), 397–414. https://doi.org/10.51903/ijgd.v3i2.3698
Kong, W., Jiang, Z., Sun, S., Guo, Z., Cui, W., Liu, T., Lou, J., & Zhang, D. (2023). Aesthetics++: Refining graphic designs by exploring design principles and human preference. IEEE Transactions on Visualization and Computer Graphics, 29(6), 3093–3104. https://doi.org/10.1109/TVCG.2022.3151617
Kuo, M.-J., Zheng, D., & Hires, J. (2025). Federated topic-preference learning for knowledge-grounded chat with differential privacy. Journal of Technology Informatics and Engineering, 4(2), 385–401. https://doi.org/10.51903/jtie.v4i2.502
Lavie, N. (2005). Distracted and confused? Selective attention under load. Trends in Cognitive Sciences, 9(2), 75–82. https://doi.org/10.1016/j.tics.2004.12.004
Li, C., Bai, J., & Wang, S. (2024). Evidence-chain reliable RAG: Word-level hallucination detection, source attribution, and provenance explanation for LLM applications. Journal of Advanced Computing Systems, 4(2), 76–92. https://doi.org/10.69987/JACS.2024.40207
Li, C., Su, W., & Zhang, E. (2023). Lightweight hallucination firewall for enterprise LLM applications: Evidence consistency, self-checking, and small-model detection on TruthfulQA. Journal of Advanced Computing Systems, 3(1), 49–65. https://doi.org/10.69987/JACS.2023.30104
Li, Y., Lu, S., & Zhao, L. (2025). LLM-as-design-critic: Aligning AI-generated UI feedback with human graphic design judgment. International Journal of Graphic Design, 3(1), 196–215. https://doi.org/10.51903/ijgd.v3i1.3661
Li, Z., Zhou, S., & Zhou, Z. (2025). Financial risk dashboard design for institutional RWA investors: Visual hierarchy, chart comprehension, and explainability in FinChart-Bench. International Journal of Graphic Design, 3(1), 196–210. https://doi.org/10.51903/ijgd.v3i1.3715
Liang, S., Liu, R., & Qian, J. (2021). Fixation prediction for advertising images: Dataset and benchmark. Journal of Visual Communication and Image Representation, 81, 103356. https://doi.org/10.1016/j.jvcir.2021.103356
Liang, S., Liu, R., & Qian, J. (2024). EAID: An eye-tracking based advertising image dataset with personalized affective tags. In B. Sheng, L. Bi, J. Kim, N. Magnenat-Thalmann, & D. Thalmann (Eds.), Advances in computer graphics: CGI 2023 (Lecture Notes in Computer Science, Vol. 14495, pp. 282–294). Springer. https://doi.org/10.1007/978-3-031-50069-5_24
Lidwell, W., Holden, K., & Butler, J. (2010). Universal principles of design (2nd ed.). Rockport Publishers.
Liu, G., He, S., & Wong, H. (2025). LLM-compatible visual brief cards for AI infrastructure capacity dashboards: A UI/UX framework for turning forecast risk into graphic design decisions. International Journal of Graphic Design, 3(1), 196–213. https://doi.org/10.51903/ijgd.v3i1.3723
Lu, S., & Zhou, D. (2023). LLM-augmented customer representation learning for next-purchase prediction in online retail. Journal of Advanced Computing Systems, 3(3), 50–64. https://doi.org/10.69987/JACS.2023.30305
Lu, Y., Mu, J., & Tran, T. (2024). Uncertainty-aware uplift modeling for safer marketing targeting: Conformal prediction and Bayesian calibration with LCB policies. Journal of Advanced Computing Systems, 4(5), 84–101. https://doi.org/10.69987/JACS.2024.40507
Lu, Y., Zhou, H., & Zhang, Y. (2025). A constrained, data-driven budgeting framework integrating macro demand forecasting and marketing response modeling. Journal of Technology Informatics and Engineering, 4(3), 493–520. https://doi.org/10.51903/jtie.v4i3.466
Lupton, E. (2010). Thinking with type: A critical guide for designers, writers, editors, & students (2nd ed.). Princeton Architectural Press.
Meng, S., Chen, J., & Zheng, I. (2026). LLM-inspired offline reranking for financial search: Query rewriting, hybrid retrieval, and listwise relevance ranking on FiQA. Journal of Technology Informatics and Engineering, 5(1), 361–378. https://doi.org/10.51903/jtie.v5i1.537
Mu, J., Lu, Y., & Hwang, E. (2026). Structured visual brief interfaces for advertising design: A UI/UX framework for turning creative intentions into designer-editable graphic design cards. International Journal of Graphic Design, 4(1), 192–208. https://doi.org/10.51903/ijgd.v4i1.3702
Mu, J., Lu, Y., & Smith, M. (2023). LLM-assisted incrementality (uplift) modeling for email advertising: From feature interactions to interpretable audience-creative-channel policies. Journal of Advanced Computing Systems, 3(1), 31–48. https://doi.org/10.69987/JACS.2023.30103
Mu, J., Ye, T., & Patel, P. (2025). Offline counterfactual evaluation for advertising and recommendation slot policies: A reproducible study on the Open Bandit Dataset (Small). Journal of Technology Informatics and Engineering, 4(3), 521–543. https://doi.org/10.51903/jtie.v4i3.500
Nie, J., Liu, G., Li, C., & Zou, T. (2026). Evidence-constrained incident visualization cards for distributed cloud logs: A UI/UX framework for turning Hadoop, OpenStack, and ZooKeeper logs into actionable SRE design interfaces. International Journal of Graphic Design, 4(1), 179–185. https://doi.org/10.51903/ijgd.v4i1.3703
O'Donovan, P., Agarwala, A., & Hertzmann, A. (2014). Learning layouts for single-page graphic designs. IEEE Transactions on Visualization and Computer Graphics, 20(8), 1200–1213. https://doi.org/10.1109/TVCG.2014.48
Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., Vanderplas, J., Passos, A., Cournapeau, D., Brucher, M., Perrot, M., & Duchesnay, E. (2011). Scikit-learn: Machine learning in Python. Journal of Machine Learning Research, 12, 2825–2830.
Pieters, R., & Wedel, M. (2004). Attention capture and transfer in advertising: Brand, pictorial, and text-size effects. Journal of Marketing, 68(2), 36–50. https://doi.org/10.1509/jmkg.68.2.36.27794
Pieters, R., Wedel, M., & Batra, R. (2010). The stopping power of advertising: Measures and effects of visual complexity. Journal of Marketing, 74(5), 48–60. https://doi.org/10.1509/jmkg.74.5.48
Rosenholtz, R., Li, Y., & Nakano, L. (2007). Measuring visual clutter. Journal of Vision, 7(2), Article 17. https://doi.org/10.1167/7.2.17
Solikhan, M., Priyadi, A., MacDonald, D., & Noramai, A. (2026). Generative AI as a Live Design Mentor: A Mixed-Reality Approach to Graphic Design Education. International Journal of Graphic Design, 4(1), 123–140. https://doi.org/10.51903/IJGD.V4I1.2375
Su, W., Chen, S., & Qian, E. (2026). Narrative-aware scientific claim verification agent with evidence ranking for ClimateCheck. Journal of Technology Informatics and Engineering, 5(1), 327–340. https://doi.org/10.51903/jtie.v5i1.549
Su, W., Chen, S., & Zhao, C. (2025). Budgeted multi-hop retrieval agent for compositional question answering: A retrieval-policy evaluation on the official MultiHop-RAG benchmark. Journal of Technology Informatics and Engineering, 4(3), 649–662. https://doi.org/10.51903/jtie.v4i3.543
Su, W., Rao, H., & Ma, E. (2026). Privacy and data-integrity risk cards for LLM agents: A UI/UX design framework for secure human oversight under prompt-injection attacks. International Journal of Graphic Design, 4(1), 186–191. https://doi.org/10.51903/ijgd.v4i1.3699
Sun, J., Xiao, X., & Dong, H. (2023). From tell-to-design to healthcare test-fit constraint checking: An LLM-compatible requirements-to-constraints framework. Journal of Advanced Computing Systems, 3(11), 52–67. https://doi.org/10.69987/JACS.2023.31105
Treisman, A. M., & Gelade, G. (1980). A feature-integration theory of attention. Cognitive Psychology, 12(1), 97–136. https://doi.org/10.1016/0010-0285(80)90005-5
Tu, H., Zhao, S., & Zhou, A. (2025). Visual brief cards for advertising design: A structured UI/UX framework for turning creative intentions into graphic design decisions. International Journal of Graphic Design, 3(1), 210–226. https://doi.org/10.51903/ijgd.v3i1.3714
Tufte, E. R. (2001). The visual display of quantitative information (2nd ed.). Graphics Press.
Wang, H., Ren, Y., & Chang, X. (2025). Layout-aware progressive PDF rendering: AI prioritization of PDF slices to reduce time-to-functional-first-frame on FUNSD. Journal of Technology Informatics and Engineering, 4(2), 425–446. https://doi.org/10.51903/jtie.v4i2.523
Wang, H., Shimose, Y., & Takamatsu, S. (2025). BannerAgency: Advertising banner design with multimodal LLM agents. In C. Christodoulopoulos, T. Chakraborty, C. Rose, & V. Peng (Eds.), Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing (pp. 4304–4329). Association for Computational Linguistics. https://doi.org/10.18653/v1/2025.emnlp-main.214
Wedel, M., & Pieters, R. (2008). A review of eye-tracking research in marketing. In N. K. Malhotra (Ed.), Review of marketing research (Vol. 4, pp. 123–147). Emerald Group Publishing. https://doi.org/10.1108/S1548-6435(2008)0000004009
Williams, R. (2014). The non-designer's design book (4th ed.). Peachpit Press.
Wolfe, J. M., & Horowitz, T. S. (2004). What attributes guide the deployment of visual attention and how do they do it? Nature Reviews Neuroscience, 5, 495–501. https://doi.org/10.1038/nrn1411
Wu, Q., Meng, S., & Zhao, J. (2025). Text-grounded LLM-assisted design rationale interfaces: Turning advertising layout metadata into explainable UI/UX decision cards. International Journal of Graphic Design, 3(1), 216–240. https://doi.org/10.51903/ijgd.v3i1.3713
Xin, Q. (2026). Self-supervised customer representation learning for segmentation and next-purchase prediction on UCI Online Retail. Journal of Information and Technology, 14(1), 20–37. https://doi.org/10.32664/j-intech.v14i01.2229
Xu, H., Chen, Y., & Med, A. (2025). Automatic detection and explanation of dark patterns from interface microcopy: Empirical comparison of BERT-style encoders, RoBERTa-style encoders, and LLM-style decoders on the ec-darkpattern Dataset. Journal of Technology Informatics and Engineering, 4(3), 590–612. https://doi.org/10.51903/jtie.v4i3.491
Xu, K., Zhou, H., Zheng, H., Zhu, M., & Xin, Q. (2024). Intelligent classification and personalized recommendation of e-commerce products based on machine learning. In Proceedings of the 6th International Conference on Computing and Data Science (pp. 143–149). https://doi.org/10.54254/2755-2721/64/20241365
Ye, T., Mu, J., & Hunter, J. (2026). Off-policy evaluation and conservative policy selection for slot-level dynamic bidding and ranking on the Open Bandit Dataset (Small). Journal of Technology Informatics and Engineering, 5(1), 178–199. https://doi.org/10.51903/jtie.v5i1.503
Zhang, B., Ren, Y., & Zou, J. (2025). LLM-style explainable e-commerce recommendation cards: A UI/UX design framework for trust-calibrated product recommendation. International Journal of Graphic Design, 3(2), 381–396. https://doi.org/10.51903/ijgd.v3i2.3697
Zhang, Y., & Zhang, H. (2025a). A therapist-facing session copilot for live counseling support: Reasoning-guided retrieval and ranking from multi-turn counseling dialogues. Journal of Technology Informatics and Engineering, 4(2), 464–486. https://doi.org/10.51903/jtie.v4i2.547
Zhang, Y., & Zhang, H. (2025b). Visualizing the right counseling support: Evidence-linked recommendation cards for explainable mental health intake interfaces. International Journal of Graphic Design, 3(1), 214–229. https://doi.org/10.51903/ijgd.v3i1.3722
Zheng, D., & Li, C. (2024). Behavior-level jailbreak resistance via multi-stage refusal + utility preservation. Journal of Advanced Computing Systems, 4(1), 83–99. https://doi.org/10.69987/JACS.2024.40107
Zheng, D., Li, C., & Davidson, H. (2023). Continual red-teaming for in-the-wild jailbreaks via online guardrail updates and guardrail distillation. Journal of Advanced Computing Systems, 3(2), 35–49. https://doi.org/10.69987/JACS.2023.30203
Zheng, D., Zhang, B., & Geibel, J. (2024). VerifySafe: Toxicity-safe agent responses under adversarial prompts with evidence-based self-verification. Journal of Advanced Computing Systems, 4(1), 67–82. https://doi.org/10.69987/JACS.2024.40106
Zhong, Z. S., Chen, J., Zhong, E., & Sun, X. (2025). Evidence-calibrated RAG for unanswerable question answering: Retrieval coverage, abstention calibration, and hallucination-proxy analysis on SQuAD 2.0. Journal of Technology Informatics and Engineering, 4(2), 502–520. https://doi.org/10.51903/jtie.v4i2.536
Zhong, Z. S., Li, C., & Rao, H. (2026). Trajectory reliability prediction for generalist AI agents: Tool-use failure analysis and success forecasting on ZClawBench. Journal of Technology Informatics and Engineering, 5(1), 341–360. https://doi.org/10.51903/jtie.v5i1.539
Zhou, B., Li, C., & Liu, L. (2025). Risk-calibrated patient-facing AI safety cards: A UI/UX design framework for rubric-based medical risk communication. International Journal of Graphic Design, 3(2), 365–380. https://doi.org/10.51903/ijgd.v3i2.3696
Zhou, B., Wang, H., & Chang, X. (2025). Distilling VMAF into an edge-deployable quality predictor: A pilot shot-level proxy with LLM-ready quality tokens. Journal of Technology Informatics and Engineering, 4(2), 447–463. https://doi.org/10.51903/jtie.v4i2.522
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Xiaochen Li, Fiona Wang

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.









5.png)
