Grounded AI Pointer Cards for Browser Agents: Evidence-Aware Web Element Reranking and Action Explanation on WebLINX

Authors

  • Yinchen Shi Computer Science, New York University, NY, USA
  • Yuxuan Ren Chemical Engineering & Data Science, University of Washington, WA, USA

DOI:

https://doi.org/10.51903/ijgd.v4i1.4002

Keywords:

browser agents, web element grounding, pointer cards, learning to rank, human–AI interaction

Abstract

AI browser agents increasingly promise to turn the mouse pointer into a situated interface for selecting page content, asking follow-up questions, scrolling to relevant sections, and opening contextual panels. For graphic design research, this shift matters because the pointer card is both a computational output and a visual communication device: it must make the target, supporting evidence, and next action legible at the moment of use. This paper evaluates evidence-aware web element reranking as the grounding layer for AI pointer cards. The empirical study uses the WebLINX candidate-reranking data, including validation, same-distribution test, and unseen-website test partitions. Each query is paired with candidate DOM/page elements and binary relevance labels. Nine reranking methods were evaluated, ranging from random and DOM-salience controls to BM25 variants, a transparent hybrid formula, a logistic reranker, and a depth-limited decision-tree reranker. Across 5,883 queries and 1,556,655 query–element pairs, the tree reranker achieved the highest same-distribution nDCG@10 (0.272) and Recall@10 (0.473), while the fixed hybrid achieved the highest unseen-website Top-1 accuracy (0.104). Combining lexical and DOM evidence improved ranking, whereas layout features transferred less consistently across websites. Because the first-ranked target was correct in only 10.4% of unseen-website queries, the findings support a confirm-before-act grounding prototype with visible alternatives and correction controls rather than autonomous execution. The reported metrics evaluate target ranking rather than card usability.

References

Amershi, S., Weld, D., Vorvoreanu, M., Fourney, A., Nushi, B., Collisson, P., Suh, J., Iqbal, S., Bennett, P. N., Inkpen, K., Teevan, J., Kikin-Gil, R., & Horvitz, E. (2019). Guidelines for human-AI interaction. Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems, 1–13. https://doi.org/10.1145/3290605.3300233

Baranes, A., & Marchant, R. (2026, May 12). Reimagining the mouse pointer for the AI era. Google DeepMind. https://deepmind.google/blog/ai-pointer/

Bertin, J. (1983). Semiology of graphics: Diagrams, networks, maps. University of Wisconsin Press.

Card, S. K., Mackinlay, J. D., & Shneiderman, B. (Eds.). (1999). Readings in information visualization: Using vision to think. Morgan Kaufmann.

Chang, X., Lu, Y., & Zhong, Z. S. (2026). Review-grounded explainable recommendation with faithfulness evaluation on Amazon reviews. JEECS (Journal of Electrical Engineering and Computer Sciences), 11(1), 9–22. https://doi.org/10.54732/jeecs.v11i1.2

Chang, X., Ye, T., & Luo, S. (2023). LLM-as-reranker for personalized recommendation: Popularity bias mitigation and faithful natural-language explanations on MovieLens 100K. Journal of Advanced Computing Systems, 3(8), 61–78. https://doi.org/10.69987/JACS.2023.30806

Chen, Y., & Li, M. (2025). From hand-drawn sketches to interactive web prototypes: A reproducible vision-language approach with structural and visual consistency evaluation. Journal of Technology Informatics and Engineering, 4(2), 364–384. https://doi.org/10.51903/jtie.v4i2.490

Chen, Y., & Xu, H. (2026). Trust-calibrated multilingual RAG for humanitarian information platforms: Empirical evaluation on OMoS-QA for migration information access. International Journal of Graphic Design, 4(1), 141–164. https://doi.org/10.51903/ijgd.v4i1.3552

Deng, X., Gu, Y., Zheng, B., Chen, S., Stevens, S., Wang, B., Sun, H., & Su, Y. (2023). Mind2Web: Towards a generalist agent for the web. Advances in Neural Information Processing Systems, 36, 28091–28114.

Doshi-Velez, F., & Kim, B. (2017). Towards a rigorous science of interpretable machine learning [Preprint]. arXiv. https://arxiv.org/abs/1702.08608

Google Labs. (2026). Disco. Retrieved July 23, 2026, from https://labs.google/disco

Heer, J., & Shneiderman, B. (2012). Interactive dynamics for visual analysis. Communications of the ACM, 55(4), 45–54. https://doi.org/10.1145/2133806.2133821

Järvelin, K., & Kekäläinen, J. (2002). Cumulated gain-based evaluation of IR techniques. ACM Transactions on Information Systems, 20(4), 422–446. https://doi.org/10.1145/582415.582418

Jin, J. (2025a). Evidence-chain reliable RAG: Hallucination detection, source attribution, and deterministic provenance explanations. Journal of Technology Informatics and Engineering, 4(2), 520–533. https://doi.org/10.51903/jtie.v4i2.535

Jin, J. (2025b). LLM-style evidence cards for scientific search interfaces: A UI/UX design framework for retrieval transparency, ranking trust, and visual evidence hierarchy. International Journal of Graphic Design, 3(2), 397–414. https://doi.org/10.51903/ijgd.v3i2.3698

Koh, J. Y., Lo, R., Jang, L., Duvvur, V., Lim, M. C., Huang, P.-Y., Neubig, G., Zhou, S., Salakhutdinov, R., & Fried, D. (2024). VisualWebArena: Evaluating multimodal agents on realistic visual web tasks. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 881–905). Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.acl-long.50

Li, C., Bai, J., & Wang, S. (2024). Evidence-chain reliable RAG: Word-level hallucination detection, source attribution, and provenance explanation for LLM applications. Journal of Advanced Computing Systems, 4(2), 76–92. https://doi.org/10.69987/JACS.2024.40207

Li, C., Liu, G., & Zhao, Z. (2026). Cost-aware LLM-style routing for AIOps log analysis: Log parsing, anomaly detection, fault diagnosis, and incident summarization on LogEval task files. Journal of Technology Informatics and Engineering, 5(2), 91–103. https://doi.org/10.51903/jtie.v5i2.538

Li, C., Su, W., & Zhang, E. (2023). Lightweight hallucination firewall for enterprise LLM applications: Evidence consistency, self-checking, and small-model detection on TruthfulQA. Journal of Advanced Computing Systems, 3(1), 49–65. https://doi.org/10.69987/JACS.2023.30104

Li, C., Zhou, B., & Gao, K. (2025). Risk-calibrated patient-facing AI safety cards: A UI/UX benchmark for explainable medical AI response interfaces. International Journal of Graphic Design, 3(2), 381–394. https://doi.org/10.51903/ijgd.v3i2.3709

Li, Y., Lu, S., & Zhao, L. (2025). LLM-as-design-critic: Aligning AI-generated UI feedback with human graphic design judgment. International Journal of Graphic Design, 3(1), 196–215. https://doi.org/10.51903/ijgd.v3i1.3661

Li, Z., Zhou, S., & Zhou, Z. (2025). Financial risk dashboard design for institutional RWA investors: Visual hierarchy, chart comprehension, and explainability in FinChart-Bench. International Journal of Graphic Design, 3(1), 196–210. https://doi.org/10.51903/ijgd.v3i1.3715

Liu, G., He, S., & Wong, H. (2025). LLM-compatible visual brief cards for AI infrastructure capacity dashboards: A UI/UX framework for turning forecast risk into graphic design decisions. International Journal of Graphic Design, 3(1), 196–213. https://doi.org/10.51903/ijgd.v3i1.3723

Liu, G., Li, C., & Zhang, E. (2024). OpsLLM for cloud incident triage: Bilingual RAG-based root cause analysis and alert summarization for AI infrastructure operations. Journal of Advanced Computing Systems, 4(4), 97–111. https://doi.org/10.69987/JACS.2024.40408

Liu, T.-Y. (2009). Learning to rank for information retrieval. Foundations and Trends in Information Retrieval, 3(3), 225–331. https://doi.org/10.1561/1500000016

Lù, X. H., Kasner, Z., & Reddy, S. (2024). WebLINX: Real-world website navigation with multi-turn dialogue. In Proceedings of the 41st International Conference on Machine Learning (Vol. 235, pp. 33007–33056). PMLR. https://proceedings.mlr.press/v235/lu24e.html

Meng, S., Chen, J., & Zheng, I. (2026). LLM-inspired offline reranking for financial search: Query rewriting, hybrid retrieval, and listwise relevance ranking on FiQA. Journal of Technology Informatics and Engineering, 5(1), 361–378. https://doi.org/10.51903/jtie.v5i1.537

Mu, J., Lu, Y., & Hwang, E. (2026). Structured visual brief interfaces for advertising design: A UI/UX framework for turning creative intentions into designer-editable graphic design cards. International Journal of Graphic Design, 4(1), 192–208. https://doi.org/10.51903/ijgd.v4i1.3702

Mu, J., Ye, T., & Patel, P. (2025). Offline counterfactual evaluation for advertising and recommendation slot policies: A reproducible study on the Open Bandit Dataset (Small). Journal of Technology Informatics and Engineering, 4(3), 521–543. https://doi.org/10.51903/jtie.v4i3.500

Muennighoff, N., Tazi, N., Magne, L., & Reimers, N. (2023). MTEB: Massive text embedding benchmark. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, 2014–2037. https://doi.org/10.18653/v1/2023.eacl-main.148

Nakano, R., Hilton, J., Balaji, S., Wu, J., Ouyang, L., Kim, C., Hesse, C., Jain, S., Kosaraju, V., Saunders, W., Jiang, X., Cobbe, K., Eloundou, T., Krueger, G., Button, K., Knight, M., Chess, B., & Schulman, J. (2021). WebGPT: Browser-assisted question-answering with human feedback [Preprint]. arXiv. https://arxiv.org/abs/2112.09332

Nie, J., Liu, G., Li, C., & Zou, T. (2026). Evidence-constrained incident visualization cards for distributed cloud logs: A UI/UX framework for turning Hadoop, OpenStack, and ZooKeeper logs into actionable SRE design interfaces. International Journal of Graphic Design, 4(1), 179–185. https://doi.org/10.51903/ijgd.v4i1.3703

Nie, J., & Zheng, D. (2023). Ambiguity-aware HDFS log anomaly detection with retrieval-augmented failure narratives and selective refusal. Journal of Advanced Computing Systems, 3(1), 66–80. https://doi.org/10.69987/JACS.2023.30105

Nielsen, J. (1994). Usability engineering. Morgan Kaufmann.

Norman, D. A. (2013). The design of everyday things. Basic Books.

Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., Vanderplas, J., Passos, A., Cournapeau, D., Brucher, M., Perrot, M., & Duchesnay, É. (2011). Scikit-learn: Machine learning in Python. Journal of Machine Learning Research, 12, 2825–2830.

Ribeiro, M. T., Singh, S., & Guestrin, C. (2016). Why should I trust you? Explaining the predictions of any classifier. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 1135–1144. https://doi.org/10.1145/2939672.2939778

Robertson, S., & Zaragoza, H. (2009). The probabilistic relevance framework: BM25 and beyond. Foundations and Trends in Information Retrieval, 3(4), 333–389. https://doi.org/10.1561/1500000019

Shneiderman, B. (2022). Human-centered AI. Oxford University Press. https://doi.org/10.1093/oso/9780192845290.001.0001

Su, W., Chen, S., & Qian, E. (2026). Narrative-aware scientific claim verification agent with evidence ranking for ClimateCheck. Journal of Technology Informatics and Engineering, 5(1), 327–340. https://doi.org/10.51903/jtie.v5i1.549

Su, W., Chen, S., & Zhao, C. (2025). Budgeted multi-hop retrieval agent for compositional question answering: A retrieval-policy evaluation on the official MultiHop-RAG benchmark. Journal of Technology Informatics and Engineering, 4(3), 649–662. https://doi.org/10.51903/jtie.v4i3.543

Su, W., Rao, H., & Ma, E. (2026). Privacy and data-integrity risk cards for LLM agents: A UI/UX design framework for secure human oversight under prompt-injection attacks. International Journal of Graphic Design, 4(1), 186–191. https://doi.org/10.51903/ijgd.v4i1.3699

Sultana, M., Sumaiya, U. K., Saha, R. R., & Tudu, R. (2025). Critical Speculations In AI Design: Ethical Narratives And Visual Experiments In Post-Human Aesthetic Practices. International Journal of Graphic Design, 3(1), 81–101. https://doi.org/10.51903/IJGD.V3I1.2873

Thakur, N., Reimers, N., Rücklé, A., Srivastava, A., & Gurevych, I. (2021). BEIR: A heterogeneous benchmark for zero-shot evaluation of information retrieval models. Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks, 1.

Tu, H., Zhao, S., & Zhou, A. (2025). Visual brief cards for advertising design: A structured UI/UX framework for turning creative intentions into graphic design decisions. International Journal of Graphic Design, 3(1), 210–226. https://doi.org/10.51903/ijgd.v3i1.3714

Tufte, E. R. (2001). The visual display of quantitative information (2nd ed.). Graphics Press.

Wang, H., Ren, Y., & Chang, X. (2025). Layout-aware progressive PDF rendering: AI prioritization of PDF slices to reduce time-to-functional-first-frame on FUNSD. Journal of Technology Informatics and Engineering, 4(2), 425–446. https://doi.org/10.51903/jtie.v4i2.523

Wang, Y., Wang, L., Li, Y., He, D., Chen, W., & Liu, T.-Y. (2013). A theoretical analysis of NDCG ranking measures. Proceedings of Machine Learning Research, 30, 25–54. https://proceedings.mlr.press/v30/Wang13.html

Ware, C. (2013). Information visualization: Perception for design (3rd ed.). Morgan Kaufmann.

Wu, Q., Bai, J., & Zhou, X. (2023). Evidence-grounded financial RAG: Reducing numerical hallucination in LLM-generated corporate risk memos. Journal of Advanced Computing Systems, 3(3), 65–84. https://doi.org/10.69987/JACS.2023.30306

Wu, Q., Meng, S., & Zhao, J. (2025). Text-grounded LLM-assisted design rationale interfaces: Turning advertising layout metadata into explainable UI/UX decision cards. International Journal of Graphic Design, 3(1), 216–240. https://doi.org/10.51903/ijgd.v3i1.3713

Xin, Q. (2025a). Explaining OpenStack failure-injection log anomalies with retrieved normal prototypes. Emerging Information Science and Technology, 6(2), 125–146. https://doi.org/10.18196/eist.v6i2.31232

Xin, Q. (2025b). Uncertainty-aware late fusion for 3D perception (confidence calibration + fusion rule learning). Journal of Technology Informatics and Engineering, 4(1), 215–238. https://doi.org/10.51903/jtie.v4i1.485

Xin, Q. (2026). Log anomaly detection with conformal alert control and evidence-grounded incident ticket generation. AVITEC, 8(2), 247–264. https://doi.org/10.28989/avitec.v8i2.3974

Xu, H., Chen, Y., & Med, A. (2025). Automatic detection and explanation of dark patterns from interface microcopy: Empirical comparison of BERT-style encoders, RoBERTa-style encoders, and LLM-style decoders on the ec-darkpattern dataset. Journal of Technology Informatics and Engineering, 4(3), 590–612. https://doi.org/10.51903/jtie.v4i3.491

Ye, T., Mu, J., & Hunter, J. (2026). Off-policy evaluation and conservative policy selection for slot-level dynamic bidding and ranking on the Open Bandit Dataset (Small). Journal of Technology Informatics and Engineering, 5(1), 178–199. https://doi.org/10.51903/jtie.v5i1.503

Zhang, B., Rao, H., & Zhao, D. (2024). Evidence-grounded RAG for cloud-native DevOps: Hallucination-resistant AIOps question answering over private operations documents. Journal of Advanced Computing Systems, 4(3), 109–125. https://doi.org/10.69987/JACS.2024.40308

Zhang, B., Ren, Y., & Zou, J. (2025). LLM-style explainable e-commerce recommendation cards: A UI/UX design framework for trust-calibrated product recommendation. International Journal of Graphic Design, 3(2), 381–396. https://doi.org/10.51903/ijgd.v3i2.3697

Zhang, K., Meng, S., & Zhou, E. (2023). Evidence-grounded trading desk risk memos over SEC filings: Retrieval-augmented generation with XBRL numeric verification. Journal of Advanced Computing Systems, 3(2), 60–76. https://doi.org/10.69987/JACS.2023.30205

Zhang, R., Wen, Z., Wang, C., Tang, C., Xu, P., & Jiang, Y. (2025). Quality analysis and evaluation prediction of RAG retrieval based on machine learning algorithms [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2511.19481

Zhang, Y., & Zhang, H. (2025a). A therapist-facing session copilot for live counseling support: Reasoning-guided retrieval and ranking from multi-turn counseling dialogues. Journal of Technology Informatics and Engineering, 4(2), 464–486. https://doi.org/10.51903/jtie.v4i2.547

Zhang, Y., & Zhang, H. (2025b). Visualizing the right counseling support: Evidence-linked recommendation cards for explainable mental health intake interfaces. International Journal of Graphic Design, 3(1), 214–229. https://doi.org/10.51903/ijgd.v3i1.3722

Zheng, B., Gou, B., Kil, J., Sun, H., & Su, Y. (2024). GPT-4V(ision) is a generalist web agent, if grounded. In Proceedings of the 41st International Conference on Machine Learning (Vol. 235, pp. 61349–61385). PMLR. https://proceedings.mlr.press/v235/zheng24e.html

Zheng, D., & Li, C. (2024). Behavior-level jailbreak resistance via multi-stage refusal + utility preservation. Journal of Advanced Computing Systems, 4(1), 83–99. https://doi.org/10.69987/JACS.2024.40107

Zheng, D., Zhang, B., & Geibel, J. (2024). VerifySafe: Toxicity-safe agent responses under adversarial prompts with evidence-based self-verification. Journal of Advanced Computing Systems, 4(1), 67–82. https://doi.org/10.69987/JACS.2024.40106

Zhong, Z. S., Chen, J., Zhong, E., & Sun, X. (2025). Evidence-calibrated RAG for unanswerable question answering: Retrieval coverage, abstention calibration, and hallucination-proxy analysis on SQuAD 2.0. Journal of Technology Informatics and Engineering, 4(2), 502–520. https://doi.org/10.51903/jtie.v4i2.536

Zhong, Z. S., Wu, Q., & Mi, G. (2025). Uncertainty-aware medical image explanation cards: LLM-generated visual explanations for AI-assisted radiology interfaces. International Journal of Graphic Design, 3(2), 415–436. https://doi.org/10.51903/ijgd.v3i2.3616

Zhou, B., Jin, J., & Zhao, D. (2025). Calibrated resume-job matching for trustworthy LLM-assisted recruiter screening: Pairwise matching, probability calibration, and selective refusal on two public recruitment datasets. Journal of Technology Informatics and Engineering, 4(3), 625–648. https://doi.org/10.51903/jtie.v4i3.529

Zhou, B., Wang, H., & Chang, X. (2025). Distilling VMAF into an edge-deployable quality predictor: A pilot shot-level proxy with LLM-ready quality tokens. Journal of Technology Informatics and Engineering, 4(2), 447–463. https://doi.org/10.51903/jtie.v4i2.522

Zhou, S., Xu, F. F., Zhu, H., Zhou, X., Lo, R., Sridhar, A., Cheng, X., Bisk, Y., Fried, D., Alon, U., & Neubig, G. (2024). WebArena: A realistic web environment for building autonomous agents. International Conference on Learning Representations.

Downloads

Published

2026-04-30

How to Cite

Grounded AI Pointer Cards for Browser Agents: Evidence-Aware Web Element Reranking and Action Explanation on WebLINX. (2026). International Journal of Graphic Design, 4(1), 224-245. https://doi.org/10.51903/ijgd.v4i1.4002