Context-Enriched Figure Indexing for Scientific Multi-Modal RAG: An Industry Case Study at Scale

Abstract

Scientific figures encode key experimental evidence and domain knowledge, yet remain underutilized in retrieval-augmented generation (RAG). Directly indexing visual embeddings often misses fine-grained scientific semantics and misaligns with text-centric retrieval pipelines, while author-provided captions are frequently sparse and insufficiently contextualized. In this paper, we report on the design, implementation, and evaluation of a context-enriched figure indexing system built for a large-scale scientific digital library. Rather than introducing specialized multi-modal indexing infrastructure, our system converts figures into context-enriched captions—generated by fusing visual content with local article context via vision–language models—and indexes them alongside text in a unified embedding space using standard text retrievers. We evaluate this design on SciCap and a production corpus of approximately 113{,}000 figures drawn from Elsevier’s ScienceDirect, comparing open-source (MiniCPM) and proprietary (GPT-4o) captioning models. Enriched captions consistently outperform direct image embeddings and original author captions for figure retrieval and downstream question answering. Although GPT-4o produces higher-quality captions under intrinsic evaluation, MiniCPM achieves comparable retrieval utility at no API cost and runs on a single commodity GPU—a trade-off that favors open-source models for large-scale indexing. We share the practical lessons, cost–performance trade-offs, and remaining open challenges from building this system, supporting context-enriched captioning as a scalable and easily adoptable foundation for scientific multi-modal RAG in production.

Publication
The 2026 Annual International ACM SIGIR Conference on Research and Development in Information Retrieval in the Asia Pacific Region (SIGIR-AP 2026)
Yongkang Li
Yongkang Li
PhD Student

I am currently a PhD student in IR LAB, the University of Amsterdam, working with Prof. Evangelos Kanoulas. Before that, I got my master degree at Southern University of Science and Technology, Department of Computer Science and Engineering, SUSTech-UTokyo Joint Research Center on Super Smart City Lab, where I am supervised by Prof. Xuan Song in SUSTech and Prof. Zipei Fan at the University of Tokyo. What’s more, I received a B.E. degree in the School of Information and Communication Engineering, Beijing University of Posts and Telecommunications in 2020.