Scientific figures encode key experimental evidence and domain knowledge, yet remain underutilized in retrieval-augmented generation (RAG). Directly indexing visual embeddings often misses fine-grained scientific semantics and misaligns with text-centric retrieval pipelines, while author-provided captions are frequently sparse and insufficiently contextualized. In this paper, we report on the design, implementation, and evaluation of a context-enriched figure indexing system built for a large-scale scientific digital library. Rather than introducing specialized multi-modal indexing infrastructure, our system converts figures into context-enriched captions—generated by fusing visual content with local article context via vision–language models—and indexes them alongside text in a unified embedding space using standard text retrievers. We evaluate this design on SciCap and a production corpus of approximately 113{,}000 figures drawn from Elsevier’s ScienceDirect, comparing open-source (MiniCPM) and proprietary (GPT-4o) captioning models. Enriched captions consistently outperform direct image embeddings and original author captions for figure retrieval and downstream question answering. Although GPT-4o produces higher-quality captions under intrinsic evaluation, MiniCPM achieves comparable retrieval utility at no API cost and runs on a single commodity GPU—a trade-off that favors open-source models for large-scale indexing. We share the practical lessons, cost–performance trade-offs, and remaining open challenges from building this system, supporting context-enriched captioning as a scalable and easily adoptable foundation for scientific multi-modal RAG in production.