Tongyi Lab Introduces VimRAG, an Open-Source Framework for All-Modal Knowledge Bases


Connecting large language models (LLMs) to enterprise knowledge bases using Retrieval-Augmented Generation (RAG) has become a standard practice, enabling AI to provide verifiable answers rather than generating speculative content. However, traditional RAG solutions face limitations when knowledge bases evolve from purely text-based documents to include images and videos.
Consider a manufacturing enterprise's knowledge base, which might contain 100,000 PDF technical documents, 50,000 CAD design drawings and production line photos, and thousands of operational training videos, each 30 to 60 minutes long. Answering complex queries, such as identifying product design changes from the previous year's third quarter and locating discussions about these changes in meeting recordings, becomes challenging. This requires AI to process information across multiple modalities and understand implicit connections, like linking text in a PDF to annotations in CAD drawings or specific dialogue in a video.

Engineer viewing CAD drawing on tablet, with production line and video screen in background, illustrating multi-modal data challenge.
This scenario highlights the difficulties in implementing all-modal, long-context RAG. To address this, Tongyi Lab has open-sourced VimRAG, a unified RAG framework designed for hybrid knowledge bases that combine text, images, and video. VimRAG replaces linear data concatenation with dynamic memory graphs, enabling AI to identify key information, understand context, and perform cross-modal verification, thereby managing complex queries more effectively.
The VimRAG project is available on Arxiv, GitHub, and Hugging Face.
Challenges in Multi-Modal Long-Context Tasks
Current approaches to hybrid-modal retrieval-augmented generation often fall into two categories. The first is a text-centric method, where images are converted to text via optical character recognition (OCR) and videos are transcribed into subtitles. This process leads to a loss of critical information such as layout, color, and spatial relationships. For instance, describing "the red emergency button in the bottom left corner with a 3mm chamfered shadow" in text alone is cumbersome, and identifying an engineer's gesture while speaking in a video becomes impossible.
The second approach is a brute-force method, which involves establishing separate libraries for text, images, and videos. During retrieval, each library is searched independently, and the results are then combined. This can lead to issues where text refers to a diagram that the AI cannot locate, or a video instructs an action based on a section of text that the AI has forgotten.
A more fundamental issue is the confusion in cross-modal reasoning paths. Existing AI agents typically process context linearly. However, in a hybrid-modal scenario, a single retrieval might yield a paragraph of text, three images, and two video clips. As the number of steps increases, this linear context management can cause the model to lose track of which modalities it has examined, how different modalities corroborate each other, or whether to explore a video further or re-examine text, potentially leading to repetitive retrieval loops.

Abstract comparison of linear data path versus complex, interconnected dynamic memory graph.
VimRAG's Structured Memory Mechanism
VimRAG addresses these challenges by adopting a "structured memory" approach, inspired by human cognition. Just as humans recall a movie by remembering key plot points and visual highlights rather than replaying it frame by frame, VimRAG upgrades the AI agent's context from a linear history to a dynamic directed acyclic graph (DAG). This allows for multi-modal memory reconstruction before each action, preserving critical information while discarding ineffective searches.
The dynamic memory graph enables traceable retrieval and supports a trial-and-error mechanism. VimRAG moves beyond a linear "think-act-retrieve" sequence by building a DAG that originates from the user's query. Each retrieval generates a new node, encapsulating a text summary, visual evidence, and topological position. Redundant paths are automatically marked as dead ends, while effective paths are highlighted. This tree-like topology allows the AI to differentiate between exploratory searches and conclusive verifications, preventing repetitive queries.
VimRAG also employs a visual energy allocation strategy. Based on the memory graph's topological structure, the framework intelligently allocates visual memory quotas for each node. Core nodes and new evidence retain high-definition visual tokens, ensuring critical details are preserved, while peripheral nodes are downgraded to text descriptions or pruned. This dynamic strategy processes information efficiently, ensuring that effective information reaches the model with minimal token consumption, similar to how humans retain original copies of core documents but review summaries for secondary materials.

Abstract visualization of data nodes with varying levels of detail and brightness, representing visual energy allocation.
To make the memory paradigm trainable, VimRAG introduces Graph-Guided Policy Optimization (GGPO). This allows for fine-grained contribution assessment, where training rewards or penalizes specific actions within the graph rather than the entire trajectory. GGPO precisely backtracks based on the graph topology, pruning unhelpful dead ends in positive samples and protecting nodes where retrieval actions were effective but the answer was incorrect in negative samples. This mechanism reduces gradient variance, helping the model quickly internalize structured memory logic and improving training stability and efficiency.
Experimental Validation
To evaluate VimRAG, researchers created a rigorous test environment by combining all text, multi-modal documents, multi-element images, and long/short videos into a single, unified multi-modal corpus. This required the model to perform precise retrieval, memory, and generative understanding across all modalities.
In end-to-end evaluations using the Qwen3-VL-8B model, VimRAG achieved an average accuracy of 50.1%, outperforming various baselines. This demonstrated its effectiveness in addressing information sparsity in multi-modal, long-context scenarios.
Analysis of retrieval performance showed that VimRAG significantly outperformed Mem1 and ReAct baselines in general text, image, visual document, and video categories. By explicitly modeling the DAG of reasoning states, VimRAG avoided the state loss and repetitive queries common in traditional methods.
The entropy curve revealed that VimRAG with GGPO exhibited stable policy convergence after exploring a solvable distribution, indicating that GGPO's fine-grained credit assignment effectively reduces gradient variance during training.
Despite introducing perceptual memory actions, VimRAG improved overall inference efficiency by maintaining a structured reasoning topology, which reduced ineffective searches.
Case Demonstration
Consider a user query: "What is the complete solution process and mathematical proof for Lagrange multipliers in Chapter 4 of Dr. Smith's Calculus?"
Traditional RAG approaches might either convert the entire course video to text, losing formulas and blackboard notes, or retrieve from separate text, image, and video libraries, potentially overlooking important details.
VimRAG's process would involve:
The agent initially retrieves Chapter 3, identifies it as discussing "single-variable extrema," and prunes it as irrelevant to Lagrange multipliers.
Using topological localization, it directly targets Section 4.3 of Chapter 4, confirming it as the core chapter for "constrained optimization."
Within Section 4.3, it extracts the mathematical definition of the Lagrange formula (text), associates blackboard screenshots (images), and locates the complete derivation video for Example 4.3.2, "Maximizing Box Volume."
The final result synthesizes the formula, theorem, and example into a comprehensive answer by following the critical path: root node → node 2 → (node 3, node 4) → node 5.
VimRAG's approach demonstrates that moving beyond modality-specific processing and context blind spots, towards dynamic memory graphs and structured reasoning, allows large models to bridge the modality gap in complex, real-world knowledge bases. This framework offers a new, trainable, and iterative path for all-modal retrieval in business scenarios, transforming multi-modal knowledge into a precisely retrievable, deeply understood, and reliably generated business asset.

Alibaba Cloud Bailian Knowledge Base interface, showing multi-modal data integration.
Alibaba Cloud Bailian Knowledge Base is integrating VimRAG's core mechanisms to support multi-modal retrieval and generation for text, tables, images, audio, and video. This enables users to build exclusive RAG services for enterprise document Q&A, product image search, or audio/video content retrieval with minimal configuration.
Stay Ahead of the AI Curve
Join 50,000+ subscribers getting the latest AI tools, trends, and tutorials delivered to their inbox weekly.
No spam, unsubscribe at any time.