EverMind Releases Memory-Sparse Attention Code for Long-Context AI Models

Open Source Talent Scout
Abstract representation of AI memory and sparse attention within a neural network.

EverMind has open-sourced the code for its "Memory-Sparse Attention" (MSA) system, which enables AI models to process up to 100 million tokens—equivalent to about a thousand books—with minimal performance degradation. The release follows a paper published by EverMind several weeks ago detailing the technology.

The MSA system, utilizing a 4-billion-parameter model, reportedly surpassed Retrieval-Augmented Generation (RAG) systems built on models 58 times its size in multiple benchmarks. Instead of relying on external database searches, MSA integrates memory functionality directly into the model's architecture, allowing it to learn end-to-end which information to retain or discard without a separate retrieval process. The project's GitHub repository accumulated 2,500 stars within a day of its release, indicating significant industry interest.

Visual representation of data flow in a sparse attention AI architecture.

Visual representation of data flow in a sparse attention AI architecture.

Project Overview

Traditional large language models are often limited to context lengths of 128K to 1M tokens due to the computational demands of full attention mechanisms. Existing solutions, such as hybrid linear attention, fixed-size state memory, and external storage like RAG, face challenges including accuracy degradation, latency increases, lack of end-to-end differentiability, or complex engineering.

Abstract digital maze symbolizing limitations of traditional AI models.

Abstract digital maze symbolizing limitations of traditional AI models.

MSA addresses these limitations through an end-to-end trainable, scalable sparse implicit state memory framework. Key features include:

  • Scalable Sparse Attention and Document-level RoPE: This combination achieves near-linear complexity during both training and inference.

  • KV Cache Compression with Memory-Parallel Inference Engine: This enables a throughput of 100 million tokens on two A800 GPUs.

  • Memory Interleaving Mechanism: This supports multi-turn, multi-hop reasoning across dispersed memory fragments.

In long-context question answering and Needle In A Haystack (NIAH) benchmarks, MSA outperformed RAG systems using the same backbone, state-of-the-art RAG solutions, and leading long-context models. The system demonstrated less than 9% performance degradation across an unprecedented range of 16K to 100 million tokens, suggesting a method to decouple memory capacity from reasoning ability. MSA integrates top-k selection with sparse attention, allowing for document decoupling during inference while maintaining end-to-end differentiability. On the MS MARCO dataset, MSA's performance degradation was less than 9%.

Core Contributions

The development of MSA includes several core contributions:

  • Memory-Sparse Attention (MSA): An end-to-end trainable, scalable sparse attention layer, combined with document-level RoPE, achieves O(L) complexity with less than 9% performance degradation from 16K to 100 million tokens.

  • KV Cache Compression + Memory Parallelism: This involves hierarchical storage, distributed scoring, and on-demand transfer, enabling 100 million token inference on two A800 GPUs.

  • Memory Interleaving: This adaptively alternates between "generative retrieval," "context expansion," and "generation," enhancing multi-hop reasoning across documents.

  • Comprehensive Evaluation: MSA demonstrated superior stability and accuracy at scale in long-context question answering and NIAH benchmarks, outperforming RAG with the same backbone, state-of-the-art RAG solutions, and leading long-context models.

Close-up of A800 GPUs with holographic data overlay, illustrating memory parallelism.

Close-up of A800 GPUs with holographic data overlay, illustrating memory parallelism.

Architectural Design and Inference

MSA integrates retrieval and generation into a differentiable closed loop. Document implicit states (K/V/Kᵣ) are compressed via chunked mean pooling. A routing projector uses cosine similarity to calculate relevance, selecting Top-k documents. Their compressed K/V are then concatenated with the query's local K/V for autoregressive decoding. Routing is applied only to upper layers, while lower layers maintain independent document processing for hierarchical alignment.

The system uses Parallel (document-level) RoPE, where the position for each document resets from zero, preventing positional drift between short training and long inference, enabling extrapolation from 64K training to 100 million tokens. Global RoPE (active context) offsets the query's starting index by 'k' (number of Top-k retrieved blocks), maintaining causal order: Background → Query → Generation.

The MSA inference pipeline operates in three stages:

  1. Global Memory Encoding (offline): A forward pass on the corpus caches chunked and pooled (K̄, V̄, K̄ᵣ).

  2. Online Routing and Context Assembly: The query is projected to Qᵣ, matched with K̄ᵣ to select Top-k documents, and then only the selected K̄/V̄ are loaded and concatenated with the local context.

  3. Sparse Generation: Autoregressive generation occurs on the sparse context.

Memory parallelism shards K̄ᵣ across multiple GPUs, broadcasting the query, performing local scoring, and conducting global reduction. Content K̄/V̄ are stored in host memory and pulled asynchronously when selected, balancing VRAM and throughput for 100 million token deployment.

Three-stage diagram of the MSA inference pipeline: Encoding, Routing, Generation.

Three-stage diagram of the MSA inference pipeline: Encoding, Routing, Generation.

Experimental Results

Experiments were conducted on nine question-answering datasets and eight NIAH subtasks. The backbone model used was Qwen3-4B-Instruct-2507, compared against RAG with the same backbone and state-of-the-art RAG solutions.

In comparisons against RAG with the same Qwen3-4B backbone, MSA achieved an average score of 3.760, representing a 16.0% improvement over standard RAG, an 11.5% improvement over RAG with reranking, and a 14.8% improvement over HippoRAG2. MSA led on all datasets except NarrativeQA within this group.

When compared to state-of-the-art RAG solutions using larger backbone models (KaLMv2 + Qwen3-235B and KaLMv2 + Llama-3.3-70B), MSA achieved the highest score on four out of nine datasets. Its average score of 3.760 represented improvements of 7.2%, 5.0%, 10.7%, and 5.4% over the strongest configurations of these larger models. Discrepancies on some datasets, such as MuSiQue, were attributed to differences in parameter count and inherent reasoning capabilities.

For NIAH stability (32K to 1 million tokens), MSA maintained 94.84% accuracy at 1 million tokens. The unmodified backbone model showed a sharp decline after 128K, reaching only 24.69% at 1 million tokens. Hybrid linear attention models also degraded significantly at or above 128K/256K. While external memory agents like RL-MemoryAgent-14B were relatively stable, their absolute accuracy and decay rate were lower than MSA's.

Ablation studies indicated that curriculum extension, memory interleaving, continuous pre-training, and injecting raw text all contributed significantly to MSA's performance; their removal led to performance degradations ranging from 5% to 37%.

Abstract visual showing parts of a digital structure dimmed or removed, symbolizing ablation.

Abstract visual showing parts of a digital structure dimmed or removed, symbolizing ablation.

Implementation Details

The training involved 158.95 billion tokens of continuous pre-training, utilizing auxiliary routing loss, followed by a two-stage Supervised Fine-Tuning (SFT) process with curriculum learning from 8K to 64K.

The project's code and instructions are available on GitHub.

ToolMesh
ToolMesh Weekly

Stay Ahead of the AI Curve

Join 50,000+ subscribers getting the latest AI tools, trends, and tutorials delivered to their inbox weekly.

No spam, unsubscribe at any time.