What happened?
Saleh Almohaimeed and four colleagues introduced a framework called Sensitive Entity Alias Generator (SEAG) that addresses the privacy problem in retrieval-augmented generation (RAG) systems. The study was published on arXiv on August 13, 2026. SEAG uses a lightweight model to detect sensitive entities (such as names, organizations, or personal data) and generates corresponding aliases, building a replacement table.
This table is used to replace sensitive words in the user query and retrieved documents before the data is sent to a third-party external generator model. This way, the external model can still produce a meaningful and accurate response without accessing the actual sensitive information.
Why does it matter?
Privacy research in RAG systems using large language models (LLM) has so far mostly focused on preventing unauthorized users from accessing sensitive data. However, according to the researchers, an overlooked issue was that external generator models could directly access the query and retrieved documents, creating a risk of misuse or unintended access to hidden information. SEAG allows users to benefit from powerful third-party generator models without having to share their sensitive information.
Findings
- On the user metric, meaning the model's ability to answer users accurately while hiding sensitive information from the external generator, all SEAG models achieved over 80% accuracy.
- The Qwen-3-based SEAG model achieved 77.83% overall accuracy in hiding all sensitive entities in documents.
- The LLaMA-3.2-based model achieved 76.73%, and the Phi-4-based model achieved 74.91% accuracy.
- The researchers created one dataset to train the SEAG models and a separate dataset to evaluate the framework.
What's next?
The study has been submitted to Knowledge-Based Systems and is expected to undergo peer review. On the arXiv page, PDF, HTML, and TeX source files are publicly available; details about code and data sharing have not yet been clarified.
What 77 percent concealment means
The numbers deserve a closer look. Even in the best-performing model, entity concealment succeeds 77.83 percent of the time — meaning roughly one in four sensitive entities reaches the external model unchanged. For an accuracy metric, 78 percent may be a respectable figure; for a privacy metric, an average means something different. A name, an organisation or a patient record is either concealed or it is not. A 22 percent leak does not read as “partly protected” to the organisation using the system; it reads as “I do not know what leaked”.
The paper also offers no baseline. It does not report how many sensitive entities pass through with SEAG absent, or what a simpler masking approach would achieve. Without that, there is no way to isolate where the framework's contribution comes from.
Relocating the risk
The approach itself raises a question too. The alias table is a file mapping real entities one-to-one onto their substitutes — which makes that table the most sensitive part of the system. Information protected from the external model becomes a new asset held inside your own infrastructure. That is not a flaw but an architectural trade-off; with no limitations section, though, it goes undiscussed.
One more point matters in practice: alias substitution conceals only the name of the sensitive information. Even when a person's name in a document is swapped for an alias, the dates, locations and roles in that same document can be enough to identify them. This is a known problem in the anonymisation literature and it applies to any method operating at entity level.