Security Aalto Researchers Reconstruct 47% of Words From Stored ColPali Indexes Alone, and Re-Identify 98.4% of Source Pages AI Agents Hark Pro Goes Wide: Brett Adcock's $6 Billion AI Startup Opens a Free Assistant That Runs Your Computer News Nvidia Has Discussed Another $1 Billion Bet on Figure as the Humanoid Startup Raises at a $38 Billion Pre-Money Valuation Business TSMC Books $16.03 Billion for September, Up 54.6% Year Over Year, as Third-Quarter Revenue Hits a Record $46.7 Billion Coding Assistants Seven Clean-Room Adobe Clones in Pure Rust, Co-Written by Claude Opus 5.5, Hit GitHub - and Adobe Shares Wobbled News Microsoft Puts a Petaflop in Your Lap: Surface Laptop Ultra Ships Oct 16 From $2,599 With 128GB Unified Memory and Local 120B-Parameter Models AI Agents Docker Is Now Shipping AI Agents: docker-agent Turns a YAML File and an OCI Registry Into an Agent Distribution Channel Gemini Google DeepMind's EmbeddingGemma 2 Packs Text, Image, Video and Audio Into One 740M-Parameter Model That Runs in 191MB of RAM Security Aalto Researchers Reconstruct 47% of Words From Stored ColPali Indexes Alone, and Re-Identify 98.4% of Source Pages AI Agents Hark Pro Goes Wide: Brett Adcock's $6 Billion AI Startup Opens a Free Assistant That Runs Your Computer News Nvidia Has Discussed Another $1 Billion Bet on Figure as the Humanoid Startup Raises at a $38 Billion Pre-Money Valuation Business TSMC Books $16.03 Billion for September, Up 54.6% Year Over Year, as Third-Quarter Revenue Hits a Record $46.7 Billion Coding Assistants Seven Clean-Room Adobe Clones in Pure Rust, Co-Written by Claude Opus 5.5, Hit GitHub - and Adobe Shares Wobbled News Microsoft Puts a Petaflop in Your Lap: Surface Laptop Ultra Ships Oct 16 From $2,599 With 128GB Unified Memory and Local 120B-Parameter Models AI Agents Docker Is Now Shipping AI Agents: docker-agent Turns a YAML File and an OCI Registry Into an Agent Distribution Channel Gemini Google DeepMind's EmbeddingGemma 2 Packs Text, Image, Video and Audio Into One 740M-Parameter Model That Runs in 191MB of RAM

Aalto Researchers Reconstruct 47% of Words From Stored ColPali Indexes Alone, and Re-Identify 98.4% of Source Pages

Aalto Researchers Reconstruct 47% of Words From Stored ColPali Indexes Alone, and Re-Identify 98.4% of Source Pages

A paper submitted October 7 by Aalto University researchers shows the roughly 1,000 patch vectors a multi-vector visual document retriever stores per page can be inverted back into readable images: 47% of words and 45% of sensitive tokens recovered, and 98.4% of reconstructed pages ranked first when queried against the index.

Multi-vector visual document retrievers such as ColPali have become the default way to search scanned PDFs, slide decks and forms, because they sidestep fragile OCR pipelines. A paper submitted to arXiv on October 7 by Zhuchenyang Liu, Yao Zhang and Yu Xiao of Aalto University argues that the infrastructure underneath them has a blind spot: the stored index itself leaks the documents it encodes.

The mechanics are the vulnerability. These systems keep roughly one thousand patch vectors per page, generated in raster order by a vision-language model that was pre-trained to read documents, and they typically live in vector databases run by third parties. Because no human can read a page from its vectors, the index is routinely treated as less sensitive than the page. The authors hypothesized that whoever runs or breaches the store can reproduce the page from the index alone.

They frame inversion as conditional document image generation. A flow-matching inverter learns to render a page from the stored vectors, inferring from the index itself what it needs: which encoder produced the vectors, the shape of the page, and, in the case of shuffled vectors, their original order.

On the ViDoRe v3 benchmark, pages reconstructed from raw indices recovered 47% of words and 45% of sensitive tokens, a category the paper defines as numbers, acronyms and capitalized terms. More damaging for access-control assumptions, when the reconstructions were used as queries against the same stored indices, they ranked their source page first 98.4% of the time.

The team also tested two cheap protections. Token pooling and vector shuffling both cut word recall to about 8%, which sounds like a fix. It is not: shuffling alone dropped source-page re-identification to 3.8%, but a learned position model that restores the order of a shuffled index pushed that figure back up to 93.5%. Inverting a pooled index, by contrast, remains an open problem, making pooling the only tested defense that held.

To check generalization, the researchers applied the same attack, unchanged, to a second multi-vector retriever. Inverted pages still ranked their source page first 70.2% of the time, although word recall there stayed below a nearest-neighbor baseline, suggesting the attack transfers more reliably as a re-identification tool than as a full-content recovery method.

The paper is explicit about what it did not evaluate, listing noise injection, quantization, keyed random projections and encrypted retrieval as untested defenses, and it directs its remediation advice at store operators rather than model vendors. The practical exposure it identifies is hosted vector databases with separated access tiers, where a tenant or an attacker with read access to the index gets far more than search results.

For teams running confidential documents through hosted visual RAG, the operational takeaway is blunt: treat the vector store as if it contained the source documents, because functionally it does. The paper has not yet been peer-reviewed, but the numbers are reproducible from the released method, and the burden of proof now sits with the platforms selling multi-vector retrieval as a privacy-preserving layer.

Comments (0)

Log in to join the discussion

Log In

No comments yet