Multi-vector visual document retrievers such as ColPali have become the default way to search scanned PDFs, slide decks and forms, because they sidestep fragile OCR pipelines. A paper submitted to arXiv on October 7 by Zhuchenyang Liu, Yao Zhang and Yu Xiao of Aalto University argues that the infrastructure underneath them has a blind spot: the stored index itself leaks the documents it encodes.
The mechanics are the vulnerability. These systems keep roughly one thousand patch vectors per page, generated in raster order by a vision-language model that was pre-trained to read documents, and they typically live in vector databases run by third parties. Because no human can read a page from its vectors, the index is routinely treated as less sensitive than the page. The authors hypothesized that whoever runs or breaches the store can reproduce the page from the index alone.
They frame inversion as conditional document image generation. A flow-matching inverter learns to render a page from the stored vectors, inferring from the index itself what it needs: which encoder produced the vectors, the shape of the page, and, in the case of shuffled vectors, their original order.
On the ViDoRe v3 benchmark, pages reconstructed from raw indices recovered 47% of words and 45% of sensitive tokens, a category the paper defines as numbers, acronyms and capitalized terms. More damaging for access-control assumptions, when the reconstructions were used as queries against the same stored indices, they ranked their source page first 98.4% of the time.
The team also tested two cheap protections. Token pooling and vector shuffling both cut word recall to about 8%, which sounds like a fix. It is not: shuffling alone dropped source-page re-identification to 3.8%, but a learned position model that restores the order of a shuffled index pushed that figure back up to 93.5%. Inverting a pooled index, by contrast, remains an open problem, making pooling the only tested defense that held.
To check generalization, the researchers applied the same attack, unchanged, to a second multi-vector retriever. Inverted pages still ranked their source page first 70.2% of the time, although word recall there stayed below a nearest-neighbor baseline, suggesting the attack transfers more reliably as a re-identification tool than as a full-content recovery method.
The paper is explicit about what it did not evaluate, listing noise injection, quantization, keyed random projections and encrypted retrieval as untested defenses, and it directs its remediation advice at store operators rather than model vendors. The practical exposure it identifies is hosted vector databases with separated access tiers, where a tenant or an attacker with read access to the index gets far more than search results.
For teams running confidential documents through hosted visual RAG, the operational takeaway is blunt: treat the vector store as if it contained the source documents, because functionally it does. The paper has not yet been peer-reviewed, but the numbers are reproducible from the released method, and the burden of proof now sits with the platforms selling multi-vector retrieval as a privacy-preserving layer.
Comments (0)
Log in to join the discussion
Log InNo comments yet