NeoMME Cuts Visual Retrieval Index Storage While Keeping Benchmark Accuracy
NeoMME, released on Hugging Face, pairs 260M and 800M multilingual multimodal encoders with a retrieval design that cuts late-interaction index storage from about 1.5 MB to 6 kB per page.

H Company’s NeoMME release, described in its Hugging Face post, puts a compact multimodal encoder family on the visual-document retrieval benchmark frontier while attacking one of the biggest operating costs in late-interaction search: the amount of index storage needed for each page.
The H Company-authored Hugging Face post introduces NeoMME as 260-million- and 800-million-parameter multilingual multimodal encoders built without the usual separate vision tower or causal language model.
A single bidirectional Transformer processes text tokens and raw image patches, and the full model is trained from scratch with a masked discrete-diffusion objective.
Many recent retrievers adapt pretrained generative visual-language models, using a vision encoder, a projector and a causal decoder.
NeoMME instead starts from the retrieval, classification and token-labeling workload, where the model needs representations that can be compared efficiently rather than autoregressive text generation.
The retrieval version, NeoMME-Retriever, was fine-tuned using ColPali’s page-image approach.
It returns dense and late-interaction embeddings in a single forward pass, so the same output can support fast first-stage search and more detailed token-level matching.
That dual-head design tries to keep late-interaction accuracy while reducing the storage and serving penalty that can make visual RAG systems hard to deploy.
H Company reports that both NeoMME sizes sit on the ViDoRe v3 Pareto frontier for nDCG@10 and model size.
The team reports that at a matched 2048×2048 image input size on an NVIDIA L40S GPU, the 260-million-parameter model encodes about 51 pages per second, roughly twice ColModernVBERT’s throughput in the comparison used for the release.
The developers report that NeoMME-Retriever uses hierarchical pooling across tokens together with asymmetric quantization to shrink a late-interaction page index from about 1.5 MB to 6 kB.
That is a 255× reduction, while the model keeps more than 95% of the baseline nDCG@10 result cited for the uncompressed setup.
The ViDoRe v3 tests cover two compression configurations.
H Company reports that a pooling factor of 10 with int8 queries and documents brings storage to 39 kB per page, a 39-fold reduction that retains more than 99 percent of baseline nDCG@10.
The smaller 6 kB configuration uses a pooling factor of eight, int8 queries and binary document representations.
These are different settings, so the storage figures should be read with their respective quality-retention results.
The developers also describe a retrieval sequence for very large collections.
A dense embedding first retrieves a small candidate set through an approximate-nearest-neighbor index.
Late-interaction representations can then rerank those candidates.
The same forward pass supplies both representations, allowing the two stages to use different matching methods.
For teams indexing scanned reports, forms, research papers or enterprise document collections, that shift changes the trade-off between storing rich page-level visual representations and keeping infrastructure costs manageable.
Late-interaction methods can improve matching quality because they preserve more page detail, but the saved representations become expensive when every page carries a large index footprint.
NeoMME is available in Hugging Face Transformers, and all model checkpoints have been released under the Apache 2.0 license.
The post also points users toward fine-tuning with Sentence Transformers, making the model family usable for teams that already build dense retrieval, reranking or visual RAG pipelines around the Hugging Face ecosystem.
The reported results are bounded by H Company’s published evidence.
NeoMME-Retriever’s strongest claims come from the ViDoRe v3 setup, the L40S throughput test and storage compression measurements, not from an independent customer deployment.
The next material question is how the compact encoder behaves on organization-specific document sets, where page layouts, languages, scan quality and query patterns can differ sharply from public benchmarks.




















