SendTech Times
News
SYSTEMS SHIFT:

NeoMME Cuts Visual Retrieval Index Storage While Keeping Benchmark Accuracy

Newsroom brief

NeoMME, released on Hugging Face, pairs 260M and 800M multilingual multimodal encoders with a retrieval design that cuts late-interaction index storage from about 1.5 MB to 6 kB per page.

Verified against source materialEdited by SendTech Times AI & Enterprise DeskSource: Hugging Face Blog
NeoMME Cuts Visual Retrieval Index Storage While Keeping Benchmark Accuracy
Image source: Hugging Face

H Company’s NeoMME release, described in its Hugging Face post, puts a compact multimodal encoder family on the visual-document retrieval benchmark frontier while attacking one of the biggest operating costs in late-interaction search: the amount of index storage needed for each page.

The H Company-authored Hugging Face post introduces NeoMME as 260-million- and 800-million-parameter multilingual multimodal encoders built without the usual separate vision tower or causal language model.

A single bidirectional Transformer processes text tokens and raw image patches, and the full model is trained from scratch with a masked discrete-diffusion objective.

Many recent retrievers adapt pretrained generative visual-language models, using a vision encoder, a projector and a causal decoder.

NeoMME instead starts from the retrieval, classification and token-labeling workload, where the model needs representations that can be compared efficiently rather than autoregressive text generation.

The retrieval version, NeoMME-Retriever, was fine-tuned using ColPali’s page-image approach.

It returns dense and late-interaction embeddings in a single forward pass, so the same output can support fast first-stage search and more detailed token-level matching.

That dual-head design tries to keep late-interaction accuracy while reducing the storage and serving penalty that can make visual RAG systems hard to deploy.

H Company reports that both NeoMME sizes sit on the ViDoRe v3 Pareto frontier for nDCG@10 and model size.

The team reports that at a matched 2048×2048 image input size on an NVIDIA L40S GPU, the 260-million-parameter model encodes about 51 pages per second, roughly twice ColModernVBERT’s throughput in the comparison used for the release.

The developers report that NeoMME-Retriever uses hierarchical pooling across tokens together with asymmetric quantization to shrink a late-interaction page index from about 1.5 MB to 6 kB.

That is a 255× reduction, while the model keeps more than 95% of the baseline nDCG@10 result cited for the uncompressed setup.

The ViDoRe v3 tests cover two compression configurations.

H Company reports that a pooling factor of 10 with int8 queries and documents brings storage to 39 kB per page, a 39-fold reduction that retains more than 99 percent of baseline nDCG@10.

The smaller 6 kB configuration uses a pooling factor of eight, int8 queries and binary document representations.

These are different settings, so the storage figures should be read with their respective quality-retention results.

The developers also describe a retrieval sequence for very large collections.

A dense embedding first retrieves a small candidate set through an approximate-nearest-neighbor index.

Late-interaction representations can then rerank those candidates.

The same forward pass supplies both representations, allowing the two stages to use different matching methods.

For teams indexing scanned reports, forms, research papers or enterprise document collections, that shift changes the trade-off between storing rich page-level visual representations and keeping infrastructure costs manageable.

Late-interaction methods can improve matching quality because they preserve more page detail, but the saved representations become expensive when every page carries a large index footprint.

NeoMME is available in Hugging Face Transformers, and all model checkpoints have been released under the Apache 2.0 license.

The post also points users toward fine-tuning with Sentence Transformers, making the model family usable for teams that already build dense retrieval, reranking or visual RAG pipelines around the Hugging Face ecosystem.

The reported results are bounded by H Company’s published evidence.

NeoMME-Retriever’s strongest claims come from the ViDoRe v3 setup, the L40S throughput test and storage compression measurements, not from an independent customer deployment.

The next material question is how the compact encoder behaves on organization-specific document sets, where page layouts, languages, scan quality and query patterns can differ sharply from public benchmarks.

Share this article
inXf

Related articles

More
Hugging Face Adds 4-Bit Nunchaku Loading To Diffusers
Chips & Semiconductors

Hugging Face Adds 4-Bit Nunchaku Loading To Diffusers

Hugging Face added Nunchaku Lite support to Diffusers, letting developers load 4-bit diffusion checkpoints with `from_pretrained()` while using Hub-delivered CUDA kernels.

Supermicro 160-Bay NVMe Server Reaches 20PB In 4U
Chips & Semiconductors

Supermicro 160-Bay NVMe Server Reaches 20PB In 4U

ServeTheHome examined Supermicro’s ASG-4116S-NU160R at FMS 2026, a single-socket AMD EPYC storage server with 160 U.2 NVMe bays, roughly 19PB to 20PB of capacity in 4U, and redundant 2.6kW power supplies.

Kioxia GP1 SSD Demo Pushes PCIe Gen6 Storage To 10 Million IOPS
Chips & Semiconductors

Kioxia GP1 SSD Demo Pushes PCIe Gen6 Storage To 10 Million IOPS

ServeTheHome showed Kioxia’s GP1 SSD running just above 10 million IOPS at FMS 2026, using PCIe Gen6 and second-generation XL-FLASH for a fast server storage tier.

Samsung $15 Billion Power Prepayment Proposal Ties Fabs to Grid Funding
Chips & Semiconductors

Samsung $15 Billion Power Prepayment Proposal Ties Fabs to Grid Funding

Reuters reports that KEPCO asked Samsung to consider a roughly $15 billion electricity prepayment as South Korea weighs grid funding for chip plants and AI data centres.

AMD Prices 256-Core EPYC 9996 At $14,904 For Server Buyers
Chips & Semiconductors

AMD Prices 256-Core EPYC 9996 At $14,904 For Server Buyers

TechRadar reports that AMD’s 6th Gen EPYC 9006 “Venice” lineup includes a 256-core EPYC 9996 with 512 threads, 1GB of L3 cache, a 600W default power rating and a $14,904 list price for 1,000-unit orders.

AI Compute Shortage Shifts Chip Strategy Towards Efficiency
Capital & Policy

AI Compute Shortage Shifts Chip Strategy Towards Efficiency

Semiconductor Engineering analysed how AI data-centre growth is being shaped by foundry, memory, power and optical-link bottlenecks as chip buyers look for more tokens from scarce capacity.

Lenovo Whitsett AI Server Line Doubles Plant To 890,000 Square Feet
Chips & Semiconductors

Lenovo Whitsett AI Server Line Doubles Plant To 890,000 Square Feet

A Tom's Hardware tour details Lenovo's expanded Whitsett AI server line, including 890,000 square feet, up to 20 MW of power and liquid-cooling tests for GB300 and Vera Rubin racks.

Memory Prices Push US PC Shipments Down 7%
Chips & Semiconductors

Memory Prices Push US PC Shipments Down 7%

Omdia data cited by Tom Hardware showed US PC shipments fell to 15.8 million units in the first quarter of 2026, as memory and storage chip shortages hit entry-level laptops and pushed the market toward a forecast 14.4% contraction.

Keep Reading

More Stories

Latest
Yokogawa Opens Singapore Hub For Industrial Cyber ResilienceCapital & PolicyOct 6, 2026Yokogawa Opens Singapore Hub For Industrial Cyber ResilienceYokogawa Engineering Asia has launched a Singapore center focused on OT cyber resilience, training, response planning and recovery coordination for Southeast Asia, Oceania and Taiwan.ClickFix Attack Uses Browser Cache To Hide Malware PayloadCybersecurityOct 6, 2026ClickFix Attack Uses Browser Cache To Hide Malware PayloadMicrosoft Threat Intelligence traced a ClickFix cache-smuggling method that preloads malware into browser caches, then uses file size checks and a pasted Run command to launch later credential-theft stages.VOA Tests Six-Month Startup Buildout Before Funding DecisionsFintech & Digital PaymentsOct 6, 2026VOA Tests Six-Month Startup Buildout Before Funding DecisionsTechCabal’s interview with VOA Venture Partners founder Victoria Olayide Adesanya describes a six-month build programme that lets the firm work inside African financial-infrastructure startups before deciding whether to invest.Bitcoin Holds $86,000 As Dollar Index Hits 18-Month HighCrypto/Web3Oct 6, 2026Bitcoin Holds $86,000 As Dollar Index Hits 18-Month HighCoinDesk reported that bitcoin stayed near $86,000 while the U.S. Dollar Index reached about 102.5, with U.S. rate expectations and European political risks strengthening the dollar backdrop.Google Freezes OSS Bug Bounty Reports After AI Submission FloodCybersecurityOct 6, 2026Google Freezes OSS Bug Bounty Reports After AI Submission FloodGoogle has stopped accepting new product vulnerability reports in its OSS VRP after invalid automated submissions swamped reviewers, while older reports and some Cloud VRP routes remain open.Fleuret AI Raises €4M For Continuous AI Pentesting PlatformCybersecurityOct 6, 2026Fleuret AI Raises €4M For Continuous AI Pentesting PlatformTech.eu reported that French startup Fleuret AI raised €4 million in pre-seed funding to develop an agentic-AI platform that turns penetration testing into a continuous security process.GFT Analysis Says AI Documentation Can Cut Maintenance Work 30%Fintech & Digital PaymentsOct 6, 2026GFT Analysis Says AI Documentation Can Cut Maintenance Work 30%A GFT Technologies analysis says AI-linked software documentation can cut maintenance effort and speed developer onboarding when knowledge assets stay synchronized with code changes.Schneider Electric Lines Up $22.6 Billion PTC DealAIOct 5, 2026Schneider Electric Lines Up $22.6 Billion PTC DealSchneider Electric plans to buy PTC in a cash transaction valuing the US engineering software provider’s equity at about $22.6 billion, adding product-lifecycle software to its industrial AI push.Aggarwal Pledges Ola Electric Stake To Fund ₹1,000 Cr Rights IssueCapital & PolicyOct 5, 2026Aggarwal Pledges Ola Electric Stake To Fund ₹1,000 Cr Rights IssueOla Electric founder Bhavish Aggarwal pledged 20 Cr shares to finance his participation in a rights issue that forms part of a larger ₹1,500 Cr fundraising plan.Natrona Schools AI Review Puts Student Privacy Ahead Of Classroom Tool UseAIOct 5, 2026Natrona Schools AI Review Puts Student Privacy Ahead Of Classroom Tool UseNatrona County trustees questioned whether teacher AI tools expose student data, even as existing district rules already ban unauthorized generative AI use by students.New Relic Reports US$18 Million GreenOps Savings After AI CertificationCloud & Data CentersOct 5, 2026New Relic Reports US$18 Million GreenOps Savings After AI CertificationA New Relic company news item carried by iTWire says the observability vendor has earned ISO/IEC 42001 certification, joined the EU AI Pact and reported US$18 million in GreenOps savings from more than 80 engineering initiatives.VDURA V12 Targets GPU Clouds With One Storage PlaneCloud & Data CentersOct 5, 2026VDURA V12 Targets GPU Clouds With One Storage PlaneVDURA has made its V12 software generally available, adding tenant isolation, API-driven provisioning, tiering and RDMA data paths for GPU cloud and AI-factory storage fleets.