H Company Releases NeoMME: A Family of 260M and 800M Single-Tower Multimodal Encoders That Drop the Vision Tower and Causal Decoder
What changed
H Company released NeoMME, two models designed to search multilingual text and images together in one system, without a separate image-processing component or text-generation component. The feed reports that the smaller model scored 0.523 on a document-and-image retrieval test, indexed 51.3 pages per second on one L40S chip, and reduced index size by 255 times; the authors also noted weaker text-search results.
What this means for you
NeoMME could help build faster, smaller search systems for documents that mix writing and images. The feed does not say whether the models are publicly available, under what license, or what they cost, so practical access remains unclear.