Zyphra Research Releases Zamba2-VL Vision-Language Model

View organization page for Zyphra

3,813 followers

Zyphra Research continues to explore architecture innovations beyond standard Transformers. Today we’re releasing Zamba2-VL, our second series of vision-language models, extending our prior Zamba2 work on hybrid SSM-Transformer architectures into the visual domain. Zamba2-VL is released at three scales: 1.2B, 2.7B, and 7B parameters. It is competitive with the leading open Transformer-based vision-language models of comparable scale across image understanding, reasoning, OCR, grounding, and counting benchmarks, while delivering roughly an order of magnitude faster time-to-first-token at every scale. Prior SSM-based vision-language models were either distilled from existing Transformer models or built on weaker open SSM bases. Zamba2-VL demonstrates that the inference-efficiency advantages of hybrid state-space LLMs carry cleanly into the multimodal setting. Pure SSMs struggle to recall specific details from earlier in a long input. The hybrid design keeps that ability by mixing in a small number of attention layers where it matters. We are releasing Zamba2-VL as a research artifact for the open community.  Zamba2-VL is released under Apache 2.0 with weights freely available on Hugging Face. Read the blog: https://lnkd.in/giZqTUSK Read the technical report:  arxiv.org/abs/2606.00390 Model weights on Hugging Face: https://lnkd.in/gRiYxRvz Inference code: https://lnkd.in/guGNFn9v

  • No alternative text description for this image

To view or add a comment, sign in

Explore content categories