Models

Liquid AI Launches d1-3B Multimodal Decision Model

Liquid AI has launched d1-3B, an open-weight multimodal decision model that bypasses text generation to deliver structured outputs in as little as eight milliseconds.

AlphaSignal18 hrs agoModels
Image: AlphaSignal

Liquid AI has introduced d1-3B, an open-weight 3.1-billion-parameter multimodal decision model designed to bypass the traditional text generation process. By eliminating the sequential decoding loop typical of generative transformers, the model processes inputs and returns structured predictions in a single forward pass with zero output tokens. This architecture allows developers to run classification, ranking, and scoring pipelines without the latency and parsing overhead of generating text.

The model achieves remarkable speed across various hardware configurations. It records a latency of just 8 milliseconds on an Nvidia RTX 4090, 9 milliseconds on an AMD MI325X, 16 milliseconds on a Jetson AGX Thor, and 50 milliseconds on a Jetson Orin Nano. In terms of capability, d1-3B scored 48.57 on the Decision Index 0.2.1 benchmark, outperforming all other models under 10 billion parameters and matching the performance of the much larger Decider 35B-A3B.

Built on the LFM2.5-VL-3B foundation model, d1-3B was developed using weight averaging, multi-seed fine-tuning, and checkpoint merging. Liquid AI also released an experimental, smaller sibling model called d1-omni-600M. The d1-3B model supports three specific API primitives for structured output: Noul, which provides a yes-or-no answer with a calibrated probability; Choice, which selects from a list of named options; and Score, which rates inputs against an ordered rubric.

For AI practitioners, this release changes how high-throughput classification tasks are handled. Instead of deploying massive generative models that require complex parsing logic and high computational budgets, developers can use d1-3B to run multiple decision queries simultaneously on a single input state containing text, JSON, or images. The model is currently available on Hugging Face with support for vLLM and SGLang, making it easy to integrate into existing edge and cloud deployment pipelines.

This is our own summary of reporting by AlphaSignal

More in Models