Models

Jev-Omni Classifier Scores Multimodal Decisions

The new open-weight Jev-Omni model classifies decisions across text, image, audio, and video inputs in milliseconds, eliminating the need for developers to parse complex generative text outputs.

AlphaSignal3 days agoModels
Image: AlphaSignal

Developers looking to streamline structured decision-making can now utilize Jev-Omni, a 12-billion-parameter open-weight classifier. Built by fine-tuning the Gemma 4 12B IT model on a dataset of 30,000 questions, this Apache-2.0 licensed tool is designed to evaluate a given context alongside a question and pre-defined options. Instead of generating conversational text that requires complex parsing, Jev-Omni directly assigns calibrated probabilities to the provided choices, supporting yes-or-no, multiple-choice, and scoring formats.

The model demonstrates strong performance across several benchmarks, scoring 87.57% on DecisionBench Medium, 63.10% on MMAU, and 53.10% on MVBench. It also achieves an expected calibration error of 0.0400 across 10 confidence bins on the DecisionBench Medium split, indicating a tight alignment between its predicted confidence levels and actual accuracy. This level of calibration ensures that the probability scores returned by the model are highly reliable for downstream applications.

In terms of speed, Jev-Omni delivers rapid inference times. When tested on a warm Nvidia H200 GPU, the model records latencies of 83 milliseconds for text, 26 milliseconds for images, 31 milliseconds for audio, and 504 milliseconds for a 16-frame video. To run the model, practitioners will need approximately 50GB of FP32 storage, though inference can be executed in BF16 format on a CUDA-enabled GPU. It performs best when evaluating 20 or fewer options.

For AI practitioners, Jev-Omni represents a significant shift away from traditional generative pipelines. By bypassing the need for regular expressions, constrained decoding, or secondary judge models to interpret text outputs, it simplifies the architecture of multimodal applications. The model is independent of TypeSafe AI's Jev and is currently available to test for free on a hosted Space.

This is our own summary of reporting by AlphaSignal

More in Models