Intern-Decision-2B scores options without writing an answer
InternLM has released Intern-Decision-2B, which assigns probabilities to answer options for named questions about a shared state in one model pass instead of writing a free-form reply. It can also take images as input. A code repository linked from the model card became public after initially returning an error, exposing the inference and evaluation paths. Its published benchmark and speed figures remain the maker’s measurements.
Artificial Intelligence··Morning
InternLM releases a model that scores options
InternLM released Intern-Decision-2B on September 26 as a model for structured decisions. The model card on Hugging Face identifies it as a multimodal model fine-tuned from Qwen3.5-2B; OrcaRouter also reports the release. Instead of composing a free-form answer, it assigns probabilities to options supplied for named questions about a shared state. Images may accompany that state. The request specifies the questions and their candidate answers before the model runs. A single pass then produces a distribution for each question, leaving the choices in the form the requester supplied. The output includes the probabilities, a confidence value taken from the highest probability, and the option selected by that value. This is a release of weights and a defined inference method, rather than a demonstration that the model independently decides which questions ought to be asked.[1], [2]
One pass reads scores before each placeholder
The card describes how the request becomes a prompt: a system instruction, the shared state, the question schema, and an assistant-shaped response with a placeholder for each field. Each candidate option is mapped to a single-token symbol. The model reads the next-token scores immediately before a field's placeholder, keeps only the symbols allowed for that field, and converts their scores into probabilities. Calibration is then applied before those symbols are mapped back to the original options. This path does not call the usual text-generation function or sample a string of words. A request can carry one to sixteen questions, with up to sixty-two options for each and as many as eight images. The default input limit is 8,192 tokens; an overlong request is rejected without being silently cut down. Those limits define the interface a developer must supply, not independently measured performance.[1]
The linked training and inference code becomes public
A code repository linked from the model card was initially inaccessible and became public on the same day, OrcaRouter reports. The available repository covers training, inference, a demonstration and evaluation, according to the model card and the separate report. That gives developers a route to inspect how the option scores are computed and checked, beyond using the released weights alone. The card lists Apache 2.0 alongside the continuing Qwen license. Its own benchmark table and speed figures are publisher-supplied measurements, so the release does not by itself establish those results through an independent rerun. The card also describes a calibration value chosen on a separate set of cases: it changes the confidence assigned to candidate options while preserving the option with the highest probability. The code repository is now accessible alongside the model weights and the described decision interface. Published performance figures remain the model maker’s measurements.[1], [2]
Related columns
For more information on this topic, you can read the related columns.