Where do 600 billion and 27 billion sit?

StepFun says Step 5 Preview's sparse mixture-of-experts architecture contains 600 billion total parameters and activates 27 billion parameters for each token. The company also announced a 1 million-token context window together with text and image input. These three figures measure different capacities: 600 billion describes the complete weight set, 27 billion the subset selected while processing one token, and 1 million the context length a request can carry.[1]

The active-parameter count shows that the full set of 600 billion weights does not enter computation for every token. That figure alone does not state the model's memory footprint; the weight precision, compression format and placement of experts across devices have not been disclosed. Reading 27 billion as the deployment size, or 600 billion as the load used for every request, would therefore combine two different denominators.[1]

The delivered layer is the API

The product available today is the API; the announced date for open weights is 15 October. In its own 24-hour GPU-kernel trial, StepFun reports that an MLA kernel reached 508 TFLOPS of peak performance, compared with 493 TFLOPS for Claude Opus 5. This measurement on a company-selected workload is not independent and does not provide production throughput per server, latency or power use.[1]

A team using the API can measure response time and its bill against its own traffic today. A team planning local deployment cannot calculate required memory and sustained throughput until weight files, precision options and hardware placement are published. If those details arrive on 15 October, the two parameter counts can become inputs to a real capacity plan; detailed API measurements could also provide a limited preliminary estimate before the weights. For now, delivered capacity is observable within the boundaries of the API operated by StepFun.[1]