New tests put AI models through real work and vision
Artificial Analysis opened a service for testing models on a user’s own data and task costs, while Moonshot’s PerceptionBench found no frontier model exceeded 60 per cent on visual tasks. The developments shift attention from generic rankings to practical limits.
Artificial Intelligence··Midday
Optima tests a user's own work
Artificial Analysis has opened Optima for comparing models on a user's own data and use case. According to The Decoder, users can provide three kinds of input: existing evaluation sets, agent traces from tools such as Arize or Langfuse, or only a description of the use case. Optima uses those inputs to compare current models on answer quality, cost per task and time per task. The service offers two scoring routes: evaluating an answer against a rubric, or placing two answers side by side and asking which is preferred. Billing is based on the actual token costs of the models used, with no markup; it charges 0.125 dollars per criterion per model and 0.375 dollars per pairwise comparison. Early users built tests for finance and accounting agents and tried matching a legal writing style. The tool is available now, allowing organizations to view results, time and cost together on their own examples instead of relying only on a general ranking.[1]
PerceptionBench measures visual limits
Moonshot AI published PerceptionBench, a test of visual perception with 3,000 tasks spread across ten skill areas. In the results reported by The Decoder, none of the 16 frontier models exceeded 60 per cent. GPT-5.6 Sol had the highest result at 59.7 per cent; Kimi K3 scored 58.5 per cent and Claude Fable 5 scored 57.2 per cent. The areas cover counting, visual relations, attributes, depth and three dimensions, localization, comparison, fine-grained recognition, context integration, OCR and hallucination. Hallucination was the weakest area for every model. The tasks and scoring code are available in the MoonshotAI/PerceptionBench repository. The results also carry an important limit: Moonshot AI, which published the test, owns the second-placed Kimi K3 model. The ranking therefore counts as the company's own measurement rather than an independent evaluation. The results describe a low ceiling on these visual tasks, while the source of the model ranking needs to remain explicit.[2]
The two tests ask different practical questions
Optima and PerceptionBench make model performance concrete at different scales. Optima, opened by Artificial Analysis, lets an organization use its own examples to measure which model produces the desired answer, how long a task takes and the cost per task. PerceptionBench, published by Moonshot AI, applies the same visual tasks to 16 frontier models and compares results across defined skill areas; no model exceeds 60 per cent. The first approach creates an evaluation that varies with the user and task. The second provides a cross-model comparison on a shared task set. The sources also have different limits: an Optima result depends on the uploaded data, rubric and preference method, while the PerceptionBench ranking is a company measurement because it was published by Moonshot AI, the owner of Kimi K3. Together, they show that model selection can examine both results, time and cost on an organization's own work and performance on shared tasks for a defined capability. Optima works with specific user examples, while PerceptionBench applies the same visual task set to every participant. Those different inputs define the scope of the results.[1], [2]