What the new measures actually measure
CapQuiz stops asking how closely a generated description resembles a reference text. It scores the description by whether it supports answers to human-verified multiple-choice questions drawn from the video, across 10 question types and 24 video domains, and folds a factuality component and a coverage component into one figure. The reason given is plain: reference matching penalises a description for wording differences, and one video admits many valid descriptions. Swapping the measure changes which property a score is about, before any two models are compared.[1]
DiscoSign makes the same move somewhere else. Sign language systems have mostly worked one sentence at a time, which loses the structure a signed passage depends on: an entity that keeps its place in signing space, a clause form carrying a particular discourse function, a stable mapping between an English concept and its sign. The framework addresses 3 of these phenomena, and because ordinary translation scores do not register discourse-level quality, the work brings a measure for each dimension it claims to handle.[2]
The instrument and the result come from one place
The two papers share something beyond their subject matter. In each, the team that redefined the measurement also reports the improvement measured by it: CapQuiz is said to correlate significantly better with human judgement than existing metrics, and discourse-aware processing is said to improve spatial consistency and entity tracking against sentence-only translation. Both statements may well hold. Read as evidence, though, they stay at the level of internal validation, because the yardstick and the reading of it come from the same place.[1], [2]
There is a friendlier reading, and it deserves saying. A new measure can be right where the old one was wrong, and a group outside the authors' own could reproduce the same ordering with its own measurements; then the instrument would not be carrying the result. The CapQuiz paper is careful on exactly this point: the comparison with human judgement is the authors' own, and the work has not been through external validation by another group. That is an honest description of where the claim sits.[1]
What would move this up the evidence ladder?
The next thing worth watching is whether either measure gets picked up by people who did not build it: a group outside the original teams running CapQuiz, or the DiscoSign discourse metrics, on models the original work did not cover. If that happens, the ranking each measure produces becomes checkable against a second hand, and the question of whether the instrument or the method produced the gain gets an answer. Until then the fair summary is that two papers have proposed a better question and answered it themselves.[1], [2]