How model and agent quality is measured; in 2026 the bottleneck is less whether a model can do a task and more whether the task is well-defined and well-measured.
Tensions
Yao Shunyu warns that every evaluation framework is easy to hack, because you can always make a metric look good without the underlying result being real (Goodhart). His proposed defense is not a better metric but dependable people who instinctively ask whether a good-looking number is actually good. This pushes the hard, differentiating work upstream into problem definition.