The runtime around a language model that manages state, memory, tool integrations, and the agentic loop — e.g. Claude code, Codex, Grok build.
Tensions
Benchmark scores rank base models; product quality is largely set by the harness. A frontier model with a weak harness ships as a chatbot; a mid-tier model with a strong harness ships as a useful agent. This makes "frontier" an ambiguous category — frontier model and frontier system are different things, and they need to be co-developed.