Because systemic properties appear only at the whole-system level, evaluations and optimizations that isolate individual components (prompts, models, tools) cannot be trusted to improve or even predict the behavior of complete agent systems; system-level emergent outcomes need to be evaluated as a whole.
Many agent pipelines are debugged and benchmarked component-by-component, yet whole-system failure can arise from interactions that component-level metrics never expose. Whole-agent scenarios, global traces, and emergent-outcome probes become necessary, not optional.
Any artifact comprised of interacting software agents must be instrumented and evaluated at the ensemble level before micro-optimizing parts.
systemic properties only appear at the system level
Reductionism studies objects in isolation by taking them apart from the bottom up