Treat AI model selection as an experiment

Benchmark rankings can narrow an AI model shortlist, but operators should test validity, controls, replication, and workflow outcomes before committing.