Benchmark
A standardised test or evaluation framework designed to measure AI system performance across specific tasks, enabling objective comparison between different models or approaches. Benchmarks provide quantitative metrics for capabilities like accuracy, reasoning, language understanding, or domain-specific skills, helping businesses make informed decisions when selecting AI solutions. Organisations use benchmarks to evaluate AI tools before adoption, track performance improvements over time, and validate that systems meet operational requirements. Common business-relevant benchmarks assess capabilities like document understanding, customer query handling, code generation quality, and decision-making accuracy across industry-specific scenarios.