Agent benchmarking: run standardized evals, custom task sets, compare across models, track performance over time, leader