Score agent response confidence using hedging/definitive language signals, citations and task-type adjustments.