Entailment scoring: does the provided context support this claim? Research agents use this to self-check outputs.