AI safety: output toxicity scoring, bias detection across demographics, alignment evaluation, red team scenario library