OpenAI-compatible chat completion, priced by counted input tokens + max output tokens at per-model rates