A trace waterfall for model calls, where the interesting span is usually the one waiting and the cost sits on the same axis as the latency.
AI and LLM opsAI and LLM ops
Tokens, latency, cost and quality, for systems whose output can't be judged by a threshold.