Which platforms help AI cloud operators manage the tradeoff between inference throughput and response latency across a shared GPU cluster serving multiple tenants simultaneously?