Artículo: LLMaaS Pricing: Discovery through computational limits
Archivos
Fecha
Autores
Editor
Publicado en
Licencia Creative Commons
Resumen
Large Language Models as a Service (LLMaaS) are shift toward on-premise deployments for data sovereignty. These changes introduce significant challenges regarding resource elasticity. In these fixed-capacity environments, providers must establish regulations to prevent service over-utilization, while users require transparency regarding their usage limits. In this paper, we employ a framework to to identify the maximum sustainable throughput and maximum instantaneous throughput of a locally deployed LLMaaS. These metrics are used to derive user-centric metrics, such as requests-per-minute and tokens-per-day, and to introduce the consumption unit, a normalized ``computational currency'' that quantifies how individual interactions deplete finite system capacity. This currency enables the elaboration of a first iteration of a pricing generator, that translates these raw hardware constraints into tiered subscription models. Finally, we demonstrate the framework's viability by applying it to a real-world LLMaaS deployed on a high-performance computing cluster.


