Skip to main navigation Skip to search Skip to main content

Where the Time Goes: Analysis of a Public LLM Serving System

  • HES-SO Valais-Wallis
  • University of Lausanne

Research output: Conference Article in Proceeding or Book/Report chapterArticle in proceedingsResearchpeer-review

Abstract

In this study, we present a characterization of serving traces collected from Public Al's serving of Apertus, an open source Large Language Model (LLM). The trace spans roughly five months (September 2025-January 2026) and contains 337K requests. We analyzed request sizes, token and timing behaviour, latency, model-size effects, and temporal patterns. Our findings show insights that do not align with common assumptions; (1) time-to-first-token is often driven by queuing rather than prefill compute, especially for small requests; (2) the 8B and 70B models show nearly the same user-perceived latency despite a 9× parameter gap; (3) a substantial fraction of requests are prefill/queuing-dominated rather than decode-dominated; and; (4) observable input features are weak predictors of output, which makes size-aware scheduling difficult at arrival time. As a contribution to the research community, we will publish this anonymized trace along with its analysis.
Original languageEnglish
Title of host publicationProceedings of the Sixth European Workshop on Machine Learning and Systems, EuroMLSys 2026, Edinburgh, Scotland, UK, April 27-30, 2026
Number of pages12
PublisherAssociation for Computing Machinery
Publication date28 Apr 2026
Pages171-182
ISBN (Print)979-8-4007-2605-7
DOIs
Publication statusPublished - 28 Apr 2026
EventComputer Systems - Edinburgh, United Kingdom
Duration: 27 Apr 202630 Apr 2026
Conference number: 21

Conference

ConferenceComputer Systems
Number21
Country/TerritoryUnited Kingdom
CityEdinburgh
Period27/04/202630/04/2026

Fingerprint

Dive into the research topics of 'Where the Time Goes: Analysis of a Public LLM Serving System'. Together they form a unique fingerprint.

Cite this