Etalon: Holistic Performance Evaluation Framework for LLM Inference Systems
- Type
- paper
- Venue
- arXiv:2407.07000 (cs.LG), submitted 9 Jul 2024, v2 revised 30 Aug 2024
- Year
- 2024
- Source
- arxiv
- Access
- free
- Language
- English
- Added
- 2026-10-02
- Verified
- 2026-10-02
Summary
Argues that conventional LLM serving metrics (TTFT, TBT, normalised latency, TPOT) fail to capture the nuances of streamed inference and user-facing performance for real-time applications like chat and translation. Etalon is a comprehensive evaluation framework built around per-request token arrival traces and deadline-based evaluation: its fluidity-index measures the fraction of per-token deadlines met within a request (tokens arriving early build slack; a missed deadline counts missed slots and resets subsequent deadlines from that token's arrival), and the fluid token generation rate is the highest playback rate inferred from evaluated requests that meets a chosen fluidity target for a specified share of requests. The framework connects scheduling behavior (prefill/decode interference, chunked prefill, speculative decoding with client buffering) to service targets, and supports capacity search: estimating the maximum request rate a replica can sustain under specified targets to plan deployment size. The authors evaluate open-source platforms and model-as-a-service offerings with Etalon, exposing skewed token-generation patterns that averages hide.
Keywords
Etalon · LLM inference systems · fluidity-index · TTFT · TBT · TPOT · token arrival traces · deadline-based evaluation · capacity search · chunked prefill · speculative decoding
Topics
LLM inference, serving, performance evaluation, latency metrics, fluidity-index, TTFT, capacity planning
Research notes
- Discovery: @wafer_ai X thread 2026-10-01 (https://x.com/wafer_ai/status/2105837274403676452), part 8 of its AI performance engineering reading series
- Docs: https://project-etalon.readthedocs.io/en/latest/tutorials/metrics_used.html, .../prefill_profiler.html, .../capacity_search.html; code linked from the arXiv page
- Thread takeaways: a long prefill can delay ongoing decode work; chunked prefill divides prompt processing across batches, and the amount of prefill work per batch affects both first-token latency and interruptions to existing streams
- TTFT includes scheduling delay plus prompt processing; proposed method profiles isolated requests across prompt lengths and adds a scheduling allowance to set first-token deadlines
- TPOT can hide generation pauses, and TBT percentiles omit where pauses occurred within a response; per-request token arrival traces preserve the order and gaps the client observes
- Speculative decoding: several tokens can arrive in one burst; client buffering for steady display lets work completed ahead of the consumption rate cover later gaps
- For controlled deployment comparisons, keep model, hardware, workload, and service targets fixed, then use capacity search to compare serving configurations
- Figure highlighted in the thread: Etalon evaluation of commercial offerings for Mixtral-8x7B and Llama3-70B over 24 hrs (fluidity-index, request volume, TTFT/TBT distributions)
- Series resource also linked: github.com/wafer-ai/gpu-perf-engineering-resources
- License not stated on the arXiv page (generic 'view license' shown)