← Back to explorer

Automated Tensor Scheduling for Hybrid CPU-GPU LLM Inference on Consumer Devices

Type
paper
Venue
arXiv / Nanjing University
Year
2026
Source
arxiv
Access
free
Language
en
Added
2026-08-14T19:15:00Z
Verified
2026-08-14T19:15:00Z

Summary

ATSInfer (Nanjing University) extends llama.cpp (~15k C++ lines) for consumer hybrid inference. Static knapsack placement uses measured empirical performance density (latency per byte) plus backend-switch costs; load-aware dynamic transfer promotes CPU-resident tensors when PCIe overlap can hide the copy; async coordination splits compute vs copy-engine streams. Vs llama.cpp under the same VRAM budget: prefill up to 1.94×, decode up to 3.29× (laptop) / 3.12× (RTX 4090), ~70% higher average decode SM utilization. Also beats vLLM offload and KTransformers on the models those support. Hardware: RTX 3060 6GB+32GB (Qwen3-14B INT4, Qwen3-30B-A3B INT4, GPT-OSS-20B MXFP4, GLM-Z1-9B FP16) and RTX 4090 24GB+64GB (Llama 3.1-70B INT4, Qwen3-Next-80B-A3B INT4, Qwen3.5-122B-A10B INT4, GPT-OSS-120B MXFP4). No public code linked in the paper.

Keywords

atsinfer · llama-cpp · offloading · hybrid-inference · moe · nju · x

Topics

LLM inference, hybrid CPU-GPU offloading, llama.cpp

Research notes

  • Primary: arxiv abs 2607.10183. Discord/X https://x.com/TeksEdge/status/2078695993063920025 via fxtwitter (third-party recap; notes no public repo). Equal-contribution Liu listed first. License left blank (HTML © none; no SPDX). Inference system paper, not a hosted corpus, so no datasets_local row.