Automated Tensor Scheduling for Hybrid CPU-GPU LLM Inference on Consumer Devices
- Type
- paper
- Venue
- arXiv / Nanjing University
- Year
- 2026
- Source
- arxiv
- Access
- free
- Language
- en
- Added
- 2026-08-14T19:15:00Z
- Verified
- 2026-08-14T19:15:00Z
Summary
ATSInfer (Nanjing University) extends llama.cpp (~15k C++ lines) for consumer hybrid inference. Static knapsack placement uses measured empirical performance density (latency per byte) plus backend-switch costs; load-aware dynamic transfer promotes CPU-resident tensors when PCIe overlap can hide the copy; async coordination splits compute vs copy-engine streams. Vs llama.cpp under the same VRAM budget: prefill up to 1.94×, decode up to 3.29× (laptop) / 3.12× (RTX 4090), ~70% higher average decode SM utilization. Also beats vLLM offload and KTransformers on the models those support. Hardware: RTX 3060 6GB+32GB (Qwen3-14B INT4, Qwen3-30B-A3B INT4, GPT-OSS-20B MXFP4, GLM-Z1-9B FP16) and RTX 4090 24GB+64GB (Llama 3.1-70B INT4, Qwen3-Next-80B-A3B INT4, Qwen3.5-122B-A10B INT4, GPT-OSS-120B MXFP4). No public code linked in the paper.
Keywords
atsinfer · llama-cpp · offloading · hybrid-inference · moe · nju · x
Topics
LLM inference, hybrid CPU-GPU offloading, llama.cpp
Research notes
- Primary: arxiv abs 2607.10183. Discord/X https://x.com/TeksEdge/status/2078695993063920025 via fxtwitter (third-party recap; notes no public repo). Equal-contribution Liu listed first. License left blank (HTML © none; no SPDX). Inference system paper, not a hosted corpus, so no datasets_local row.