← Back to explorer

SIE: Superlinked Inference Engine

Type
repo
Venue
Superlinked (GitHub)
Year
2026
Source
github
Access
free
Language
en
Added
2026-08-14T19:50:00Z
Verified
2026-08-14T19:50:00Z

Summary

Apache-2.0 inference server (superlinked/sie). Argument: agent pipelines now run 4–5 small models, but one-server-per-model (vLLM + TEI + FastAPI wrappers) bills idle GPU time. SIE puts encode()/score()/extract()/generate() behind one API, loads models on first request, and evicts LRU so one GPU rotates the set. Catalog ~112 models with packaged serving configs. Plugs into Qdrant/Weaviate/Chroma/LanceDB/LangChain/LlamaIndex; OpenAI-compatible URL. Discord/X is Akshay Pachaar’s recap plus X article “How to serve 5 models on one GPU” walking a flood-insurance claim through Docling, GLiNER2, bge-reranker-v2-m3, Grounding DINO, and Qwen3.5-4B. pip sie-server[local]. Python.

Keywords

sie · superlinked · inference · multi-model · vllm · tei · x

Topics

multi-model serving, GPU sharing, inference engines

Research notes

  • Primary: GitHub README (Apache-2.0, Python, 2786 stars / 271 forks at check). Discord/X https://x.com/akshay_pachaar/status/2085279733139558405 via fxtwitter quoting X article https://x.com/akshay_pachaar/status/2084992645966016757 (id 2084270232458420224). Example https://github.com/superlinked/sie/tree/main/examples/insurance-claims-agent. Serving engine, not a hosted corpus, so no datasets_local row.