SIE: Superlinked Inference Engine
- Type
- repo
- Venue
- Superlinked (GitHub)
- Year
- 2026
- Source
- github
- Access
- free
- Language
- en
- Added
- 2026-08-14T19:50:00Z
- Verified
- 2026-08-14T19:50:00Z
Summary
Apache-2.0 inference server (superlinked/sie). Argument: agent pipelines now run 4–5 small models, but one-server-per-model (vLLM + TEI + FastAPI wrappers) bills idle GPU time. SIE puts encode()/score()/extract()/generate() behind one API, loads models on first request, and evicts LRU so one GPU rotates the set. Catalog ~112 models with packaged serving configs. Plugs into Qdrant/Weaviate/Chroma/LanceDB/LangChain/LlamaIndex; OpenAI-compatible URL. Discord/X is Akshay Pachaar’s recap plus X article “How to serve 5 models on one GPU” walking a flood-insurance claim through Docling, GLiNER2, bge-reranker-v2-m3, Grounding DINO, and Qwen3.5-4B. pip sie-server[local]. Python.
Keywords
sie · superlinked · inference · multi-model · vllm · tei · x
Topics
multi-model serving, GPU sharing, inference engines
Research notes
- Primary: GitHub README (Apache-2.0, Python, 2786 stars / 271 forks at check). Discord/X https://x.com/akshay_pachaar/status/2085279733139558405 via fxtwitter quoting X article https://x.com/akshay_pachaar/status/2084992645966016757 (id 2084270232458420224). Example https://github.com/superlinked/sie/tree/main/examples/insurance-claims-agent. Serving engine, not a hosted corpus, so no datasets_local row.