n8loom
- Type
- repo
- Venue
- N8python (GitHub)
- Year
- 2026
- Source
- github
- Access
- free
- Added
- 2026-08-14T18:35:00Z
- Verified
- 2026-08-14T18:35:00Z
Summary
Python library (mlx_lm + Transformers) that stores per-node KV-cache fragments and concatenates the ancestor path when generating, so many ToT branches can share prefix cache without duplicating the full prompt cache at each node. Core types are Heddle (a reasoning node) and Loom (chat-root Heddle); ramify/crown helpers batch- or stream-expand children. Llama-architecture only. pip install n8loom; optional FastAPI example server. CC0.
Keywords
n8loom · tree-of-thought · kv-cache · mlx · llama · loom · heddle
Topics
tree-of-thought inference, KV cache, LLM serving
Research notes
- Primary: GitHub README + API (CC0-1.0, Python, 80 stars / 3 forks at check). Discord posted the repo (readme basic-script-example fragment). Inference library, not a new hosted corpus, so no datasets_local row.