LingBot-Vision
- Type
- repo
- Venue
- Robbyant (GitHub)
- Year
- 2026
- Source
- github
- Access
- free
- Added
- 2026-08-14T18:45:00Z
- Verified
- 2026-08-14T18:45:00Z
Summary
Robbyant/lingbot-vision ships LingBot-Vision, a family of self-supervised ViT backbones (S/16, B/16, L/16, g/16) for dense spatial perception. A ~1.1B ViT-g/16 teacher is pretrained with masked boundary modeling (boundary-centric masked targets that keep spatial structure plus semantics) and distilled to L/B/S. Drop-in encoder for PCA feature maps, depth, semantic/video object segmentation, and as the encoder init for LingBot-Depth 2.0 (RGB-D corpus scaled 3M to 150M in the tech report). Backbone-only .pt checkpoints on HF/ModelScope. Apache-2.0. Paper arXiv 2607.05247.
Keywords
lingbot-vision · masked-boundary-modeling · dinov3 · vit · depth · robbyant · self-supervised
Topics
self-supervised vision, dense spatial perception
Research notes
- Primary: GitHub README + API (Apache-2.0, Python, 908 stars / 40 forks at check; owner Robbyant). Paper arXiv 2607.05247 (cs.CV; HTML CC BY 4.0; comment Tech report, 31 pages). Project https://technology.robbyant.com/lingbot-vision. Discord posted paper.pdf; cataloged the repo URL as the leftover GitHub item. HF collection robbyant/lingbot-vision (vit-giant/large/base/small) not copied into hf_* because leftover is the GitHub repo. Weights/code item; LingBot-Depth 2.0 150M RGB-D set is described but not hosted here, so no datasets_local row.