← Back to explorer

Photon 2.0: Inference engine for Physical AI

Type
other
Venue
Moondream (M87 Labs)
Year
2026
Source
web
Access
free
Language
en
Added
2026-08-14T19:25:00Z
Verified
2026-08-14T19:25:00Z

Summary

Moondream launch blog (2026-08-03). Argument: chat engines (vLLM/SGLang) optimize large-batch tokens/s, while robots/cameras need low concurrency, tight p99, and several models sharing a GPU. A proprietary compiler traces the forward pass and emits one megakernel so the GPU runs the whole inference without CPU launch chatter. First matrix: Moondream 2/3, Qwen3.5/3.6 0.8B–9B, Gemma 4 E2B/E4B on NVIDIA H100. Matched ChartQA streams: Photon beats vLLM and SGLang on every batch 1/2/4/8 throughput test and on cold start. pip install moondream; md.photon("Qwen/Qwen3.5-4B"). Engine Apache-2.0 and megakernels free to run; compiler closed. Discord/X is vikhyatk quoting the company thread.

Keywords

photon · moondream · megakernel · inference · physical-ai · vllm · x

Topics

inference compilers, megakernels, physical AI

Research notes

  • Primary: launch blog. Discord/X https://x.com/vikhyatk/status/2084409834523476073 via fxtwitter (quotes https://x.com/moondreamai/status/2084402190840738204). pip moondream; github.com/vikhyat/moondream and m87-labs/moondream-python. Docs https://docs.moondream.ai/. Inference engine, not a hosted corpus, so no datasets_local row.