A from-scratch LLM inference harness that runs entirely in the browser: C++ compiled to WebAssembly, compute on WebGPU. Early — build system and testable core only.
Description excerpted from the original listing, which is linked below.