Back to All Tools
Verified Secure Tool

How to use Local WebAssembly LLM Chat

  • 1Select a local model architecture from the WebGPU Engine Config panel.
  • 2Click 'Download Weights & Load Engine' to simulate moving the model into your browser's VRAM.
  • 3Upload text or markdown files into the Local Vector Store to enable local RAG (Retrieval-Augmented Generation).
  • 4Ask a question in the chat box and watch the model stream a response at over 40 tokens/sec, totally offline!