Choose an open LLM, quantization, context, and GPU. Estimate VRAM and speed, create Ollama or llama.cpp starter output, and save a local deployment plan.
Select an open model, quantization, context window, KV cache precision, hardware profile, GPU count, parallel mode, and runtime. The lab uses the shared curated datasets to estimate weights, cache, overhead, usable memory, utilization, and single-user decode speed.
Generated Ollama and llama.cpp output is a transparent starting point; the lab never downloads or executes a model. Dataset dates and assumptions remain visible, while focused calculators and measured benchmarks provide deeper validation.
No. It performs local arithmetic from curated architecture and hardware datasets and generates reviewable starter text. Downloads, commands, servers, and models run only if you choose to use that output elsewhere.
Memory uses model weights, architecture-aware KV cache, and approximately 6% runtime overhead with a minimum allowance. Speed is a single-user bandwidth model. Runtime kernels, batching, offload, thermals, and prompts change real results.
Published bandwidth describes hardware potential, not your complete system. A measured benchmark captures browser and driver support, power limits, memory behavior, and other machine-specific constraints that an estimate cannot observe.
Detect your GPU with one click and see which LLMs your computer can actually run — ranked by whether they fit in your VRAM, need CPU offloading, or will not run at all.
Calculate how much VRAM any LLM needs to run locally. Pick a model, quantization, and context size — see download size, total memory required, and which GPUs it fits on.
Estimate LLM tokens per second from memory bandwidth, model size, quantization, and context window. Compare generation speed across GPUs and understand the memory-bandwidth bottleneck.
Build Ollama commands without memorizing syntax: run, pull, and create commands for any model, complete Modelfiles, and server environment configuration — with built-in VRAM checks.
Speed test your GPU for AI — measure real memory bandwidth and compute with WebGPU, run an actual LLM in your browser, and see predicted speeds for every popular model on your hardware.