MicroLLM Lab – Try 7 tiny LLM's in the browser
4 hours ago
- Objective checks (regex/exact tokens) are used, not writing quality.
- A 135M model is allowed to fail — that is the measurement.
- Pick models, then run. Estimate uses last tok/s if we have one.
- Runtime per test (ms), accuracy per test, speed (tokens/s sustained decode, suite wall) and accuracy (pass rate on objective tests) from runs in this browser.
- Numbers stay on this machine. Charts use the latest suite per model.
- Generate and download a verifiable performance certificate with device hardware, peak and sustained tokens/second, and share your score.
- Write a benchmark in JavaScript: the editor is eval()'d in this origin, then each check runs on the model’s output.