Hasty Briefsbeta

Bilingual

MicroLLM Lab – Try 7 tiny LLM's in the browser

4 hours ago
  • Objective checks (regex/exact tokens) are used, not writing quality.
  • A 135M model is allowed to fail — that is the measurement.
  • Pick models, then run. Estimate uses last tok/s if we have one.
  • Runtime per test (ms), accuracy per test, speed (tokens/s sustained decode, suite wall) and accuracy (pass rate on objective tests) from runs in this browser.
  • Numbers stay on this machine. Charts use the latest suite per model.
  • Generate and download a verifiable performance certificate with device hardware, peak and sustained tokens/second, and share your score.
  • Write a benchmark in JavaScript: the editor is eval()'d in this origin, then each check runs on the model’s output.