Rene-1: Open-weight classifier sets SOTA on Decision Index (+9 over Jev)
11 hours ago
- Rene-1 31B FP8 is a non-generative decision model that reads a document and returns calibrated probabilities for every option of typed questions in a single forward pass.
- It supports three question formats: yes/no, choice among options, and ordinal score with confidence and probability distributions.
- It can be used through Transformers via pipeline or AutoModel with trust_remote_code=True, and requires an NVIDIA GPU with FP8 support such as Ada Lovelace, Hopper, or Blackwell.
- Model weights are 33.3 GB in FP8 and the full model fits on one GPU; typical peak memory is 44.5 GB, with larger requests needing 80 GB+ GPUs.
- On Decision Index 0.2, Rene-1 achieves 64.18 balanced skill across 37 benchmarks, a median calibration error of 0.046, and a median latency of about 92 ms for 5 questions over 641 tokens on a B200.
- It is designed for classification, routing, triage, policy checking, and ordinal grading, but is not suited for free-text generation, multi-step reasoning, non-English text, or high-stakes individual decisions without human oversight.
- Known limitations include lower skill on benchmarks like ACOS, ChessBench, and SGD, weaker calibration on PhishNChips, SGD, and POP909, and it provides no explanations for its answers.