Answer like Jev 98% of the time, locally, in 15 ms
Jevstiller is a drop-in proxy that sits in front of Jev's classification API. It trains a small local model from Jev's own answers and serves requests it is confident about locally in ~15 ms on a CPU, forwarding everything else to Jev. Its distinguishing feature is a statistical contract: you set a target agreement (e.g. 98%), and the system guarantees that at least that share of requests get the answer Jev would have given — with a 95% confidence bound, verified by benchmarks and a live audit slice.
Set a target (e.g. 98%) and Jevstiller returns the label Jev would have returned on at least that share of requests, with an exact Clopper–Pearson confidence bound holding with at least 95% probability per model version. The bound is computed with fixed-sequence testing on held-out calibration rows
A frozen bge-small sentence encoder (384 dimensions, ONNX Runtime) plus a multinomial logistic regression head answers confident, in-distribution requests in about 15 ms on a CPU, versus ~300 ms for a network call to Jev at any load.
The whole model is a few hundred kilobytes, trained in seconds on a few thousand rows by cross-entropy against Jev's full probability distribution (not just the top label), and retrained every 2,000 new Jev answers. Each version is shadow-tested on live traffic before promotion.
A k-nearest-neighbour scorer over the training embeddings detects unfamiliar inputs and routes them to Jev regardless of the local model's confidence. Its cutoff is calibrated on training data, never on the calibration rows.
A task can carry a confidence_floor: requests Jev would have answered below that confidence are treated as disagreements and forwarded to Jev, so downstream code that relies on Jev's 'unsure' signal keeps working. Costs 4–12 coverage points at a floor of 0.6; off unless configured.
A fixed share of requests (starting at 2%) goes to Jev regardless of the local model's answer, giving an unbiased live view of agreement. If the audit confirms a breach, all traffic falls back to Jev and training restarts. In a 24-hour soak test where the stand-in Jev silently changed every answer a
Published numbers come from a benchmark suite covering Banking77, CLINC150, AG News, and two TweetEval tasks (20 random splits each), rerunnable locally with 'bash experiments/bench.sh --no-record' without an API key.
Each request is embedded by a frozen bge-small encoder (384 dimensions, ONNX Runtime on CPU). A multinomial logistic regression head (one linear layer + softmax) is trained by cross-entropy against Jev's full probability distribution over labels — not just the top label — with full-batch Adam and ea
A request above a confidence threshold and inside the training distribution gets the local head's answer in ~15 ms; everything else goes to Jev. The threshold is picked from a fixed grid, tested strictest-first on held-out IID calibration rows using exact Clopper–Pearson bounds (fixed-sequence testi
The obvious recipe — sweep thresholds and keep the loosest one whose measured disagreement fits the budget — exceeded a 2% budget on 6–12 of 20 splits per task across five public tasks, because picking the loosest passing threshold selects on noise. Jevstiller's bound-based rule broke the budget onc
It sits in front of Jev's classification API, learns a small local model from Jev's own answers (a frozen bge-small encoder plus a logistic regression head), and answers confident, in-distribution requests locally in about 15 ms on a CPU. Everything else — low-confidence or out-of-distribution requests — is forwarded to Jev, which takes about 300 ms.
Per task, over a window of traffic, the share of all requests that end up with an answer Jev would not have given is kept under your budget (e.g. 2 in 100 at a 98% target), with the bound holding with at least 95% probability per model version. It is a statement about agreement with Jev over all requests, not about the local model's accuracy on the requests it chose to answer, and not about accuracy against the truth.
The local model answers in about 15 ms on a CPU. Requests it declines go to Jev at roughly 300 ms per call, at any load.
No GPU: the model runs on CPU via ONNX Runtime. The whole model is a few hundred kilobytes and trains in seconds on a few thousand rows.
Not claimed. Where the local model differs from Jev it is about as often right as Jev was — on the five benchmark tasks, system accuracy against the datasets' own labels stayed within a point of Jev's across targets from 99% down to 90% — but the page notes this held on those tasks and is not a law.
A fixed share of requests (starting at 2%) is always sent to Jev as an audit slice, the only unbiased view of production once routing begins. If the audit's bound confirms a breach, every request goes back to Jev and training restarts from that point. In a 24-hour soak where a stand-in Jev silently changed every answer at hour twelve, local share fell from 90% to 9% within four minutes and recovered to 90% within 49 minutes, with no manual intervention.
Yes, since version 0.4.0. A task can carry a confidence_floor: local answers are counted as disagreements when Jev would have answered below that floor, so such requests are forwarded to Jev and its real confidence comes back. It costs 4–12 coverage points at a floor of 0.6 on the benchmark tasks and is off unless you set it.
The project is on GitHub, the design decisions are documented in DESIGN.md (§7.3–§7.6 cover the routing code paths), and the published benchmark numbers can be rerun locally with 'bash experiments/bench.sh --no-record' without an API key. No specific license is stated on this page.