s1-pro
Rune · EU
Calibrated text decisions, with up to three questions per call. Rune is the main EU model.
Models · text, images and routing
Surogate Rune 26B-A4B v3 is the main EU model behind s1-pro and s1-vision. Choose s1-fast for short text decisions or the CPU auto-router to select an LLM. Peer-to-Peer offers the lowest launch price; EU keeps inference in the EU. Model-reference benchmarks below describe the tested reference implementations; a hosted capability refresh is pending.
Rune · EU
Calibrated text decisions, with up to three questions per call. Rune is the main EU model.
Plumb-4B · short text
Short text, low latency. Runs on a batched engine, so probabilities for identical requests can differ slightly.
Rune · EU · one image
Choices with one image attached.
CPU
Classifies category, difficulty and stakes to help your router choose an LLM.
When Rune is unavailable, s1-pro text falls back to Winnow-12B Q8; on the EU tier and on Peer-to-Peer without worldwide opt-in s1-vision has no fallback and returns a retryable 503; Peer-to-Peer keys with worldwide processing opt-in may be served by on-demand worldwide GPUs running the Gemma 4 12B image implementation when capacity is busy. On Peer-to-Peer, Rune serves when spare queue/admission capacity permits, with a compatible single-question fallback; bundles are best-effort and may return retryable HTTP 503. Full model list: /models · Licence.
Input limits
Limits apply per request and are counted with the selected model’s own tokenizer after the state, question instructions and answer options are combined. Code, numbers and IDs use more tokens per character than ordinary prose. The answer is a typed decision, so there is no separate output budget. The state is at most 16 KiB of serialized JSON for s1-fast and 250 KiB for s1-pro and s1-vision; the request body is at most 6 MiB.
| Model | Max input per request | Above the limit |
|---|---|---|
s1-fast | 4,096 tokens | HTTP 400 invalid_request; nothing is truncated or billed |
s1-pro | 32,000 tokens, all questions of the request together | HTTP 400 invalid_request; nothing is truncated or billed |
s1-vision | 32,000 tokens, including the image (a 64×64 test image added about 250 tokens; larger images use more) | HTTP 400 invalid_request; nothing is truncated or billed |
s1-llm-auto-router | Reads at most 512 tokens: the first 600 characters of context (if any), then the request, cut so the total is 512 | Accepted; the rest is ignored and not billed |
Requests above 16 KiB of state are always served by the EU Rune node . Requests between 4,096 tokens and that size are served by Rune in the EU tier; on Peer-to-Peer or during an EU fallback a 4,096-token node may refuse them with HTTP 400 (not billed) — retry on EU. Requests above 16 KiB of state may queue: when long-input capacity is busy the API answers HTTP 429 with Retry-After: 3; the request is not billed, just send it again. Measured on 7 October 2026 on our own test pod (not on the production API, one request without load): about 0.4 s at 8,000, 0.9 s at 16,000 and 2.3 s at 32,000 tokens. Measurements.
Earlier measurement on the production API on 6 October 2026 (the limit was then 4,096 tokens for every model): requests of 4,095 tokens (s1-fast) and 4,094 tokens (s1-pro) were accepted; requests a few tokens longer were rejected. Earlier measurements.
Rune scores close to Jev on both public reference benchmarks. Our EU input rates are below Jev’s list price. Compare measured latency in the chart below. Surogate Rune 26B-A4B v3 is the main EU model, served as s1-pro for calibrated text decisions. The indexes below are public reference results — not hosted Decision Models scores — and JevBench is run by the founder of Decision Models. Method and source data: /benchmarks · download the data (JSON).
Balanced, chance-corrected index / 100 ↑ · v0.2.1
Snapshot: Decision Index v0.2.1, 2 October 2026 · see the current board
Ours – s1-pro · reference modelJev 1.13.0 (reference API)Other tested models
Scale 0–100
Source: Decision Index · data · retrieved 2 October 2026
Public default balanced index; equal weight across five areas. Rows retain the source order and use sequential display ranks, including tied scores. Jev is the reference API; other entries are tested model implementations. Reference results are not a hosted System1 measurement.
Mean of Intelligence and Calibration / 100 ↑ · v1.5.5 · cap-admitted
Snapshot: JevBench Capability Score v1.5.5, 2 October 2026 · see the current board
Ours – s1-pro · reference modelJev 1.13.0 (reference API)Other tested models
Scale 0–100
Source: JevBench Capability Score · data · retrieved 2 October 2026
Only models inside the cost and median latency caps (each 2× Jev 1.13.0) are ranked. Benchmark run by the founder of System1 Models. Tested reference engines; the hosted System1 engine has not been scored in this release.
Medium text · one question · warm HTTP/2 · median ms ↓
Decision ModelsJev 1.13.0
Scale 0–300 ms
Measured from Nuremberg, Germany, 2 October 2026. 33 rotated rounds on identical medium synthetic support tickets, one question, HTTP/2 persistent clients; first 3 rounds excluded as warm-up; 30 measured calls per target. p50=sample median; p95=nearest rank ceil(0.95*n). End-to-end successful-response latency; every failure retained. s1-pro has a lower median but a higher p95 than Jev. Small sample; results vary with workload and network.
Source: Artifact (JSON) · retrieved 2 October 2026
The measured latency comparison stands in the speed chart above — medians, p95 and success counts from the linked measurement artifact. Public index results are reference-only — not hosted Decision Models scores.
Public synthetic text; real customer API end-to-end latency, successful-response quantiles; failures tracked separately; no retained customer payloads. p95 uses the nearest sample percentile.
$0–0.05 per M input tokens
Decision ModelsJev 1.13.0
$0–0.05 per M input tokens
Launch pricing · output free · taxes excluded.
Source: Decision Models pricing · Jev 1.13.0 price · retrieved 2 October 2026
$0–0.012 per 1,000 decisions
Decision ModelsJev 1.13.0
$0–0.012 per 1,000 decisions
Measured 3-question cost estimates. EU and Jev: 30/30 successes. Peer-to-Peer: 14/30 from Finland, 12/30 from Virginia; estimates cover successful calls only; Peer-to-Peer bundles are best-effort and may return retryable HTTP 503.
Source: Artifact (JSON) · retrieved 2 October 2026
2 October 2026. Medium text, 3 questions per call, from Finland. Mean billable input tokens × USD list rate / 3 decisions × 1,000. EU/Jev 10/10 successful medium calls; Peer-to-Peer 5/10, estimates cover successful calls only. Internal trial requests were not paid debits. Across all S/M/L cells: EU/Jev 30/30, Peer-to-Peer 14/30 in Finland and 12/30 in Virginia. Peer-to-Peer bundles are best-effort and may return retryable HTTP 503. Shared state billing reduces cost; large bundles can be slower than Jev.
Two approaches
Same model profiles and API request shape. Implementations depend on available capacity. Choose per API key or per request with one header.
Built to scale worldwide. Peer-to-Peer starts on spare EU capacity. When it is busy, keys you opted in can use on-demand GPU capacity rented through Lium, operated by independent third-party providers in various countries, including outside the EU/EEA.
Today
EU requests have priority. Enable “Allow worldwide processing” for each key that may leave the EU. Without opt-in, Peer-to-Peer stays on EU capacity and may be queued or rejected when busy. Existing customers keep EU-only processing unless they opt in. Request and response content is not stored on any node.
For teams with the highest data-protection demands or strict EU regulatory requirements. EU requests run only on contracted operators inside the EU, and inference data never leaves the EU. EU requests have priority, even if a key allows worldwide Peer-to-Peer processing.
Sign-in (Google) and payments (Stripe) are handled by the sub-processors listed on our sub-processor page; they never receive your prompts.
Public benchmark context
Each benchmark has equal visual weight. Read its metric and coverage separately; different tasks, adapters and hardware can produce different trade-offs.
| Model profile | Decision Index0.2.1Chance-corrected index / 100Independent community | Workflow Evalsdatasets 2026-09-28 · code 0ac3b8aModel-reference agreementTypeSafe · reference vendor | Jev Rerank Bench2026-09-25Dataset-macro nDCG@10 / 1Independent author | JevBench1.5.1Composite score / 100Same founder as Decision Models | ImageJevBench0.1.4Composite score / 100Same founder as Decision Models |
|---|---|---|---|---|---|
| s1-fastPlumb-4B | Not measured¹ | Not measured¹ | Not measured¹ | 71.56 | Not measured¹ |
| s1-proWinnow-12B Q8 | 50.02 | Not measured¹ | 0.634 | 73.23 | 43.56 |
| TypeSafe referenceJev 1.13.0 | 57.91 | 67.8 | 0.670 | 72.13 | Not measured¹ |
¹ Not measured means no comparable published measurement was verified for this model; it is not a zero. Results concern the tested model implementations, not the Decision Models service. These published results were measured on the implementations that served these profiles at publication time (Plumb-4B, Winnow-12B Q8 and the Winnow Q8 image path). Since 2 October 2026 the EU tier serves s1-pro and s1-vision on Surogate Rune 26B-A4B v3; Peer-to-Peer-tier requests may use Rune when spare capacity is available; Winnow-12B Q8 is the fallback for s1-pro (text), s1-vision has none on the EU tier or on Peer-to-Peer without worldwide opt-in (Peer-to-Peer keys with worldwide opt-in may be served by on-demand worldwide GPUs running the Gemma 4 12B image implementation when busy), and a hosted benchmark refresh on the current EU implementation is pending. ImageJevBench uses a separate image path; customer image input is available with s1-vision.
Conflict of interest: JevBench and ImageJevBench are run by the founder of Decision Models. Workflow Evals is published by TypeSafe AI, the vendor of the reference model. It measures model-reference agreement, not human-labelled accuracy. No overall score is published while comparable coverage is incomplete. Method and sources →
Image decisions
Send exactly one PNG, JPEG or WebP image (base64 data URL or public HTTPS URL, up to 4 MiB and 2 megapixels) with model s1-vision. Model credits are on the model licences page. The image evaluation is one equal column in the matrix above. Its result belongs to a separate Winnow Q8 image implementation. Image input for Decision Models customers is available with s1-vision; s1-fast and s1-pro accept text.
Jev is a trademark of TypeSafe AI; references to Jev are nominative compatibility references only. Decision Models is not affiliated with TypeSafe AI and does not resell Jev.