Run locally / Model details

Jeff-1.0-Large

31B text model.

Fully sovereign. Maximum data privacy: it runs on your hardware or in your own cloud account, and your data never leaves your machines.

Installer licence

Free for small teams. USD 1,000 + USD 100/month for larger companies.

Free for individuals and companies with up to 10 employees and up to USD 1M ARR.

Larger companies: USD 1,000 one-time + USD 100/month (self-declared). That applies above either threshold: more than 10 employees or more than USD 1 million ARR.

This is the licence for the Decision Models installer. The model's own weights licence is separate and shown further down.

Personal and non-commercial use only.

This installer entry is not cleared for commercial use. Read the model card before downloading or serving it.

Hardware

Minimum and recommended hardware

31B parameters. Sized for decision readouts with inputs up to about 8k tokens.

RequirementMinimumRecommended
Precisionbf16bf16
GPU memory (VRAM)70 GB80 GB
Example GPUs——
System RAM80 GB80 GB
Disk61.43 GB61.43 GB
CPU only
No The installer catalogue has no CPU variant.
Apple Silicon
Not yet The installer catalogue has no Apple Silicon variant.

How we computed this: weights size × bytes per value for the precision, plus 10 % and room for the context. For mixture-of-experts models all parameters must be in memory. Rounded up to common GPU sizes. Installer figures come from the reviewed catalogue.

Variants

Will it run on my machine?

Requirements are recorded per variant. A dash means the catalogue does not provide that value.

VariantPrecisionRuntimeBenchmarkedInstall recipeMinimum / recommended VRAMRAMDiskPlatformsExpected speed
bf16-jeff-servebf16TransformersYesTested install recipe70 GB / 80 GB80 GB61.43 GBlinux-nvidia, wsl2-nvidiaAbout 0.3 s per decision on the measured GPU (board latency 0.31 s). Not measured elsewhere.
bf16-jeff-servebf16
Runtime
Transformers
Install recipe
Tested install recipe
VRAM
70 GB minimum / 80 GB recommended
RAM
80 GB
Disk
61.43 GB
Platforms
linux-nvidia, wsl2-nvidia
Expected speed
About 0.3 s per decision on the measured GPU (board latency 0.31 s). Not measured elsewhere.
dm-local plan jeff-1-0-largeChecks memory, disk, platform, and runtime against the catalogue.

Quick start

Choose where to run it

Installer commands use the pinned model revision. Confirm the model terms before proceeding.

An install recipe has been tested. Platform and hardware coverage is shown per variant.

  1. 1curl -fsSL https://decisionmodels.io/local/install.sh | sh && export PATH="$HOME/.local/bin:$PATH"
  2. 2dm-local install jeff-1-0-large

Your cloud account

Run it on AWS, Azure or Google Cloud

Launch a GPU machine in your own account, run the model there, and keep everything inside your cloud account and the region you pick.

AWS

Instance for this model: on request. Pick a GPU with at least 70 GB of memory.

  1. Open the EC2 console and launch a GPU instance in an EU region of your choice, using the AWS Deep Learning AMI (GPU, Ubuntu).
  2. If the launch is refused, request GPU quota first: Service Quotas → EC2 → “Running On-Demand G and VT instances” (or P instances).
  3. Attach a security group that allows inbound SSH (port 22) only from your IP. Do not open the model port to the internet.
  4. Connect with SSH and run the one-line installer:
    curl -fsSL https://decisionmodels.io/local/install.sh | sh && export PATH="$HOME/.local/bin:$PATH"
  5. Install the model:
    dm-local install jeff-1-0-large
  6. Reach the endpoint through an SSH tunnel (ssh -L 8484:127.0.0.1:8484 user@instance) or from inside your private network, then call http://127.0.0.1:8484/v1/systemone.

Your data stays in your cloud account and region. Prices change; region, storage and traffic may add charges.

Azure

Instance for this model: on request. Pick a GPU with at least 70 GB of memory.

  1. Open the Virtual machines blade and launch a GPU instance in an EU region of your choice, using the NVIDIA GPU-Optimized VM image (Marketplace) or the Ubuntu HPC image with NVIDIA drivers.
  2. If the launch is refused, request GPU quota first: Subscription → Usage + quotas → request vCPUs for the instance family in your region (NVadsA10v5 for A10, NCadsA100v4 for A100).
  3. Attach a network security group that allows inbound SSH (port 22) only from your IP. Do not open the model port to the internet.
  4. Connect with SSH and run the one-line installer:
    curl -fsSL https://decisionmodels.io/local/install.sh | sh && export PATH="$HOME/.local/bin:$PATH"
  5. Install the model:
    dm-local install jeff-1-0-large
  6. Reach the endpoint through an SSH tunnel (ssh -L 8484:127.0.0.1:8484 user@instance) or from inside your private network, then call http://127.0.0.1:8484/v1/systemone.

Your data stays in your cloud account and region. Prices change; region, storage and traffic may add charges.

Google Cloud

Instance for this model: on request. Pick a GPU with at least 70 GB of memory.

  1. Open the Compute Engine console and launch a GPU instance in an EU region of your choice, using the Deep Learning VM (CUDA, Ubuntu).
  2. If the launch is refused, request GPU quota first: IAM & Admin → Quotas → GPUs (all regions) and the specific GPU type in your region.
  3. Attach a firewall rule that allows inbound SSH (port 22) only from your IP. Do not open the model port to the internet.
  4. Connect with SSH and run the one-line installer:
    curl -fsSL https://decisionmodels.io/local/install.sh | sh && export PATH="$HOME/.local/bin:$PATH"
  5. Install the model:
    dm-local install jeff-1-0-large
  6. Reach the endpoint through an SSH tunnel (ssh -L 8484:127.0.0.1:8484 user@instance) or from inside your private network, then call http://127.0.0.1:8484/v1/systemone.

Your data stays in your cloud account and region. Prices change; region, storage and traffic may add charges.

Other options: rent a GPU from RunPod or CoreWeave, or use any machine you control over SSH: dm-local remote user@host install jeff-1-0-large.

Fully sovereign

Fully sovereign — your data never leaves your machines

Prompts, answers and decisions are processed on hardware you control. The installer has no telemetry: it sends no prompts, no usage data and no analytics. The local service listens on loopback (127.0.0.1) unless you change it. After setup, the model runs in Hugging Face offline mode with vendor usage statistics switched off, so serving needs no internet connection.

What the installer contacts

  • GitHub — Installer releases and their signatures, the signature verifier (Sigstore cosign) if you do not have it, and pinned runtime binaries such as uv and llama.cpp. Recipe source code is fetched as an archive from GitHub.
  • Hugging Face — The model weights, at a pinned revision, checked against recorded hashes. A token is sent only if you set one for a gated model.
  • Python package indexes and container registries — Runtime packages (via uv) and container images, only for runtimes that need them, for example vLLM.
  • decisionmodels.io — The installer script download. For a paid licence, it also receives your licence key when you activate and about every 30 days. The key is the only thing sent.

Setup needs these connections once; the free tier never contacts decisionmodels.io for licensing. A bundled offline package for fully isolated (air-gapped) machines is not available yet — if you need one, ask us. Runtimes written by model authors remain subject to their own network behaviour.

Call your endpoint

Use a local typed endpoint

The API accepts typed questions and returns typed answers with probabilities.

Request

curl http://127.0.0.1:8484/v1/systemone -H 'Content-Type: application/json' -d '{"state":"Mia owns a red bicycle.","questions":{"color":{"type":"choice","instructions":"Which color is the bicycle?","criteria":{"red":null,"blue":null}}}}'

Example output

{"id":"dec_…","model":"jeff-1-0-large","answers":{"color":{"type":"choice","choice":"red","confidence":0.97,"probabilities":{"red":0.97,"blue":0.03}}},"usage":{"input_tokens":64,"output_tokens":0,"decisions":1}}

Python

import json
import urllib.request
payload = {"state": "Mia owns a red bicycle.",
           "questions": {"color": {"type": "choice", "instructions": "Which color is the bicycle?", "criteria": {"red": None, "blue": None}}}}
request = urllib.request.Request("http://127.0.0.1:8484/v1/systemone", data=json.dumps(payload).encode(), headers={"Content-Type": "application/json"})
with urllib.request.urlopen(request, timeout=30) as response:
    result = json.load(response)
print(result["answers"]["color"])
Read the hosted API reference →

Uninstall

Remove the local files

Use the installer to stop the service and remove this model's downloaded files.

dm-local uninstall jeff-1-0-large

Model terms

Licence details

cc-by-nc-4.0Commercial use: no

The model’s listed licence does not allow commercial use.

Model card

The local installer has a separate free and commercial licence.

Troubleshooting

Common setup issues

Driver or CUDA version is too old

Update to a driver/runtime combination listed for the selected variant, then run dm-local plan again.

Out of memory

Choose a smaller quantised variant when one is listed, or use a remote GPU with enough recommended memory.

NVIDIA Container Toolkit is missing

Install and configure the NVIDIA Container Toolkit for your host before retrying the container runtime.

The port is already in use

Stop the service using port 8484 or configure a different local port, then rerun the self-test.

A download was interrupted

Run the install command again. Completed downloads resume when the artifact server supports range requests.

Checksum verification failed

Do not use the file. Remove the incomplete download and retry from the pinned official source.

Apple Silicon memory pressure

Close other memory-heavy apps and choose a smaller supported variant if the catalogue lists one.

WSL2 cannot see the GPU

Check the Windows GPU driver and WSL2 GPU support, then verify nvidia-smi inside the WSL distribution.

SSH or firewall blocks the endpoint

Keep the service on loopback and use an SSH tunnel. Avoid exposing the API port directly to the public internet.

Jeff-1.0-Large local setup — Decision Models