| Requirement | Minimum | Recommended |
|---|---|---|
| Precision | bf16 | bf16 |
| GPU memory (VRAM) | 70 GB | 80 GB |
| Example GPUs | — | — |
| System RAM | 80 GB | 80 GB |
| Disk | 61.43 GB | 61.43 GB |
Run locally / Model details
Jeff-1.0-Large
31B text model.
Fully sovereign. Maximum data privacy: it runs on your hardware or in your own cloud account, and your data never leaves your machines.
Installer licence
Free for small teams. USD 1,000 + USD 100/month for larger companies.
Free for individuals and companies with up to 10 employees and up to USD 1M ARR.
Larger companies: USD 1,000 one-time + USD 100/month (self-declared). That applies above either threshold: more than 10 employees or more than USD 1 million ARR.
This is the licence for the Decision Models installer. The model's own weights licence is separate and shown further down.
This installer entry is not cleared for commercial use. Read the model card before downloading or serving it.
Hardware
Minimum and recommended hardware
31B parameters. Sized for decision readouts with inputs up to about 8k tokens.
- CPU only
- No The installer catalogue has no CPU variant.
- Apple Silicon
- Not yet The installer catalogue has no Apple Silicon variant.
How we computed this: weights size × bytes per value for the precision, plus 10 % and room for the context. For mixture-of-experts models all parameters must be in memory. Rounded up to common GPU sizes. Installer figures come from the reviewed catalogue.
Variants
Will it run on my machine?
Requirements are recorded per variant. A dash means the catalogue does not provide that value.
| Variant | Precision | Runtime | Benchmarked | Install recipe | Minimum / recommended VRAM | RAM | Disk | Platforms | Expected speed |
|---|---|---|---|---|---|---|---|---|---|
| bf16-jeff-serve | bf16 | Transformers | Yes | Tested install recipe | 70 GB / 80 GB | 80 GB | 61.43 GB | linux-nvidia, wsl2-nvidia | About 0.3 s per decision on the measured GPU (board latency 0.31 s). Not measured elsewhere. |
- Runtime
- Transformers
- Install recipe
- Tested install recipe
- VRAM
- 70 GB minimum / 80 GB recommended
- RAM
- 80 GB
- Disk
- 61.43 GB
- Platforms
- linux-nvidia, wsl2-nvidia
- Expected speed
- About 0.3 s per decision on the measured GPU (board latency 0.31 s). Not measured elsewhere.
dm-local plan jeff-1-0-largeChecks memory, disk, platform, and runtime against the catalogue.Quick start
Choose where to run it
Installer commands use the pinned model revision. Confirm the model terms before proceeding.
An install recipe has been tested. Platform and hardware coverage is shown per variant.
- 1
curl -fsSL https://decisionmodels.io/local/install.sh | sh && export PATH="$HOME/.local/bin:$PATH" - 2
dm-local install jeff-1-0-large
Your cloud account
Run it on AWS, Azure or Google Cloud
Launch a GPU machine in your own account, run the model there, and keep everything inside your cloud account and the region you pick.
AWS
Instance for this model: on request. Pick a GPU with at least 70 GB of memory.
- Open the EC2 console and launch a GPU instance in an EU region of your choice, using the AWS Deep Learning AMI (GPU, Ubuntu).
- If the launch is refused, request GPU quota first: Service Quotas → EC2 → “Running On-Demand G and VT instances” (or P instances).
- Attach a security group that allows inbound SSH (port 22) only from your IP. Do not open the model port to the internet.
- Connect with SSH and run the one-line installer:
curl -fsSL https://decisionmodels.io/local/install.sh | sh && export PATH="$HOME/.local/bin:$PATH" - Install the model:
dm-local install jeff-1-0-large - Reach the endpoint through an SSH tunnel (
ssh -L 8484:127.0.0.1:8484 user@instance) or from inside your private network, then callhttp://127.0.0.1:8484/v1/systemone.
Your data stays in your cloud account and region. Prices change; region, storage and traffic may add charges.
Azure
Instance for this model: on request. Pick a GPU with at least 70 GB of memory.
- Open the Virtual machines blade and launch a GPU instance in an EU region of your choice, using the NVIDIA GPU-Optimized VM image (Marketplace) or the Ubuntu HPC image with NVIDIA drivers.
- If the launch is refused, request GPU quota first: Subscription → Usage + quotas → request vCPUs for the instance family in your region (NVadsA10v5 for A10, NCadsA100v4 for A100).
- Attach a network security group that allows inbound SSH (port 22) only from your IP. Do not open the model port to the internet.
- Connect with SSH and run the one-line installer:
curl -fsSL https://decisionmodels.io/local/install.sh | sh && export PATH="$HOME/.local/bin:$PATH" - Install the model:
dm-local install jeff-1-0-large - Reach the endpoint through an SSH tunnel (
ssh -L 8484:127.0.0.1:8484 user@instance) or from inside your private network, then callhttp://127.0.0.1:8484/v1/systemone.
Your data stays in your cloud account and region. Prices change; region, storage and traffic may add charges.
Google Cloud
Instance for this model: on request. Pick a GPU with at least 70 GB of memory.
- Open the Compute Engine console and launch a GPU instance in an EU region of your choice, using the Deep Learning VM (CUDA, Ubuntu).
- If the launch is refused, request GPU quota first: IAM & Admin → Quotas → GPUs (all regions) and the specific GPU type in your region.
- Attach a firewall rule that allows inbound SSH (port 22) only from your IP. Do not open the model port to the internet.
- Connect with SSH and run the one-line installer:
curl -fsSL https://decisionmodels.io/local/install.sh | sh && export PATH="$HOME/.local/bin:$PATH" - Install the model:
dm-local install jeff-1-0-large - Reach the endpoint through an SSH tunnel (
ssh -L 8484:127.0.0.1:8484 user@instance) or from inside your private network, then callhttp://127.0.0.1:8484/v1/systemone.
Your data stays in your cloud account and region. Prices change; region, storage and traffic may add charges.
Other options: rent a GPU from RunPod or CoreWeave, or use any machine you control over SSH: dm-local remote user@host install jeff-1-0-large.
Fully sovereign
Fully sovereign — your data never leaves your machines
Prompts, answers and decisions are processed on hardware you control. The installer has no telemetry: it sends no prompts, no usage data and no analytics. The local service listens on loopback (127.0.0.1) unless you change it. After setup, the model runs in Hugging Face offline mode with vendor usage statistics switched off, so serving needs no internet connection.
What the installer contacts
- GitHub — Installer releases and their signatures, the signature verifier (Sigstore cosign) if you do not have it, and pinned runtime binaries such as uv and llama.cpp. Recipe source code is fetched as an archive from GitHub.
- Hugging Face — The model weights, at a pinned revision, checked against recorded hashes. A token is sent only if you set one for a gated model.
- Python package indexes and container registries — Runtime packages (via uv) and container images, only for runtimes that need them, for example vLLM.
- decisionmodels.io — The installer script download. For a paid licence, it also receives your licence key when you activate and about every 30 days. The key is the only thing sent.
Setup needs these connections once; the free tier never contacts decisionmodels.io for licensing. A bundled offline package for fully isolated (air-gapped) machines is not available yet — if you need one, ask us. Runtimes written by model authors remain subject to their own network behaviour.
Call your endpoint
Use a local typed endpoint
The API accepts typed questions and returns typed answers with probabilities.
Request
curl http://127.0.0.1:8484/v1/systemone -H 'Content-Type: application/json' -d '{"state":"Mia owns a red bicycle.","questions":{"color":{"type":"choice","instructions":"Which color is the bicycle?","criteria":{"red":null,"blue":null}}}}'Example output
{"id":"dec_…","model":"jeff-1-0-large","answers":{"color":{"type":"choice","choice":"red","confidence":0.97,"probabilities":{"red":0.97,"blue":0.03}}},"usage":{"input_tokens":64,"output_tokens":0,"decisions":1}}Python
import json
import urllib.request
payload = {"state": "Mia owns a red bicycle.",
"questions": {"color": {"type": "choice", "instructions": "Which color is the bicycle?", "criteria": {"red": None, "blue": None}}}}
request = urllib.request.Request("http://127.0.0.1:8484/v1/systemone", data=json.dumps(payload).encode(), headers={"Content-Type": "application/json"})
with urllib.request.urlopen(request, timeout=30) as response:
result = json.load(response)
print(result["answers"]["color"])Read the hosted API reference →Uninstall
Remove the local files
Use the installer to stop the service and remove this model's downloaded files.
dm-local uninstall jeff-1-0-largeModel terms
Licence details
The model’s listed licence does not allow commercial use.
The local installer has a separate free and commercial licence.
Troubleshooting
Common setup issues
Driver or CUDA version is too old
Update to a driver/runtime combination listed for the selected variant, then run dm-local plan again.
Out of memory
Choose a smaller quantised variant when one is listed, or use a remote GPU with enough recommended memory.
NVIDIA Container Toolkit is missing
Install and configure the NVIDIA Container Toolkit for your host before retrying the container runtime.
The port is already in use
Stop the service using port 8484 or configure a different local port, then rerun the self-test.
A download was interrupted
Run the install command again. Completed downloads resume when the artifact server supports range requests.
Checksum verification failed
Do not use the file. Remove the incomplete download and retry from the pinned official source.
Apple Silicon memory pressure
Close other memory-heavy apps and choose a smaller supported variant if the catalogue lists one.
WSL2 cannot see the GPU
Check the Windows GPU driver and WSL2 GPU support, then verify nvidia-smi inside the WSL distribution.
SSH or firewall blocks the endpoint
Keep the service on loopback and use an SSH tunnel. Avoid exposing the API port directly to the public internet.