Open-weights decision model

Mixedbread mxbai-rerank-base-v2

Run it privately from $0.48/hour, call it as an API, run it locally, or get it pre-installed.

Dedicated hosting

Your own private endpoint

Best for steady traffic and privacy.

We rent the right GPU, deploy the model and hand you an OpenAI-compatible endpoint with your own API key. Nobody else shares it. Stop or delete it any time.

  1. Pick EU or Global. We size the GPU for you.
  2. Confirm the price — paid from your prepaid balance.
  3. Your endpoint goes live in about 5–12 minutes. Call it with any OpenAI SDK.
  • Pinned to the exact published weights (3ea9d4d)
  • Your own API keys — create and revoke any time
  • Live status, server log, requests and spend in one dashboard
  • We don’t store your prompts or outputs

Pick a region → confirm the price → call it with any OpenAI SDK. Live in about 5–12 minutes.

curl https://hosting.decisionmodels.io/v1/chat/completions \ -H "Authorization: Bearer $DM_API_KEY" \ -H "Content-Type: application/json" \ -d '{"model":"mxbai-rerank-base-v2","messages":[{"role":"user","content":"…"}]}'

Dedicated hosting is in private preview.

Where should it run?

Lowest price. GPUs from independent and data-centre providers worldwide. Choose EU for personal or regulated data.

How should it run?

Warm 24/7. Billed per hour, prepaid.

Auto-sized
— / hour

Paid from your prepaid balance — VAT is applied when you top up.

Stops when your balance runs out — never charged beyond it.

Just trying it? costs nothing while idle.

Advanced: GPU, idle timeout
GPU

Shared servers: join the waitlist.

Deploy

Full refund if we can’t get it running. First start takes about 5–12 minutes.

Prepaid · no subscription · stop any time

Hosted API

Best for trying it or low volume.

On request

Want Mixedbread mxbai-rerank-base-v2 as pay-per-token? Tell us. We add models based on demand.

We’ll email you once when it’s live. Nothing else.

Run locally

Best for testing offline.

One installer checks your machine (or a cloud/SSH GPU node), downloads Mixedbread mxbai-rerank-base-v2 and starts a local Jev-compatible (typed answers with probabilities) endpoint. Plan for about 5 GB of GPU memory.

Get the installer The installer is free for individuals and small companies. The model’s licence still applies (Apache 2.0).

Pre-installed hardware

Best for data that must stay on site.

Ship Mixedbread mxbai-rerank-base-v2 on a box you own. Tell us your setup and we quote within one business day.

Ask for a quote We reply within one business day.

Model facts

Model
Mixedbread mxbai-rerank-base-v2
Author
Mixedbread
Base
Qwen2.5 (size not stated)
Weights
mixedbread-ai/mxbai-rerank-base-v2
Licence
Apache 2.0
JevBench
#128 on the open-weights board · capability 43.9 · details

Benchmark data from Benchmark Heaven. Model names belong to their authors; we host the published weights.

Dedicated hosting — good to know

What exactly do I get?

A private, OpenAI-compatible endpoint that serves only Mixedbread mxbai-rerank-base-v2 on GPUs reserved for you, with API keys you control, a live status page, request counts and spend. The model is pinned to an exact published revision.

How is it billed?

From your prepaid Decision Models balance, in blocks: one hour at a time for “always on”, 15 minutes at a time while warm for “scale to zero”. VAT is applied when you top up. The next block is charged just before the current one ends. When your balance can’t cover the next block, the endpoint stops — you are never billed beyond your balance. Optional daily and monthly caps.

What does “scale to zero” mean?

After a quiet period (15 minutes by default) we stop the GPU. The next request wakes it; until it is ready the endpoint answers 503 with a Retry-After header. A cold start takes a few minutes: a fresh GPU host loads the server and the weights (the estimate for your choice is shown above). Great for batch jobs and prototypes; pick “always on” for live traffic.

EU or Global?

EU — hosting in EU, built for GDPR, aligned with the EU AI Act. Runs on GPUs of an EU provider in EU data centres. Global — lowest price. GPUs from independent and data-centre providers worldwide. Choose EU for personal or regulated data.

Do you store my prompts?

No. Requests pass straight through to your model server; we keep only counts (requests, tokens, errors) for your dashboard. Your prompts are processed on the GPU that runs your endpoint — on Global that GPU belongs to an independent or data-centre provider, so choose EU for personal or regulated data.