Pricing matrix

Software-license tier by inference tier. Rows: Self-hosted licence, Managed. Columns: SF default, BYOK, Sovereign.
Inference · SF default
metered, itemized
Inference · BYOK
your Anthropic key
Inference · Sovereign
your own GPU
Self-hosted
Flat software fee
Flat fee · metered inferenceLicensed software on your servers, your database. Inference billed per token through SF and itemised on the invoice. Fits operators already running their own stack who want the inference handled. Flat fee · your Anthropic billLicensed software on your servers. Your own key, so Anthropic bills you directly and SF takes no inference margin. Fits teams with an existing model contract. Flat fee · your own GPULicensed software and a local model, both on your hardware. Nothing leaves the building and there is no inference line at all. Fits the highest-privacy verticals.
Managed
Flat / workspace / mo
Flat fee · metered inferenceWe operate it on infrastructure you authorize, with SLAs. Inference billed per token through SF, itemized. Fits teams who want the thesis without running servers. Flat fee · your Anthropic billWe operate it; your own key handles inference. Fits managed operations with an existing Anthropic contract. Flat fee · your own GPUWe operate it against your own GPU and local model. Fits a sovereignty mandate that still wants managed operations.

Software fees and the SF inference margin are confirmed per engagement, because the honest number depends on scope. The per-token markup is stated before you sign, never after — see how inference is itemised.

Straight answers

Why the pricing looks like this.

Why no per-seat pricing

Per-seat pricing taxes you for growing. It indexes your bill to headcount instead of to the work the software does, and it charges you for the colleague who barely logs in. Our costs do not rise when you hire, so our price does not either.

There is a second reason now. When the agents are doing the work, a seat count has stopped describing anything real. You are not buying logins; you are buying a department that runs.

Why inference is itemized

Bundling inference into a plan hides what it actually costs and lets the margin drift. We would rather show you the wholesale rate, show you our margin, and let you decide whether to keep buying through us or bring your own key. A bill you can read is a bill you can leave.

Why the managed tier exists

Some teams genuinely should not run servers. The managed tier is for them: the same software, the same data-ownership posture, operated by us against infrastructure you authorize. It is a convenience you pay a flat fee for, not a lever to extract more as you grow.

What you are actually licensing

The software is proprietary and licensed to you. That is a change from an earlier posture, and worth stating plainly: you get the runtime, you run it on your own hardware, and it keeps running there whether or not you renew a retainer with us. What you do not get is the right to redistribute it.

What about having it built for you

If you want the agents built, deployed and operated rather than self-served from the docs, that is SF Foundry — a setup fee plus a monthly retainer, scoped to your operation.

SF Foundry

Get started

Talk through the software fee, the managed option, and which inference tier fits.

The software is free to run today. When you want the numbers for managed operation or metered inference, a call is the fastest path.