# DecisionOne

> DecisionOne is an API from Sqwish Labs for small decision models called La Décision
> (Fat, Chub, Diet, Nibble and Crumb, largest first). `POST /v1/decide` takes some context and one
> or more questions, each with a fixed set of allowed answers, and returns a probability for
> every answer. You can fine-tune a model on labelled examples, promote a version to
> production and call it by name.

The OpenAPI file at [/openapi.json](/openapi.json) is the source of truth, so do not invent
fields. Every error has a stable code and a message that says what to change. The
interactive reference is at [/docs](/docs). Paths below are relative to your server's
address, which the examples call `$D1_API_URL`.

## Models

The model ids are `sqwish-d1-fat`, `sqwish-d1-chub`, `sqwish-d1-diet`,
`sqwish-d1-nibble` and `sqwish-d1-crumb`, largest first. `GET /v1/models` lists them all, plus
the models you have trained. A size you can call now has `"status": "ready"`. One whose
engines are all away has `"status": "unavailable"` and `"availability_reason": "no_engine"`,
and a request for it gets a retryable `engine_unavailable` or another size's answer. Fat
isn't served yet. Fine-tunes train on Diet. The research workbench runs Diet only. A multimodal model,
`sqwish-d1-multimodal`, is in early access; ask for it with
`POST /v1/early-access {"model": "sqwish-d1-multimodal", "working_on": "...", "why": "..."}`.

## Auth

Send `Authorization: Bearer d1_sk_...` with every `/v1` request. Create a key in the console
or with `POST /v1/account/keys {"name": "laptop"}`. The key is shown once. List keys with
`GET /v1/account/keys` and revoke one with `DELETE /v1/account/keys/{key_id}`. A local
research server started with `D1_AUTH_MODE=off` needs no key.

Some routes need no key. `POST /v1/playground/decide` takes the same body as
`/v1/decide`, runs base models only, allows up to 8 decisions, is rate-limited per address
and never stores anything. `POST /v1/lint` checks how questions are asked, without running a
model. `GET /v1/recipes` lists the ready-made requests, without their labelled examples.

## Run a decision

Each decision has an `id`, a `question` and its allowed outcomes. `kind` is `single` (one
named outcome, the default), `binary` (yes or no) or `ordinal` (levels `"0"`, `"1"` and so
on, with one `rubric` line per level). Ask related questions about the same context in one
request. Keep outcomes and JSON keys in their original order, because order changes what
the model reads.

```sh
curl -s "$D1_API_URL/v1/decide" \
  -H "Authorization: Bearer $D1_API_KEY" -H "Content-Type: application/json" \
  -d @- <<'EOF'
{
  "model": "sqwish-d1-diet",
  "context": "Our billing export has failed for two days. We need it for today's audit.",
  "decisions": [
    {"id": "team", "kind": "single", "question": "Which team should handle this ticket?",
     "outcomes": {"billing": "Charges, invoices, and financial exports.",
                  "technical": "Other application failures.", "sales": "New purchases."}},
    {"id": "urgent", "kind": "binary",
     "question": "Does the ticket state a deadline within one day?",
     "outcomes": {"yes": "An explicit deadline is today or within the next 24 hours.",
                  "no": "The deadline is later, absent, or unclear."}},
    {"id": "impact", "kind": "ordinal",
     "question": "How severe is the reported impact on the customer's work?",
     "outcomes": ["0", "1", "2"],
     "rubric": ["Minor inconvenience; the task can still be completed.",
                "The task is impaired, but a workable alternative is stated.",
                "An important task is blocked and no workable alternative is stated."]}
  ]
}
EOF
```

The Diet pilot (checkpoint 640) answered this request on 26 September 2026, under its
earlier id `d1-diet`, which still works as an alias. An excerpt, rounded to four places:

```json
{
  "id": "dec_963b7180cbd24e69a09870f8",
  "model": "sqwish-d1-diet",
  "decisions": {
    "team": {"kind": "single", "top": "billing", "p_top": 0.9268,
             "probabilities": {"billing": 0.9268, "technical": 0.0539, "sales": 0.0192}},
    "urgent": {"kind": "binary", "top": "yes", "p_yes": 0.9124,
               "probabilities": {"yes": 0.9124, "no": 0.0876}},
    "impact": {"kind": "ordinal", "top": "2", "expected_level": 1.7931,
               "probabilities": {"0": 0.0507, "1": 0.1056, "2": 0.8437}}
  }
}
```

Every decision also returns `margin`, `entropy`, `action` and `abstained`, and ordinal ones
add `median_level` and `legend`. `weights`, `costs` and `abstain` are optional rules applied
after the model runs. `probabilities` always stays the model's own distribution; with
`weights`, the reweighted one is in `policy_probabilities`. Probabilities are conditional on
the outcomes you declared. The base models' calibration is not measured yet, so do not read
`p_top` as an accuracy; a version you train reports its calibration error (`ece`) on its
held-out rows. `GET /v1/recipes` lists recipes, complete requests for common jobs, and
`GET /v1/decisionone` returns them with their labelled examples.

Before you ship new questions, send the same body to `POST /v1/lint`. It returns
`{"warnings": [{"decision": ..., "rule": ..., "severity": ..., "message": ...}]}` for
problems such as a negative question or a choice with no way out, and an empty list when it
finds none.

## Named models and versions

The `model` field accepts:

| `model` | What answers |
|---|---|
| `sqwish-d1-diet` | the base model |
| `router` | the production version of your model called `router` |
| `router@3` | version 3 |
| `router@candidate` | the newest version, if it is newer than production |
| `router@previous` | the version production replaced |
| `ft-ftj_...` | one trained adapter, by id |

When a named model answers, the response says which version ran:
`"model": "router", "version": "router@3"`. `GET /v1/models/router` lists every version with
its held-out results and gate, plus a summary of the rows the model remembers.

## When a model can't answer

If the model you asked for can't answer right now, for example because its adapter isn't
loaded or its engine is restarting, a base size that is ready answers instead, smartest
first: Chub, Diet, Nibble, then Crumb. The response then adds `"fallback": {"from": "router", "to": "sqwish-d1-diet", "reason": "adapter_not_loaded"}`.
A base model hasn't learned your labels, so check for `fallback` where that matters. Send
`"fallback": "none"` to get the error instead, for example when you evaluate one version, or
a list of up to four model ids to try in order. A wrong request never falls back.

## Fine-tune a named model

1. Create a dataset. `POST /v1/datasets {"name": "...", "examples": [...]}` takes JSON up to
   1 MiB. `POST /v1/datasets/upload` takes JSONL up to 16 MiB with
   `Content-Type: application/x-ndjson`. Each example is a decide request without `model`,
   plus `targets` (a probability for every outcome, summing to 1) or `rewards` (a reward for
   every outcome). Use at least two different contexts. Read the stored rows back with
   `GET /v1/datasets/{dataset_id}/rows`.
2. Train version 1: `POST /v1/fine-tuning/jobs {"dataset_id": "ds_...", "model_name": "router"}`.
   Leave out `hyperparameters.steps` for about one pass over the rows.
3. Poll `GET /v1/fine-tuning/jobs/{job_id}` until `status` is `succeeded`, `failed` or
   `cancelled`. A run takes minutes to hours.
4. Read `metrics.gate`. The new version is compared with the one it would replace on the
   same held-out decisions, split into remembered rows (older data) and new rows, so a big
   new batch cannot hide forgetting. When `passed` is false, `reasons` says why.
5. Train the next version on new data with `{"dataset_id": "ds_...", "parent": "router"}`.
   It starts from production's weights and mixes in `replay` remembered rows for each new row
   (default 1), to reduce forgetting.

Send an `Idempotency-Key` header when you create a dataset or a job, and send the same key if
you retry. A job that finishes does not go live on its own.

## Label real requests with teacher models

Teacher models can label requests you have no labels for: upload them as JSONL to
`POST /v1/labeling/sources`, start `POST /v1/labeling/jobs`, review
`GET /v1/labeling/jobs/{job_id}/results`, then send the rows you accept to
`POST /v1/labeling/jobs/{job_id}/finalize`, which makes a dataset for step 2 above.
`GET /v1/labeling/teachers` says whether this server has teachers set up.

## Tune the wording instead

With 60 or more labelled rows and no appetite for training, tune the words, not the weights.
`POST /v1/prompt-tuning/jobs {"dataset_id": "ds_...", "decision": "team"}` rewrites that
decision's question, outcome descriptions and rubric, and tests the result once on rows it
never saw. Outcome ids never change. Poll `GET /v1/prompt-tuning/jobs/{job_id}`. When
`result.verdict` is `improved`, use `result.tuned` as the decision; otherwise keep yours.
`result.label_issues` lists rows whose labels contradict the wording. A job that can't prove
an improvement on its first test of those rows is refunded; a rerun on the same rows is
stricter and charged. Base models only for now.

## Learn from production

1. Send `"store": true` with a decide request to your named model. The response `id` names
   the decision. Nothing is kept without that flag.
2. Once you know what really happened, send
   `POST /v1/feedback {"decision_id": "dec_...", "outcomes": {"team": "billing"}}`.
   Sending it again for the same decision corrects it.
3. `POST /v1/models/router/feedback-dataset` turns the new feedback into a dataset. Train
   the next version from it with `"parent": "router"`. Each piece of feedback goes into one
   dataset only. Send an `Idempotency-Key`, so a retry after a timeout returns the same
   dataset.
4. Or switch on keep learning, and the model does step 3 and the training by itself.
   `POST /v1/models/router/learning {"after": 200, "auto_promote": false}` trains the next
   version after 200 new labelled decisions (at least 150), at most once a day, and each
   round costs the same as a fine-tune. The new version waits as `@candidate`; with
   `"auto_promote": true` it is promoted if it passes its gate. `{"enabled": false}` turns
   it off. `GET /v1/models/router` shows progress under `learning`.

To train on one segment, send `"metadata": {"surface": "checkout"}` with the decisions and
the same `{"metadata": {...}}` to feedback-dataset. The model never sees metadata.

## Promote and roll back

- `POST /v1/models/router/promote {"version": "candidate"}` loads and checks that version,
  then makes it production. `version` can also be `"previous"` or a number. A version whose
  gate failed needs `"force": true`.
- `POST /v1/models/router/rollback` swaps production with the previous version. Calling it
  again swaps back.
- `POST /v1/models/router/keep-warm {"enabled": true}` keeps the production version loaded
  between requests.
- A version that hasn't been used for a while is unloaded to make room for others. The next
  decision waits up to 10 seconds for it to load, then gets a base size's answer marked
  `fallback` with reason `model_warming`, or, with `"fallback": "none"`, a retryable 503
  `model_warming` with `Retry-After`. The load carries on either way. Each version in
  `GET /v1/models/router` has a `state` (`loaded`, `loading` or `asleep`), and
  `POST /v1/models/router/wake` starts loading an asleep one and answers 202 at once.

## Moving from TypeSafe's Jev

An existing Jev request works unchanged at `POST /v1/systemone`. Point it at your DecisionOne
server and send a DecisionOne key. Jev's model ids (`jev-latest`, `jev-1.13.0`) run on
`sqwish-d1-diet`, and the answer's `model` field says so. A field DecisionOne doesn't model
is rejected with a 422 that names it, rather than ignored.

```sh
curl -s "$D1_API_URL/v1/systemone" -H "Authorization: Bearer $D1_API_KEY" \
  -H "Content-Type: application/json" -d '{"model": "sqwish-d1-diet",
  "state": "My parcel arrived broken.",
  "questions": {"refund": {"type": "noul", "instructions": "Is the customer asking for a refund?"}}}'
```

New code should use the native shape, which says the same thing in DecisionOne's words:

```sh
curl -s "$D1_API_URL/v1/decide" -H "Authorization: Bearer $D1_API_KEY" \
  -H "Content-Type: application/json" -d '{"model": "sqwish-d1-diet",
  "context": "My parcel arrived broken.",
  "decisions": [{"id": "refund", "kind": "binary", "question": "Is the customer asking for a refund?"}]}'
```

`state` becomes `context` and `questions` become `decisions`. `noul`, `choice` and `score`
become `binary`, `single` and `ordinal`, and `criteria` become outcome descriptions or an
ordinal `rubric`. `/v1/systemone` keeps Jev's `confidence` so thresholds tuned on Jev keep
working. `/v1/decide` returns every probability with `p_top`, `margin` and `entropy`, and
adds no confidence score of its own.

`POST /v1/systemone/translate` does the conversion for you. Send it a Jev request and it
returns the native one, with notes on anything that changed, including how a Jev confidence
threshold becomes a `p_top` threshold. It runs no model and needs no key.

## Errors

Every error has the same shape:
`{"error": {"code": "...", "message": "...", "retryable": false, "request_id": "...", "details": [...]}}`.
`retryable` says whether the same request can succeed later. Retry only when it is true,
after the `Retry-After` header when there is one, and with the same `Idempotency-Key`. A 503
can go either way: an engine blip is retryable, a size this server doesn't serve is not. On
a 422, `details` lists each field that failed with its path, so fix those fields rather than
sending the same body again. If a decide, promote or rollback request fails without a
response, check the model before you repeat it.

## Tools for coding agents

- An MCP server with tools to run and check decisions, tune their wording, record feedback,
  train, promote and roll back. Servers
  installed with the `mcp` extra serve it at `/mcp` (streamable HTTP, same API key as a
  bearer token; a browser session gets 401 `api_key_required` there), and
  `decisionone mcp` runs it over stdio. Setup for Claude Code, Codex and Cursor is in
  `docs/agents.md` in the workbench repository.
- Python SDK in `sdks/python`, used as `DecisionOne(url, api_key=...)`.
- TypeScript client in `sdks/typescript/decisionone.ts`, with no dependencies.
- Agent skill in `skills/decisionone/SKILL.md`.
- `POST /v1/hooks/claude-code` lets a La Décision model check Claude Code's tool calls before
  they run. Setup is in `docs/claude-code-hook.md`.

## Reference

Dataset, job and labelling lists come back newest first with `has_more` and `next_cursor`;
send that value as `cursor` to get the next page. Dataset rows and labelling results page by
`offset` and `limit` instead.

- [OpenAPI schema](/openapi.json)
- [Interactive reference](/docs)
- [ReDoc](/redoc)
