Inference API
Inference routing, execution, streaming, and OpenAI-compatible gateway endpoints.
Inference
Section titled “Inference”The daemon exposes Signet’s inference control plane over both native RPC-style routes and an OpenAI-compatible gateway. Native inference routes are intended for first-party harnesses and CLI tooling. The OpenAI-compatible gateway is for harnesses that can point at a model endpoint but cannot yet send the richer Signet routing metadata.
GET /api/inference/status
Section titled “GET /api/inference/status”Requires diagnostics permission in authenticated modes.
Returns configured accounts, targets, policies, workload bindings, and the current runtime snapshot for each route target.
Response
{ "enabled": true, "source": "explicit", "defaultPolicy": "auto", "defaultAgentId": "default", "policies": ["auto", "strict-coding"], "taskClasses": ["casual_chat", "hard_coding", "hipaa_sensitive"], "targetRefs": ["sonnet/default", "gpt/gpt54", "local/gemma4"], "workloadBindings": { "interactive": "auto", "memoryExtraction": "memory-pipeline", "aggregateRecall": "aggregation/default" }, "configIssues": [ { "severity": "warning", "field": "policies.auto.defaultTargets", "ref": "ghost/default", "message": "Target ref \"ghost/default\" referenced by policies.auto.defaultTargets does not exist." } ], "runtimeSnapshot": { "targets": { "sonnet/default": { "available": true, "health": "healthy", "circuitOpen": false, "accountState": "ready" } } }, "concurrency": { "active": { "execute": 0, "nativeStream": 1, "gatewayStream": 0, "total": 1 }, "limits": { "execute": 8, "nativeStream": 8, "gatewayStream": 16, "total": 24 } }}Inference concurrency limits can be tuned with:
SIGNET_INFERENCE_MAX_CONCURRENT_EXECUTESIGNET_INFERENCE_MAX_CONCURRENT_NATIVE_STREAMSSIGNET_INFERENCE_MAX_CONCURRENT_GATEWAY_STREAMSSIGNET_INFERENCE_MAX_CONCURRENT_TOTAL
configIssues lists broken cross-references discovered when the config was
loaded (#1005). Most entries are severity: "warning" — the config loads but
the reference degrades routing (dangling policy target lists, agent rosters,
target accounts, and stale workload policy/target pins, which the router
falls back around). A severity: "error" makes routing structurally
non-functional (only a broken defaultPolicy) — the endpoint then returns
HTTP 400 with the structured error instead of a status body:
{ "error": "Routing config has 1 broken reference(s): defaultPolicy=\"background-acpx\"", "details": { "issues": [ { "severity": "error", "field": "defaultPolicy", "ref": "background-acpx", "message": "Policy \"background-acpx\" referenced by defaultPolicy does not exist." } ], "warnings": [] }}The same references are logged (at boot and whenever the config file changes) before any route is attempted.
GET /api/inference/catalog
Section titled “GET /api/inference/catalog”Requires diagnostics permission in authenticated modes. Returns the provider
and model registry from pi-ai, the OAuth provider registry, and ACPX agent
names. modelErrors reports provider-specific catalog failures; failures are
never represented as a silent empty model list.
GET /api/inference/oauth/providers
Section titled “GET /api/inference/oauth/providers”Requires diagnostics permission in authenticated modes. Returns the OAuth
providers registered by pi-ai and whether Signet has encrypted credentials for
each provider. Access and refresh tokens are never returned.
POST /api/inference/oauth/login/:id
Section titled “POST /api/inference/oauth/login/:id”Requires admin permission. Starts a bounded, daemon-owned OAuth login and
returns a Server-Sent Events stream. The response’s
X-Signet-OAuth-Session-Id header identifies the login session.
Events include session, auth, device_code, prompt, select, manual_code,
progress, connected, error, and done. Interactive events include a
responseId that must be answered through the completion route. The daemon
stores credentials returned by pi-ai; clients never submit or receive OAuth
credentials.
POST /api/inference/oauth/complete
Section titled “POST /api/inference/oauth/complete”Requires admin permission. Answers one pending login interaction.
{ "sessionId": "a login session UUID", "responseId": "the event response UUID", "value": "the prompt response, selected option id, or manual callback URL"}Sessions and responses are single-use, expire after ten minutes, and reject empty values unless the provider explicitly allows them.
POST /api/inference/oauth/disconnect/:id
Section titled “POST /api/inference/oauth/disconnect/:id”Requires admin permission. Deletes the provider’s encrypted OAuth credential.
Subsequent inference requests treat the account as disconnected.
GET /api/inference/history
Section titled “GET /api/inference/history”Requires diagnostics permission in authenticated modes.
Returns recent local inference telemetry in a redacted, operator-friendly shape. The endpoint only returns events when telemetry is enabled.
Query parameters
| Parameter | Type | Description |
|---|---|---|
limit |
integer | Max events, default 50, max 500 |
since |
string | ISO timestamp lower bound |
until |
string | ISO timestamp upper bound |
event |
string | One inference event type to include |
failures |
1 |
Include only failed, cancelled, or fallback rows |
Response
{ "enabled": true, "events": [ { "event": "inference.fallback", "timestamp": "2026-04-10T18:12:00.000Z", "surface": "native", "agentId": "rose", "operation": "interactive", "taskClass": "interactive", "policyId": "auto", "selectedTarget": "primary/fast", "finalTarget": "backup/safe", "attemptPath": "primary/fast -> secondary/deep -> backup/safe", "failedTargets": "primary/fast,secondary/deep", "fallbackCount": 2, "errorCode": "RATE_LIMITED" } ], "summary": { "total": 1, "failures": 1, "fallbacks": 1, "cancelled": 0 }}Inference history excludes raw prompts, response text, credentials, and session references.
POST /api/inference/explain
Section titled “POST /api/inference/explain”Requires admin permission in authenticated modes.
Dry-runs a route decision without executing the request. This is the backend
used by signet route explain.
Request body
{ "agentId": "rose", "operation": "interactive", "taskClass": "hard_coding", "privacy": "restricted_remote", "promptPreview": "fix this failing bun test", "refresh": true}Boundary guards:
explicitTargetsmay contain at most 8 entries.expectedInputTokens,expectedOutputTokens, andlatencyBudgetMsare clamped to sane non-negative bounds.promptPreviewis truncated to 4,000 characters.- Oversized request bodies return
413.
Response
Returns a full RouteDecision object, including trace.candidates[] with the
ordered scoring and policy gates applied to each target.
POST /api/inference/execute
Section titled “POST /api/inference/execute”Requires admin permission in authenticated modes.
Routes and executes a prompt using the Signet inference layer.
Request body
{ "agentId": "miles", "operation": "code_reasoning", "taskClass": "hard_coding", "prompt": "explain why this Rust borrow checker error happens", "maxTokens": 1200}Boundary guards:
- Request bodies over 512 KiB return
413. promptis capped at 200,000 characters.maxTokensandtimeoutMsare clamped to sane non-negative bounds.explicitTargetsmay contain at most 8 entries.
Response
{ "text": "The borrow error happens because ...", "usage": { "inputTokens": 322, "outputTokens": 471, "cacheReadTokens": null, "cacheCreationTokens": null, "totalCost": null, "totalDurationMs": null }, "decision": { "policyId": "auto", "mode": "automatic", "taskClass": "hard_coding", "targetRef": "gpt/gpt54" }, "attempts": [ { "targetRef": "gpt/gpt54", "ok": true, "durationMs": 1840 } ]}POST /api/inference/stream
Section titled “POST /api/inference/stream”Requires admin permission in authenticated modes.
Streams a routed inference request over Server-Sent Events for first-party
Signet consumers. The request body matches POST /api/inference/execute.
SSE events
meta— includesrequestIdand the selected route decisiondelta— streamed text chunksdone— final text, usage, and attempt metadatacancelled— emitted when the stream is cancelled explicitly or by disconnecterror— emitted when a provider dies mid-stream, including partial text
The response also includes x-signet-request-id, which can be used with the
cancellation endpoint below.
DELETE /api/inference/requests/:id
Section titled “DELETE /api/inference/requests/:id”Requires admin permission in authenticated modes.
Cancels an active native or gateway inference stream by request id.
GET /v1/models
Section titled “GET /v1/models”Requires admin permission in authenticated modes.
OpenAI-compatible model listing for the Signet gateway. Returned IDs include:
signet:autopolicy:<policy-id>- explicit target refs like
gpt/gpt54
POST /v1/chat/completions
Section titled “POST /v1/chat/completions”Requires admin permission in authenticated modes.
OpenAI-compatible chat completion endpoint. Signet routes the request before
execution. The request body accepts standard OpenAI-style model, messages,
and max_tokens fields.
Signet-specific routing hints can be provided in headers:
x-signet-agent-idx-signet-task-classx-signet-privacy-tierx-signet-operationx-signet-route-policyx-signet-explicit-target
Boundary guards:
- Request bodies over 512 KiB return
413. messagesmay contain at most 128 entries and 200,000 total characters of string content.- Signet routing headers are normalized and invalid hint values return
400.
When stream: true, the gateway returns OpenAI-style SSE chunks and includes
x-signet-request-id in the response headers so operators can cancel the
stream through DELETE /api/inference/requests/:id.
GET /api/assistant/models
Section titled “GET /api/assistant/models”Requires recall permission and a resolved agent scope; accepts optional agentId.
Returns { models: [{ targetRef, model, name, provider, account }] } from the Pi AI
registry for text-capable models associated with configured targets whose Signet
account credentials resolve successfully. Credentials and secret references are
never returned. ACPX targets are excluded from this Pi model picker.
POST /api/assistant/chat
Section titled “POST /api/assistant/chat”Requires recall permission and a resolved agent scope. Send UUIDs requestId (one per turn) and conversationId (stable across turns),
optional agentId, selectedEntityId, and modelSelection: { targetRef, model }, and messages containing user or
assistant roles with string content. The final message must be from the user.
History is bounded to 32 messages, 16,000 characters per message, and 64,000
characters total; the request body is bounded to 96 KiB.
Omit modelSelection to use the existing interactive/backend assignment. An explicit
selection must match the credential-backed catalog at execution time. It applies
only to this request, preserves the configured provider/account and route eligibility,
and does not update durable model assignments. Unknown or disconnected selections
fail explicitly; they do not fall back to another model.
The response is an SSE stream of JSON data events: delta, tool, citation,
retrieval (nodeIds and evidenceRefs, at most 100 identifiers each), focus, saved, dream, done, or error. done ends a successful turn;
error ends a failed turn. Disconnecting cancels the current turn and awaits its
cleanup while retaining the idle session. Output is bounded to 1 MiB and agent execution to 90 seconds.
The daemon uses the shared Pi agent worker boundary also used by Dreaming. At most four Pi workers exist concurrently; at most three are retained chat sessions. Additional sessions fail explicitly when capacity is reached. Conversations are scoped to the authenticated subject, resolved agent, and conversation UUID. One turn may run per conversation. The daemon retains each chat’s native Pi session and tool history for 15 minutes of inactivity after its last settled turn. Continuations append only the new user message; request history seeds a session when it is first created or has expired. Switching models updates the existing Pi session. Each turn has a 64-call tool budget. Shutdown disposes all workers. Dreaming sessions remain bounded to their pass lifecycle. The worker owns model execution and the agent loop; tools execute through existing daemon capabilities and the asynchronous database owner protocol. No shell, coding, ambient extension, or direct semantic mutation tools are provided to chat.
Chat reads memory through scoped daemon capabilities: the ontology readers,
search_evidence over the full history of episodic memories, artifacts, and
transcripts, and recall_memories, which calls POST /api/memory/recall with the
resolved agent and recall surface dashboard. Recall covers memories curated by
Dreaming as well as captured ones. Its output passes the memory content-safety
projection before reaching the model; withheld or malformed rows are dropped.
A recalled memory that resolves as episodic evidence through search_evidence
carries a memory:<id> sourceRef and emits citation and retrieval events.
Dreaming-curated memories, ontology claims, and source rows carry only a
recallId and must be backed by evidence tools before they are cited.
The system prompt requires evidence citations as [[kind:exact-sourceRef-id]]
wikilinks copied verbatim from retrieval results. The dashboard renders only
references backed by retrieved citation events as clickable source pills; entity
IDs and invented references do not become evidence links.
remember_user_message can save only the exact current user message through
/api/memory/remember, with a request-based idempotency key. request_dreaming
saves that message first and invokes /api/dream/trigger using its evidence
reference. The original caller’s permissions apply to both operations. saved
means evidence was recorded; dream means a pass was accepted, not completed.
Neither action is rolled back when inference is cancelled. Chat history itself
is not persisted by this endpoint.

