> ## Documentation Index
> Fetch the complete documentation index at: https://docs.orgo.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Create chat completion

> OpenAI-compatible endpoint that lets an AI model drive a computer

Send a message to an AI model and have it drive an Orgo computer on your behalf. The model sees the screen, clicks, types, and runs shell commands until your request is satisfied, then returns the final assistant message. Use this endpoint as a drop-in replacement for `openai.chat.completions.create`.

<Info>
  The response is wire-compatible with OpenAI's chat completions. Any SDK that points at `https://www.orgo.ai/api/v1` works unchanged. The only required extension is passing a `computer_id` to bind the agent to a running computer.
</Info>

## Endpoint

```http theme={null}
POST https://www.orgo.ai/api/v1/chat/completions
```

**Auth:** `Authorization: Bearer $ORGO_API_KEY`

## Request

<ParamField body="computer_id" type="string" required>
  Identifier of a running Orgo computer: either its UUID (the `id` field from [Create computer](/api-reference/computers/create)) or its instance id. You must own the computer or be a member of its workspace with write access. A view-only member, or a caller with no access to the workspace, gets 401 `invalid_api_key`. Omitting it returns 400 `missing_computer_id`.
</ParamField>

<ParamField body="messages" type="array" required>
  Array of `{ role, content }` objects, where `content` is a plain string. Messages whose `role` is neither `user` nor `assistant` are dropped, and the first surviving message must have role `user`. An array with no usable message returns 400 `empty_messages`.

  The server does **not** load history from a thread, so a multi-turn conversation must resend the full transcript on every request. See [Threads and history](#threads-and-history).
</ParamField>

<ParamField body="model" type="string" default="claude-sonnet-5">
  Model identifier. Exactly one of `claude-sonnet-5`, `claude-opus-5.5`, `claude-opus-5`, `claude-opus-4.8`, `claude-sonnet-4.6`, or `claude-opus-4.6`. Anything else returns 400 `invalid_model`. When omitted, `claude-sonnet-5` is used. See [Models](#models) below.
</ParamField>

<ParamField body="stream" type="boolean" default="false">
  If `true`, responses are streamed as OpenAI-format Server-Sent Events. The connection stays open until the agent finishes or the client aborts the request. When omitted or `false`, the server buffers the whole agent run and returns a single JSON body.
</ParamField>

<ParamField body="thread_id" type="string">
  UUID of a thread to record this turn against. When omitted, a new thread is created and its ID is returned in `orgo.thread_id`. A `thread_id` that does not exist, or that another user created on a computer you can access, is ignored without error and a new thread is created instead. The value must be a UUID: anything else fails the request with a `500` that does not carry the error envelope.

  This selects where messages are **written**. It does not affect what the model sees. See [Threads and history](#threads-and-history).
</ParamField>

<Note>
  `max_steps` is not currently supported. The endpoint ignores any value you send, and the agent loop runs to a fixed internal ceiling of 250 steps.
</Note>

### Headers

<ParamField header="Authorization" type="string" required>
  `Bearer $ORGO_API_KEY`, your Orgo API key. Get one at [orgo.ai/settings/credentials](https://orgo.ai/settings/credentials).
</ParamField>

<ParamField header="x-anthropic-key" type="string">
  Bring-your-own Anthropic key. When present, inference bills directly to your Anthropic account, no Orgo credit hold is placed, and `orgo.cost_cents` is `0`. When absent, the request draws from the Orgo credit balance of the computer's workspace owner.
</ParamField>

## Threads and history

<Warning>
  `thread_id` does not do what its name suggests, and the difference will lose data if you assume otherwise.

  * **Prior history is not loaded.** The agent runs on exactly the `messages` you send in this request. Passing a `thread_id` with only the new user turn gives the model no context from earlier turns.
  * **Stored history is replaced, not appended.** When the run finishes, the thread's stored messages are overwritten with this request's `messages` plus the assistant replies produced in this turn. A second call that sends only the new turn destroys the first turn's transcript in the thread.

  To hold a multi-turn conversation, keep the transcript client-side and resend all of it on every request. Passing the same `thread_id` each time then keeps the stored copy complete.
</Warning>

## Response

### Non-streaming

<ResponseField name="id" type="string">
  Request identifier (`chatcmpl-` followed by a 16-character opaque id). Also returned as the `X-Request-Id` response header.
</ResponseField>

<ResponseField name="object" type="string">
  Always `chat.completion`.
</ResponseField>

<ResponseField name="created" type="integer">
  Unix timestamp (seconds).
</ResponseField>

<ResponseField name="model" type="string">
  The model ID you requested, echoed verbatim.
</ResponseField>

<ResponseField name="choices" type="array">
  Single-element array containing the final assistant message.

  <Expandable title="choice">
    <ResponseField name="index" type="integer">Always `0`.</ResponseField>
    <ResponseField name="message" type="object">`{ role: "assistant", content: string }`, holding the text of the last assistant message the agent loop produced. Empty string if the loop ended without assistant text.</ResponseField>
    <ResponseField name="finish_reason" type="string">Always `stop` on success.</ResponseField>
  </Expandable>
</ResponseField>

<ResponseField name="usage" type="object">
  Token counts for the full agent loop (all intermediate turns, not only the final message).

  <Expandable title="usage">
    <ResponseField name="prompt_tokens" type="integer">Total input tokens, excluding cache reads and cache writes.</ResponseField>
    <ResponseField name="completion_tokens" type="integer">Total output tokens.</ResponseField>
    <ResponseField name="total_tokens" type="integer">Sum of `prompt_tokens` and `completion_tokens`. Cache read and cache write tokens are billed but are not included in any of these counts.</ResponseField>
  </Expandable>
</ResponseField>

<ResponseField name="orgo" type="object">
  Orgo-specific metadata not in the OpenAI spec.

  <Expandable title="orgo">
    <ResponseField name="thread_id" type="string">ID of the thread this completion was written to. `null` if the thread could not be created.</ResponseField>
    <ResponseField name="steps" type="integer">Number of agent steps executed. Each step is one model call plus the tool calls it made.</ResponseField>
    <ResponseField name="cost_cents" type="number">Computed cost of this request in cents. `0` for BYOK. Enterprise plans see the computed cost here even though no credits are deducted.</ResponseField>
    <ResponseField name="credit_balance_cents" type="integer">Credit balance measured after the request's hold was placed and before settlement. Settlement runs immediately after the response is sent, so a later read of the balance will differ. Omitted for BYOK and enterprise plans.</ResponseField>
  </Expandable>
</ResponseField>

### Response headers

| Header | Description |
| - | - |
| `X-Request-Id` | Echoes the `id` field. Include it when reporting issues. |
| `X-Thread-Id` | Thread ID. Omitted when no thread could be created. |

## Streaming

When `stream: true`, the server returns `text/event-stream` with standard OpenAI `chat.completion.chunk` events:

```text theme={null}
data: {"id":"chatcmpl-VNJf55hQTfngVIm0","object":"chat.completion.chunk","created":1745136000,"model":"claude-sonnet-4.6","choices":[{"index":0,"delta":{"role":"assistant"},"finish_reason":null}]}

data: {"id":"chatcmpl-VNJf55hQTfngVIm0","object":"chat.completion.chunk","created":1745136000,"model":"claude-sonnet-4.6","choices":[{"index":0,"delta":{"content":"Opening Chrome"},"finish_reason":null}]}

data: {"id":"chatcmpl-VNJf55hQTfngVIm0","object":"chat.completion.chunk","created":1745136000,"model":"claude-sonnet-4.6","choices":[{"index":0,"delta":{},"finish_reason":"stop"}],"usage":{"prompt_tokens":1240,"completion_tokens":312,"total_tokens":1552}}

data: [DONE]
```

Only the model's text output is streamed. Tool calls, screenshots, and intermediate reasoning happen server-side and are not exposed as deltas. The `orgo` metadata block is not sent on this path. Read `X-Thread-Id` from the response headers instead, which are set before the first chunk arrives.

### Errors during a stream

The HTTP status is committed as 200 before the agent runs, so a mid-run failure cannot change it. Instead the stream emits a single error frame and terminates:

```text theme={null}
data: {"error":{"message":"…","type":"server_error"}}

data: [DONE]
```

Treat a stream that ends without a `finish_reason: "stop"` chunk as failed.

A run that fails mid-stream is billed for the tokens it used before failing, and its hold is refunded only when it used none, the same as the non-streaming path. A run that finishes is settled against actual usage even when the client has already disconnected.

## Examples

<CodeGroup>
  ```python Python (OpenAI SDK) theme={null}
  import os
  from openai import OpenAI

  client = OpenAI(
      base_url="https://www.orgo.ai/api/v1",
      api_key=os.environ["ORGO_API_KEY"],
  )

  response = client.chat.completions.create(
      model="claude-sonnet-4.6",
      messages=[
          {"role": "user", "content": "Open Chrome and search for 'claude computer use'"}
      ],
      extra_body={"computer_id": os.environ["COMPUTER_ID"]},
  )

  print(response.choices[0].message.content)
  print(f"Thread: {response.orgo['thread_id']}")
  ```

  ```typescript TypeScript (OpenAI SDK) theme={null}
  import OpenAI from "openai";

  const client = new OpenAI({
    baseURL: "https://www.orgo.ai/api/v1",
    apiKey: process.env.ORGO_API_KEY,
  });

  const response = await client.chat.completions.create({
    model: "claude-sonnet-4.6",
    messages: [
      { role: "user", content: "Open Chrome and search for 'claude computer use'" },
    ],
    // computer_id is an Orgo extension, so cast to bypass SDK types
    computer_id: process.env.COMPUTER_ID,
  } as any);

  console.log(response.choices[0].message.content);
  ```

  ```bash cURL theme={null}
  curl https://www.orgo.ai/api/v1/chat/completions \
    -H "Authorization: Bearer $ORGO_API_KEY" \
    -H "Content-Type: application/json" \
    -d '{
      "model": "claude-sonnet-4.6",
      "computer_id": "'"$COMPUTER_ID"'",
      "messages": [
        {"role": "user", "content": "Open Chrome and search for claude computer use"}
      ]
    }'
  ```
</CodeGroup>

### Streaming with raw SSE

<CodeGroup>
  ```python Python theme={null}
  import json, os, httpx

  with httpx.stream(
      "POST",
      "https://www.orgo.ai/api/v1/chat/completions",
      headers={"Authorization": f"Bearer {os.environ['ORGO_API_KEY']}"},
      json={
          "model": "claude-sonnet-4.6",
          "computer_id": os.environ["COMPUTER_ID"],
          "messages": [{"role": "user", "content": "Open github.com"}],
          "stream": True,
      },
      timeout=300,
  ) as r:
      for line in r.iter_lines():
          if not line.startswith("data: "):
              continue
          payload = line[6:]
          if payload == "[DONE]":
              break
          chunk = json.loads(payload)
          if "error" in chunk:
              raise RuntimeError(chunk["error"]["message"])
          delta = chunk["choices"][0].get("delta", {})
          if "content" in delta:
              print(delta["content"], end="", flush=True)
  ```

  ```bash cURL theme={null}
  curl -N https://www.orgo.ai/api/v1/chat/completions \
    -H "Authorization: Bearer $ORGO_API_KEY" \
    -H "Content-Type: application/json" \
    -d '{
      "model": "claude-sonnet-4.6",
      "computer_id": "'"$COMPUTER_ID"'",
      "messages": [{"role": "user", "content": "Open github.com"}],
      "stream": true
    }'
  ```
</CodeGroup>

### Continuing a conversation

Keep the transcript yourself and resend it in full. Pass the `thread_id` from the first response so both turns land in the same stored thread.

```python theme={null}
import os
from openai import OpenAI

client = OpenAI(base_url="https://www.orgo.ai/api/v1", api_key=os.environ["ORGO_API_KEY"])
computer_id = os.environ["COMPUTER_ID"]

history = [{"role": "user", "content": "Open Chrome and go to github.com"}]

first = client.chat.completions.create(
    model="claude-sonnet-4.6",
    messages=history,
    extra_body={"computer_id": computer_id},
)
thread_id = first.orgo["thread_id"]

# Append this turn's result, then append the next question.
history.append({"role": "assistant", "content": first.choices[0].message.content})
history.append({"role": "user", "content": "Search for 'orgo'"})

follow_up = client.chat.completions.create(
    model="claude-sonnet-4.6",
    messages=history,          # the FULL transcript, not only the new turn
    extra_body={
        "computer_id": computer_id,
        "thread_id": thread_id,
    },
)
```

### Bring your own Anthropic key

Pass `x-anthropic-key` to bill inference against your own Anthropic account. Orgo credits are not held or consumed and `orgo.cost_cents` is `0`.

```bash theme={null}
curl https://www.orgo.ai/api/v1/chat/completions \
  -H "Authorization: Bearer $ORGO_API_KEY" \
  -H "x-anthropic-key: $ANTHROPIC_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "claude-opus-4.6",
    "computer_id": "'"$COMPUTER_ID"'",
    "messages": [{"role": "user", "content": "Take a screenshot"}]
  }'
```

## Example response

```json theme={null}
{
  "id": "chatcmpl-VNJf55hQTfngVIm0",
  "object": "chat.completion",
  "created": 1745136000,
  "model": "claude-sonnet-4.6",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "Done. Chrome is open on google.com with the search results for 'claude computer use'."
      },
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 1240,
    "completion_tokens": 312,
    "total_tokens": 1552
  },
  "orgo": {
    "thread_id": "7c9e6679-7425-40de-944b-e07fc1f90ae7",
    "steps": 6,
    "cost_cents": 3.84,
    "credit_balance_cents": 9616
  }
}
```

## Models

| Model | Best for | Notes |
| - | - | - |
| `claude-sonnet-5` | Default. Fast and capable, and it handles most workflows. | Tool version `computer_20251124`. |
| `claude-opus-5.5` | The newest Opus, for complex multi-step tasks. | Tool version `computer_toolset_20260801`. |
| `claude-opus-5` | Previous Opus. | Tool version `computer_20251124`. |
| `claude-opus-4.8` | Prior-generation Opus. | Tool version `computer_20251124`. |
| `claude-sonnet-4.6` | Prior-generation Sonnet. | Tool version `computer_20251124`. |
| `claude-opus-4.6` | Prior-generation Opus. | Tool version `computer_20251124`. |

Per-model rates are on the [pricing page](https://orgo.ai/pricing).

<Note>
  Orgo's OpenAI-compatible endpoint uses dotted model IDs for point versions (`claude-sonnet-4.6`, `claude-opus-5.5`). If you call Anthropic's native SDK directly, as in the [Claude Computer Use guide](/guides/claude-computer-use), use the hyphenated form (`claude-sonnet-4-6`). That is Anthropic's canonical identifier.
</Note>

## Billing

Requests draw on the credit balance of the computer's workspace owner, not the caller. Orgo places a flat credit hold at request start and settles it against actual token usage once the agent loop finishes. Settlement is deferred until after the response is sent, so the balance reported in `orgo.credit_balance_cents` reflects the hold rather than the final cost.

The agent loop stops early once its running cost comes within 25 cents of the balance available at request start. When a workspace member's request spends the owner's balance, the balance that request may use is capped at \$5.

Holds are skipped entirely for BYOK requests and for enterprise plans. Rates are on the [pricing page](https://orgo.ai/pricing); the computed cost of a single request is returned in `orgo.cost_cents`.

## Errors

| Status | Code | When |
| - | - | - |
| 400 | `invalid_json` | Request body is not valid JSON. |
| 400 | `invalid_model` | `model` is not one of the supported identifiers (see [Models](#models)). The message lists the accepted values. |
| 400 | `missing_computer_id` | `computer_id` was not supplied. |
| 400 | `empty_messages` | `messages` contained no message with role `user` or `assistant`. |
| 400 | `invalid_message_order` | The first usable message does not have role `user`. |
| 400 | `context_overflow` | The conversation exceeds the model's context window. Prune the transcript you resend. Non-streaming only. |
| 401 | `invalid_api_key` | Every authentication or access failure: a missing or unknown key, a workspace-scoped key (this endpoint requires an account-wide key), a computer or `thread_id` in a workspace you cannot access, or view-only access to the computer's workspace. This surface returns one body for all of them, with `type` `authentication_error` and the message `Invalid API key. Get yours at https://orgo.ai/settings/credentials`. |
| 402 | `credits_exhausted` | Credit balance too low to place the hold. The body carries an extra `balance_cents` field. Add credits at [orgo.ai/settings/billing](https://orgo.ai/settings/billing). |
| 403 | `model_proxy_disabled` | The OpenAI-compatible endpoint is switched off for this account. Call your model provider directly and use the Orgo API for computer actions only. |
| 404 | `computer_not_found` | No computer matches `computer_id`, or it can no longer run: a free trial computer that has expired, or a computer paid for by its own subscription or dedicated purchase whose payment has lapsed. |
| 500 | `internal_error` | Unexpected server error. Retry with exponential backoff. Include the `X-Request-Id` header when reporting. |

Every error response has the shape:

```json theme={null}
{
  "error": {
    "type": "invalid_request",
    "message": "computer_id is required. Pass the ID of a running Orgo computer.",
    "code": "missing_computer_id"
  }
}
```

`type` is one of `invalid_request`, `authentication_error`, `insufficient_credits`, `permission_error`, `not_found`, or `server_error`.

Once a streaming response has started, failures arrive as an SSE error frame rather than an HTTP status. See [Errors during a stream](#errors-during-a-stream).


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.