> ## Documentation Index
> Fetch the complete documentation index at: https://docs.orgo.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Use Any Model

> Build computer use agents with Claude, GPT, Gemini, or any model

Orgo exposes programmatic endpoints for every computer action: screenshot, click, type, key, scroll, bash. You bring the model, Orgo provides the computer.

Any model with computer use support works. Here is how to wire them up.

<Warning>
  The examples pass `ram` and `cpu`, and take screenshots and scroll over HTTP, so they also run on older SDKs. SDK versions before 0.0.47 (Python) and 1.0.22 (TypeScript) default to 2 GB, which `POST /computers` rejects with `400`. Their `screenshot()` and `screenshot_base64()` cannot read the stored-image path the API returns by default, and their `scroll()` takes no coordinates, so the computer scrolls at the top-left corner.
</Warning>

## Coordinate space

Orgo computers boot at `1280x720x24`.

Claude and OpenAI read the screen size from the screenshots you send. They return coordinates in that same pixel space. Send an Orgo screenshot unresized and the coordinates are already the computer's pixels, so you can click with them directly.

Gemini is different. It returns a normalized 0-999 grid, which you denormalize against the computer's real resolution:

```python theme={null}
x = int(fn_args["x"] / 1000 * 1280)
y = int(fn_args["y"] / 1000 * 720)
```

<Warning>
  Downscale a screenshot before sending it, or denormalize Gemini's grid against `1024x768` on a `1280x720` computer, and every click lands short of where the model aimed. Nothing errors. The agent looks like it is confidently clicking the wrong thing.
</Warning>

If you create a computer at a non-default resolution, or resize its screen later, use that resolution instead. `GET /computers/{id}/screens` reports the current `width` and `height`.

## Anthropic Claude

Claude drives computer use through the Messages API. `computer_toolset_20260801` declares every member tool in one entry and needs no beta header.

```python theme={null}
import os

import anthropic
import requests
from orgo import Computer

API = "https://www.orgo.ai/api"
HEADERS = {"Authorization": f"Bearer {os.environ['ORGO_API_KEY']}"}


def screenshot_base64(computer):
    """The screen as base64 PNG, straight from the API."""
    response = requests.get(
        f"{API}/computers/{computer.computer_id}/screenshot",
        params={"response_format": "base64"},
        headers=HEADERS,
    )
    response.raise_for_status()
    return response.json()["image"]


def scroll(computer, x, y, direction, amount):
    """Scroll at (x, y). Orgo scrolls up or down only."""
    if direction not in ("up", "down"):
        raise ValueError(f"Unsupported scroll direction: {direction}")
    response = requests.post(
        f"{API}/computers/{computer.computer_id}/scroll",
        json={"x": x, "y": y, "direction": direction, "amount": amount},
        headers=HEADERS,
    )
    response.raise_for_status()


# Pass a size: SDKs before 0.0.47 default to 2 GB, which the API rejects.
computer = Computer(ram=4, cpu=1)
client = anthropic.Anthropic()

messages = [{"role": "user", "content": "Open Chrome and search for AI news"}]

tools = [
    {"type": "computer_toolset_20260801"},
    {"type": "bash_20250124", "name": "bash"},
]

response = client.messages.create(
    model="claude-sonnet-5",
    max_tokens=8192,
    tools=tools,
    messages=messages,
)

# Claude can return several member calls per turn. Run them in order.
for block in response.content:
    if block.type != "tool_use":
        continue

    if block.name == "bash":
        output = computer.bash(block.input["command"])
    elif getattr(block, "toolset_name", None) == "computer":
        if block.name == "screenshot":
            image = screenshot_base64(computer)
            # Return as a tool_result with a base64 image
        elif block.name == "left_click":
            computer.left_click(*block.input["coordinate"])
        elif block.name == "type":
            computer.type(block.input["text"])
        elif block.name == "key":
            computer.key(block.input["text"])
        elif block.name == "scroll":
            scroll(computer, *block.input["coordinate"], block.input["scroll_direction"], block.input["scroll_amount"])

computer.destroy()
```

**Models:** `claude-opus-5-5`, `claude-sonnet-5-5`, `claude-opus-5`, `claude-sonnet-5`, `claude-opus-4-8`
**Tools:** `computer_toolset_20260801` + `bash_20250124`
**Beta header:** none
**Older models:** Opus 4.7 and Sonnet 4.6 use `computer_20251124` with the `computer-use-2025-11-24` beta
**Guide:** [Claude Computer Use](/guides/claude-computer-use)
**Docs:** [Anthropic Computer Use](https://platform.claude.com/docs/en/agents-and-tools/tool-use/computer-use-tool)

***

## OpenAI GPT

OpenAI's computer use works through the Responses API with a built-in `computer` tool. The tool takes no configuration, and each turn returns a batch of actions.

```python theme={null}
import os

import requests
from openai import OpenAI
from orgo import Computer

API = "https://www.orgo.ai/api"
HEADERS = {"Authorization": f"Bearer {os.environ['ORGO_API_KEY']}"}


def scroll(computer, x, y, direction, amount):
    """Scroll at (x, y). Orgo scrolls up or down only."""
    if direction not in ("up", "down"):
        raise ValueError(f"Unsupported scroll direction: {direction}")
    response = requests.post(
        f"{API}/computers/{computer.computer_id}/scroll",
        json={"x": x, "y": y, "direction": direction, "amount": amount},
        headers=HEADERS,
    )
    response.raise_for_status()


# Pass a size: SDKs before 0.0.47 default to 2 GB, which the API rejects.
computer = Computer(ram=4, cpu=1)
client = OpenAI()

response = client.responses.create(
    model="gpt-5.6",
    tools=[{"type": "computer"}],
    input="Open Chrome and search for AI news",
)

# Process computer_call actions, in order
for item in response.output:
    if item.type != "computer_call":
        continue
    for action in item.actions:
        if action.type == "click":
            computer.left_click(action.x, action.y)
        elif action.type == "double_click":
            computer.double_click(action.x, action.y)
        elif action.type == "type":
            computer.type(action.text)
        elif action.type == "keypress":
            # Orgo presses keys with xdotool, which expects X keysym names.
            # The OpenAI guide below maps the rest, such as the arrow keys.
            names = {"ENTER": "Return", "ESC": "Escape", "SPACE": "space", "TAB": "Tab", "BACKSPACE": "BackSpace"}
            computer.key("+".join(names.get(k.upper(), k.lower()) for k in action.keys))
        elif action.type == "scroll":
            scroll_y = getattr(action, "scroll_y", 0)
            scroll(computer, action.x, action.y, "down" if scroll_y > 0 else "up", max(1, abs(scroll_y) // 100))

computer.destroy()
```

**Models:** `gpt-5.6`, `gpt-5.5`
**Tool:** `computer`
**Older integration:** `computer-use-preview` with the `computer_use_preview` tool
**Guide:** [OpenAI Computer Use](/guides/openai-computer-use)
**Docs:** [OpenAI Computer Use](https://developers.openai.com/api/docs/guides/tools-computer-use)

***

## Google Gemini

Computer use is a built-in tool on Gemini 3.x, called through the Interactions API. Declare `environment: "desktop"` for an Orgo computer.

```python theme={null}
import base64
import io
import os

import requests
from google import genai
from orgo import Computer
from PIL import Image

# Pass a size: SDKs before 0.0.47 default to 2 GB, which the API rejects.
computer = Computer(ram=4, cpu=1)
client = genai.Client(api_key="YOUR_GEMINI_API_KEY")


def get_screenshot_png(computer: Computer) -> str:
    """Gemini accepts PNG only: fetch the screenshot as base64 and re-encode it as PNG."""
    response = requests.get(
        f"https://www.orgo.ai/api/computers/{computer.computer_id}/screenshot",
        params={"response_format": "base64"},
        headers={"Authorization": f"Bearer {os.environ['ORGO_API_KEY']}"},
    )
    response.raise_for_status()
    image = Image.open(io.BytesIO(base64.b64decode(response.json()["image"])))
    png_buffer = io.BytesIO()
    image.save(png_buffer, format="PNG")
    return base64.b64encode(png_buffer.getvalue()).decode("utf-8")


screenshot_png = get_screenshot_png(computer)

interaction = client.interactions.create(
    model="gemini-3.7-flash",
    input=[
        {"type": "text", "text": "Open Chrome and search for AI news"},
        {"type": "image", "data": screenshot_png, "mime_type": "image/png"},
    ],
    tools=[{
        "type": "computer_use",
        "environment": "desktop",
        "enable_prompt_injection_detection": True,
    }],
)

# Coordinates are normalized 0-999, so scale them to the screen
for step in interaction.steps:
    if step.type != "function_call":
        continue
    args = step.arguments
    if step.name == "click":
        computer.left_click(int(args["x"] / 1000 * 1280), int(args["y"] / 1000 * 720))
    elif step.name == "type":
        computer.type(args["text"])

computer.destroy()
```

**Model:** `gemini-3.7-flash`
**Tool:** `computer_use` with `environment` of `browser`, `desktop`, or `mobile`
**SDK:** `google-genai` 2.7.0 or later
**Coordinates:** normalized 0-999, scale to pixel resolution
**Legacy:** `gemini-2.5-computer-use-preview-10-2025` with `generate_content`
**Guide:** [Gemini Computer Use](/guides/gemini-computer-use)
**Docs:** [Gemini Computer Use](https://ai.google.dev/gemini-api/docs/computer-use)

***

## Other models

Any model that outputs structured actions (click, type, screenshot) can work with Orgo. The pattern is always the same:

1. **Create** a computer with `Computer(ram=4, cpu=1)`
2. **Screenshot** with `GET /computers/{id}/screenshot?response_format=base64`
3. **Send** the screenshot and instruction to your model
4. **Execute** the model's actions with `computer.left_click()`, `computer.type()`, and so on
5. **Loop** until the task is done

Open-source computer-use agents follow the same pattern. [Agent S3](/guides/agent-s2) is one worked example.

## Orgo action reference

| Action | Orgo call | Description |
| - | - | - |
| Screenshot | `GET /computers/{id}/screenshot?response_format=base64` | Capture screen as base64 PNG |
| Left click | `computer.left_click(x, y)` | Click at coordinates |
| Right click | `computer.right_click(x, y)` | Right-click |
| Double click | `computer.double_click(x, y)` | Double-click |
| Type | `computer.type("text")` | Type text |
| Key press | `computer.key("Return")` | Press key or combo (`ctrl+c`), using xdotool key names |
| Scroll | `POST /computers/{id}/scroll` with `x`, `y`, `direction`, `amount` | Scroll up or down at coordinates |
| Bash | `computer.bash("ls -la")` | Run terminal command |
| Wait | `computer.wait(2)` | Pause in seconds |

Full API details: [API Reference](/api-reference/introduction)


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.