> ## Documentation Index
> Fetch the complete documentation index at: https://docs.orgo.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# OpenAI Computer Use

> Control an Orgo computer with GPT

OpenAI's computer tool drives interfaces through the Responses API. This guide wires it to an Orgo computer.

<Info>
  The `computer` tool is the generally available integration and runs on the GPT-5.5 and GPT-5.6 families. The older `computer_use_preview` tool and its `computer-use-preview` model still work, and are covered at the end of this page. Use the GA tool for anything new.
</Info>

## Coordinate space

Orgo computers boot at `1280x720x24`.

The `computer` tool takes no display size. The model reads the screen dimensions from the screenshots you send, and returns `x` and `y` in that same pixel space. Send an Orgo screenshot unresized and the coordinates come back as the computer's pixels. Pass them straight to `left_click`.

Send `"detail": "original"` with each screenshot so the API keeps the resolution the model is measuring against.

<Warning>
  Downscale a screenshot and click the raw coordinates and every click lands short of where the model aimed. Nothing errors. The agent looks like it is confidently clicking the wrong thing. If you must resize, keep the scale factor and map each coordinate back to the original resolution.
</Warning>

`GET /computers/{id}/screens` reports the current `width` and `height` if you created the computer at another resolution or resized its screen later.

## Quick start

<Warning>
  The examples pass `ram` and `cpu`, and take screenshots and scroll over HTTP, so they also run on older SDKs. SDK versions before 0.0.47 (Python) and 1.0.22 (TypeScript) default to 2 GB, which `POST /computers` rejects with `400`. Their `screenshot()` and `screenshot_base64()` cannot read the stored-image path the API returns by default, and their `scroll()` takes no coordinates, so the computer scrolls at the top-left corner.
</Warning>

<Steps>
  <Step title="Install packages">
    ```bash theme={null}
    pip install orgo openai python-dotenv
    ```
  </Step>

  <Step title="Set up API keys">
    ```bash theme={null}
    export ORGO_API_KEY=your_orgo_api_key
    export OPENAI_API_KEY=your_openai_api_key
    ```
  </Step>

  <Step title="Run your first task">
    ```python theme={null}
    from openai import OpenAI
    from orgo import Computer

    # Initialize
    client = OpenAI()
    # Pass a size: SDKs before 0.0.47 default to 2 GB, which the API rejects.
    computer = Computer(ram=4, cpu=1)

    # Create request with task
    response = client.responses.create(
        model="gpt-5.6",
        tools=[{"type": "computer"}],
        input="Open Chrome and search for OpenAI",
    )

    # Execute the suggested actions, in order
    call = next(
        (item for item in response.output if item.type == "computer_call"),
        None,
    )
    if call:
        for action in call.actions:
            if action.type == "click":
                computer.left_click(action.x, action.y)
            elif action.type == "type":
                computer.type(action.text)

    # Clean up
    computer.destroy()
    ```
  </Step>
</Steps>

## Complete example

A full agent loop. The model returns a batch of actions per turn, so run every entry in `actions[]` in order before you screenshot.

<CodeGroup>
  ```python example.py expandable icon="python" theme={null}
  import os
  import time
  import requests
  from openai import OpenAI
  from orgo import Computer
  from dotenv import load_dotenv

  load_dotenv()

  API = "https://www.orgo.ai/api"
  HEADERS = {"Authorization": f"Bearer {os.environ['ORGO_API_KEY']}"}

  MODEL = "gpt-5.6"
  TOOLS = [{"type": "computer"}]

  INSTRUCTIONS = """You are controlling a desktop computer.
  - Always double-click desktop icons to open applications.
  - Use keyboard shortcuts as single commands (e.g. 'ctrl+c', not separate keys)."""


  def run_computer_task(task, computer_id=None, max_turns=20):
      """Execute a task using the OpenAI computer tool with Orgo."""

      client = OpenAI()
      # Pass a size for a new computer: SDKs before 0.0.47 default to 2 GB, which the API rejects.
      computer = Computer(computer_id=computer_id, ram=4, cpu=1)
      print(f"Computer ID: {computer.computer_id}")

      response = client.responses.create(
          model=MODEL,
          tools=TOOLS,
          instructions=INSTRUCTIONS,
          input=task,
          reasoning={"summary": "concise"},
      )

      for _ in range(max_turns):
          # Display progress
          for item in response.output:
              if item.type == "reasoning":
                  for summary in getattr(item, "summary", []):
                      print(f"[thinking] {getattr(summary, 'text', '')}")

          call = next(
              (item for item in response.output if item.type == "computer_call"),
              None,
          )

          # No computer call means the task is done.
          if call is None:
              print("Task completed")
              break

          # Run the whole batch, in order.
          for action in call.actions:
              print(f"-> {action.type}")
              execute_action(computer, action)
          time.sleep(1)  # Allow the UI to settle

          screenshot = screenshot_base64(computer)

          response = client.responses.create(
              model=MODEL,
              tools=TOOLS,
              previous_response_id=response.id,
              input=[{
                  "type": "computer_call_output",
                  "call_id": call.call_id,
                  "output": {
                      "type": "computer_screenshot",
                      "image_url": f"data:image/png;base64,{screenshot}",
                      "detail": "original",
                  },
              }],
              reasoning={"summary": "concise"},
          )

      return computer


  def screenshot_base64(computer):
      """The screen as base64 PNG, straight from the API."""
      response = requests.get(
          f"{API}/computers/{computer.computer_id}/screenshot",
          params={"response_format": "base64"},
          headers=HEADERS,
      )
      response.raise_for_status()
      return response.json()["image"]


  def scroll(computer, x, y, direction, amount):
      """Scroll at (x, y). Orgo scrolls up or down only."""
      if direction not in ("up", "down"):
          raise ValueError(f"Unsupported scroll direction: {direction}")
      response = requests.post(
          f"{API}/computers/{computer.computer_id}/scroll",
          json={"x": x, "y": y, "direction": direction, "amount": amount},
          headers=HEADERS,
      )
      response.raise_for_status()


  # Orgo presses keys with xdotool, which expects X keysym names.
  XDOTOOL_KEYS = {
      "enter": "Return", "return": "Return", "esc": "Escape", "escape": "Escape",
      "backspace": "BackSpace", "tab": "Tab", "space": "space", "delete": "Delete",
      "arrowup": "Up", "arrowdown": "Down", "arrowleft": "Left", "arrowright": "Right",
      "pageup": "Page_Up", "pagedown": "Page_Down", "home": "Home", "end": "End",
  }


  def to_xdotool(key):
      return XDOTOOL_KEYS.get(key.lower(), key.lower())


  def execute_action(computer, action):
      """Execute one computer action using Orgo."""

      match action.type:
          case "click":
              if getattr(action, "button", "left") == "right":
                  computer.right_click(action.x, action.y)
              else:
                  computer.left_click(action.x, action.y)

          case "double_click":
              computer.double_click(action.x, action.y)

          case "type":
              computer.type(action.text)

          case "keypress":
              keys = getattr(action, "keys", [])
              if keys:
                  # One combo, e.g. ["CTRL", "C"] becomes "ctrl+c"
                  computer.key("+".join(to_xdotool(k) for k in keys))

          case "scroll":
              scroll_y = getattr(action, "scroll_y", 0)
              direction = "down" if scroll_y > 0 else "up"
              scroll(computer, action.x, action.y, direction, max(1, abs(scroll_y) // 100))

          case "wait":
              computer.wait(2)

          case "screenshot":
              pass  # The loop screenshots after every batch


  if __name__ == "__main__":
      computer = run_computer_task("Open a terminal and list files")

      # Always clean up
      computer.destroy()
  ```
</CodeGroup>

## Usage examples

### Basic tasks

```python theme={null}
# Open a browser
computer = run_computer_task("Open Chrome")

# Navigate to a website
computer = run_computer_task("Go to github.com and search for orgo")

# Fill out a form
computer = run_computer_task("Fill out the contact form with test data")

# Always clean up
computer.destroy()
```

### Complex workflows

```python theme={null}
# Multi-step task
task = """
1. Open a text editor
2. Write a Python hello world program
3. Save it as hello.py
4. Open a terminal
5. Run the program
"""
computer = run_computer_task(task)
computer.destroy()
```

### Reusing sessions

```python theme={null}
# First task
computer = run_computer_task("Open VS Code")
computer_id = computer.computer_id

# Continue in same session
computer = run_computer_task(
    "Create a new Python file", 
    computer_id=computer_id
)

# Clean up when done
computer.destroy()
```

## Key concepts

### The agent loop

1. **Request** Send the task to the model.
2. **Actions** The model returns a `computer_call` holding a batch of actions.
3. **Execute** Your code runs every action in the batch, in order.
4. **Screenshot** Capture the result and return it as a `computer_call_output`.
5. **Repeat** Continue until a response has no `computer_call`.

### Action types

| Action | Description | Example |
| - | - | - |
| `click` | Click at coordinates | Click button at (100, 200) |
| `double_click` | Double-click | Open desktop icon |
| `move` | Move the cursor | Hover a menu |
| `drag` | Drag between points | Move a file |
| `type` | Type text | Enter username |
| `keypress` | Press key(s) | Press Enter, Ctrl+C |
| `scroll` | Scroll page | Scroll down 3 units |
| `wait` | Pause execution | Wait 2 seconds |
| `screenshot` | Take screenshot | Capture current state |

### Safety

The model acts on whatever the screen shows, and page content is untrusted input. OpenAI's guidance is to isolate the environment, restrict which sites the agent can reach, and keep a human in the loop for consequential steps. An Orgo computer covers the isolation half: the blast radius is that one computer, and `computer.destroy()` ends it.

## Best practices

### 1. Clear instructions

```python theme={null}
# Good: specific and clear
task = "Open Chrome, go to github.com, and star the orgo repository"

# Avoid: too vague
task = "Do some web stuff"
```

### 2. Error handling

```python theme={null}
def safe_run_task(task):
    """Run task with error handling."""
    computer = None
    try:
        computer = run_computer_task(task)
        return computer
    except Exception as e:
        print(f"Error: {e}")
        if computer:
            computer.destroy()
        raise
```

### 3. Session management

```python theme={null}
# Use context manager pattern
class ComputerSession:
    def __init__(self, task):
        self.task = task
        self.computer = None
        
    def __enter__(self):
        self.computer = run_computer_task(self.task)
        return self.computer
        
    def __exit__(self, *args):
        if self.computer:
            self.computer.destroy()

# Usage
with ComputerSession("Open calculator") as computer:
    print(f"Session ID: {computer.computer_id}")
```

### 4. Timing considerations

```python theme={null}
# Add delays for UI updates
time.sleep(1)  # After clicks
time.sleep(2)  # After opening applications
time.sleep(0.5)  # After typing
```

## Migrating from the preview tool

The preview integration still runs. These are the differences:

| | Preview | GA |
| - | - | - |
| Model | `computer-use-preview` | `gpt-5.6` |
| Tool type | `computer_use_preview` | `computer` |
| Display size | `display_width`, `display_height` | Removed, inferred from screenshots |
| Environment | `environment: "linux"` | Removed |
| Actions | One `action` per call | Batched `actions[]` array |
| Truncation | `truncation: "auto"` required | Not needed |

The preview tool declares its own display size, so on a default Orgo computer set `display_width: 1280` and `display_height: 720` to match the screenshots you send.

## Comparison with Claude

| Feature | OpenAI computer tool | Claude computer use |
| - | - | - |
| API | Responses API | Messages API |
| Model | `gpt-5.6` | `claude-sonnet-5-5` |
| Tool | `computer` | `computer_toolset_20260801` |
| Beta header | None | None |
| Reasoning | Optional summaries | Thinking blocks |
| Coordinates | Screenshot pixels | Screenshot pixels |
| Batched actions | Yes | Yes |

## Next steps

<CardGroup cols={2}>
  <Card title="OpenAI Docs" icon="book" href="https://developers.openai.com/api/docs/guides/tools-computer-use">
    Official OpenAI computer use documentation
  </Card>

  <Card title="Orgo Quickstart" icon="rocket" href="/quickstart">
    Learn more about Orgo computers
  </Card>
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.