> ## Documentation Index
> Fetch the complete documentation index at: https://docs.orgo.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Gemini Computer Use

> Control an Orgo computer with Gemini

Wire Google's computer use tool to an Orgo computer. Gemini sees a screenshot, returns a batch of actions, and your loop runs them on the computer.

Computer use is a built-in tool on Gemini 3.x rather than a separate model. `gemini-3.7-flash` is the model Google recommends for it, through the Interactions API.

<Info>
  `gemini-2.5-computer-use-preview-10-2025` is now marked legacy. It used the older `generate_content` call and the `types.ComputerUse` tool object. The code below uses the current Interactions API.
</Info>

## Coordinate space

Orgo computers boot at `1280x720x24`.

Gemini does not work in pixels. It returns coordinates on a normalized 0-999 grid, scaled to the screenshot you sent. You denormalize every coordinate against the computer's real resolution before you can click with it:

```python theme={null}
SCREEN_WIDTH = 1280
SCREEN_HEIGHT = 720

def denormalize_x(x: int) -> int:
    return int(x / 1000 * SCREEN_WIDTH)

def denormalize_y(y: int) -> int:
    return int(y / 1000 * SCREEN_HEIGHT)
```

<Warning>
  These two constants are load-bearing. They are the only thing translating Gemini's grid into pixels. Set them to `1024x768` against a `1280x720` computer and every horizontal click lands roughly 25% short of where Gemini aimed. Nothing errors. The agent looks like it is confidently clicking the wrong thing.
</Warning>

There is no display size to declare on the request. If you create a computer at a non-default resolution, or resize its screen later, set the constants to that resolution instead. `GET /computers/{id}/screens` reports the current `width` and `height`.

## Setup

Install the required packages. Gemini 3.x computer use needs `google-genai` 2.7.0 or later:

```bash pip theme={null}
pip install "orgo" "google-genai>=2.7.0" pillow python-dotenv
```

Set up your API keys in a `.env` file:

```bash .env icon="file" theme={null}
ORGO_API_KEY=your_orgo_api_key
GEMINI_API_KEY=your_gemini_api_key
```

Or export them as environment variables:

```bash terminal icon="terminal" theme={null}
export ORGO_API_KEY=your_orgo_api_key
export GEMINI_API_KEY=your_gemini_api_key
```

## Complete example

An Orgo computer is a full desktop, so declare `environment: "desktop"`. The `browser` environment adds page-level actions such as `navigate` and `go_back`, and expects the current page URL back in every result.

```python example.py expandable icon="python" theme={null}
import os
import io
import json
import time
import base64
import requests
from google import genai
from orgo import Computer
from PIL import Image
from dotenv import load_dotenv

# Load environment variables
load_dotenv()

# Initialize Gemini client
client = genai.Client(api_key=os.environ.get('GEMINI_API_KEY'))

# Connect to your Orgo computer
# Use the computer UUID (the `id` returned when you create a computer)
computer = Computer(computer_id="your-computer-id")

MODEL = 'gemini-3.7-flash'

TOOLS = [{
    "type": "computer_use",
    "environment": "desktop",
    "enable_prompt_injection_detection": True,
}]

# Screen resolution
SCREEN_WIDTH = 1280
SCREEN_HEIGHT = 720

# System prompt with Ubuntu-specific instructions
SYSTEM_PROMPT = f"""You are controlling an Ubuntu Linux virtual machine with a display resolution of {SCREEN_WIDTH}x{SCREEN_HEIGHT}.

<UBUNTU_DESKTOP_GUIDELINES>
* CRITICAL: When opening applications or files on the Ubuntu desktop, you MUST USE DOUBLE-CLICK, not single-click
* Single-click only selects desktop icons but DOES NOT open them
* Desktop interactions:
  - Desktop icons (apps/folders): DOUBLE-CLICK to open
  - Menu items: SINGLE-CLICK to select
  - Taskbar/launcher icons: SINGLE-CLICK to open
  - Window buttons (close/minimize/maximize): SINGLE-CLICK
  - File browser items: DOUBLE-CLICK to open
* When you need to submit or confirm, use the 'Enter' key
</UBUNTU_DESKTOP_GUIDELINES>

<IMPORTANT_NOTES>
* Wait for applications to load before acting on them
* Batch multiple actions together when possible before checking the result
</IMPORTANT_NOTES>"""


def denormalize_x(x: int) -> int:
    """Convert a normalized x coordinate (0-999) to an actual pixel."""
    return int(x / 1000 * SCREEN_WIDTH)


def denormalize_y(y: int) -> int:
    """Convert a normalized y coordinate (0-999) to an actual pixel."""
    return int(y / 1000 * SCREEN_HEIGHT)


def get_screenshot_png() -> str:
    """Get a base64 PNG screenshot. Gemini accepts PNG only, so re-encode to be safe."""
    # Ask the API for inline base64. The SDK's screenshot methods cannot read
    # the stored-image path the API returns by default.
    response = requests.get(
        f"https://www.orgo.ai/api/computers/{computer.computer_id}/screenshot",
        params={"response_format": "base64"},
        headers={"Authorization": f"Bearer {os.environ['ORGO_API_KEY']}"},
    )
    response.raise_for_status()
    image_data = base64.b64decode(response.json()["image"])
    image = Image.open(io.BytesIO(image_data))
    png_buffer = io.BytesIO()
    image.save(png_buffer, format='PNG')
    return base64.b64encode(png_buffer.getvalue()).decode('utf-8')


def execute_action(name: str, args: dict) -> dict:
    """Run one Gemini action on the computer. Returns extra result fields."""
    if name in ("click", "double_click", "right_click"):
        x, y = denormalize_x(args["x"]), denormalize_y(args["y"])
        {
            "click": computer.left_click,
            "double_click": computer.double_click,
            "right_click": computer.right_click,
        }[name](x, y)

    elif name == "type":
        computer.type(args["text"])
        if args.get("press_enter", False):
            computer.key("Return")

    elif name == "press_key":
        computer.key(args["key"])

    elif name == "hotkey":
        computer.key("+".join(args["keys"]).lower())

    elif name == "scroll":
        # Orgo scrolls up or down in clicks, so treat the pixel magnitude as a hint.
        if args["direction"] not in ("up", "down"):
            return {"error": f"Unsupported scroll direction: {args['direction']}"}
        amount = max(1, args.get("magnitude_in_pixels", 300) // 100)
        # The SDK's scroll() takes no coordinates, so call the API with them.
        response = requests.post(
            f"https://www.orgo.ai/api/computers/{computer.computer_id}/scroll",
            json={
                "x": denormalize_x(args["x"]),
                "y": denormalize_y(args["y"]),
                "direction": args["direction"],
                "amount": amount,
            },
            headers={"Authorization": f"Bearer {os.environ['ORGO_API_KEY']}"},
        )
        response.raise_for_status()

    elif name == "wait":
        computer.wait(args.get("seconds", 2))

    elif name == "take_screenshot":
        pass  # Every result already carries a fresh screenshot

    else:
        return {"error": f"Unsupported action: {name}"}

    return {}


def execute_function_calls(interaction) -> list:
    """Run every function call in the turn, in order."""
    results = []

    for step in interaction.steps:
        if step.type != "function_call":
            continue

        print(f"  -> {step.name}: {step.arguments.get('intent', '')}")

        try:
            result = execute_action(step.name, step.arguments)
        except Exception as error:
            print(f"    Error: {error}")
            result = {"error": str(error)}

        time.sleep(1)  # Wait for the UI to update
        results.append((step.name, step.call_id, result))

    return results


def get_function_results(results: list) -> list:
    """Return one result per call, each carrying the new screen."""
    screenshot_png = get_screenshot_png()

    return [
        {
            "type": "function_result",
            "name": name,
            "call_id": call_id,
            "result": [
                {"type": "text", "text": json.dumps({"status": "completed", **result})},
                {"type": "image", "data": screenshot_png, "mime_type": "image/png"},
            ],
        }
        for name, call_id, result in results
    ]


# Define the task
task = "Open Chrome and search for 'gemini ai'"
print(f"Task: {task}\n")

interaction = client.interactions.create(
    model=MODEL,
    system_instruction=SYSTEM_PROMPT,
    input=[
        {"type": "text", "text": task},
        {"type": "image", "data": get_screenshot_png(), "mime_type": "image/png"},
    ],
    tools=TOOLS,
)

# Agent loop
for turn in range(20):
    print(f"\n--- Turn {turn + 1} ---")

    has_function_calls = any(
        step.type == "function_call" for step in interaction.steps
    )

    if not has_function_calls:
        text = " ".join(
            block.text
            for step in interaction.steps if step.type == "model_output"
            for block in step.content if block.type == "text"
        )
        print(f"Agent finished: {text}")
        break

    results = execute_function_calls(interaction)

    interaction = client.interactions.create(
        model=MODEL,
        previous_interaction_id=interaction.id,
        input=get_function_results(results),
        tools=TOOLS,
    )

# The computer keeps running. Call computer.destroy() to tear it down.
```

## Usage examples

### Basic tasks

```python theme={null}
# Change the task variable to control what Gemini does
task = "Open Chrome and search for 'gemini ai'"

# Navigate to a website
task = "Go to github.com and search for 'orgo'"

# Fill a form
task = "Fill out the contact form with test data"
```

### Complex workflows

```python theme={null}
# Multi-step task
task = """
1. Open a text editor
2. Write a Python hello world program
3. Save it as hello.py
4. Open a terminal
5. Run the program
"""
```

## Key concepts

### The agent loop

1. **Request** Send the task and a screenshot to the model.
2. **Actions** The model returns `function_call` steps.
3. **Execute** Your code runs them in order.
4. **Screenshot** Capture the result and return one `function_result` per call.
5. **Repeat** Chain turns with `previous_interaction_id` until no call comes back.

### Reading the response

Everything the model produced for a turn is in `interaction.steps`. A step of type `function_call` carries `name`, `call_id`, and `arguments`. A step of type `model_output` carries `content` blocks, and its `text` blocks are the agent's final answer.

Every action's `arguments` includes an `intent` string explaining why the model chose it. It is the cheapest debugging signal in the loop, so log it.

### System prompt

Pass the Ubuntu guidance as `system_instruction`. Without it Gemini single-clicks desktop icons and nothing opens:

```python theme={null}
interaction = client.interactions.create(
    model=MODEL,
    system_instruction=SYSTEM_PROMPT,
    input=[...],
    tools=TOOLS,
)
```

### Getting your computer ID

`computer_id` is the computer UUID: the `id` field that [Create computer](/api-reference/computers/create) returns. In the dashboard, it is the last segment of the computer page's URL, `https://www.orgo.ai/computers/{id}`. Pass it as `Computer(computer_id="your-computer-id")`.

### Image format conversion

Orgo returns screenshots as PNG by default. Gemini accepts PNG only, so the helper decodes and re-encodes each screenshot as PNG. That keeps the loop working if you ask Orgo for another screenshot format:

```python theme={null}
def get_screenshot_png() -> str:
    response = requests.get(
        f"https://www.orgo.ai/api/computers/{computer.computer_id}/screenshot",
        params={"response_format": "base64"},
        headers={"Authorization": f"Bearer {os.environ['ORGO_API_KEY']}"},
    )
    response.raise_for_status()
    image_data = base64.b64decode(response.json()["image"])
    image = Image.open(io.BytesIO(image_data))
    png_buffer = io.BytesIO()
    image.save(png_buffer, format='PNG')
    return base64.b64encode(png_buffer.getvalue()).decode('utf-8')
```

### Safety confirmations

With `enable_prompt_injection_detection` on, the model can return a `safety_decision` of `require_confirmation` on an action. Prompt your user before running it, and include a `safety_acknowledgement` in that call's result when they approve. Skip the confirmation and the action does not run.

### Action arguments

Coordinates are normalized 0-999. Denormalize `x`, `y`, `start_x`, `start_y`, `end_x`, and `end_y` before use.

| Action | Arguments |
| - | - |
| `click`, `double_click`, `triple_click`, `right_click`, `middle_click` | `x`, `y`, `intent` |
| `move` | `x`, `y`, `intent` |
| `type` | `text`, `press_enter` (optional), `intent` |
| `press_key` | `key`, `intent` |
| `hotkey` | `keys` (list), `intent` |
| `scroll` | `x`, `y`, `direction`, `magnitude_in_pixels` (optional), `intent` |
| `drag_and_drop` | `start_x`, `start_y`, `end_x`, `end_y`, `intent` |
| `wait` | `seconds` (optional), `intent` |
| `take_screenshot` | `intent` |

The `desktop` environment also exposes `mouse_down`, `mouse_up`, `key_down`, and `key_up` for held input.

## Tool compatibility

| Gemini action | Orgo call | Description |
| - | - | - |
| `click` | `computer.left_click(x, y)` | Click at coordinates |
| `double_click` | `computer.double_click(x, y)` | Double-click |
| `right_click` | `computer.right_click(x, y)` | Right-click |
| `type` | `computer.type(text)` | Type text |
| `press_key` | `computer.key(key)` | Press one key |
| `hotkey` | `computer.key(keys)` | Press a combination (e.g. "ctrl+c") |
| `scroll` | `POST /computers/{id}/scroll` with `x`, `y`, `direction`, `amount` | Scroll up or down at coordinates |
| `wait` | `computer.wait(seconds)` | Wait |
| `take_screenshot` | `GET /computers/{id}/screenshot?response_format=base64` | Capture screen (PNG) |

`drag_and_drop` maps to `computer.drag(start_x, start_y, end_x, end_y)` once you denormalize its four coordinates. The loop above does not handle it.

The loop takes screenshots and scrolls over HTTP rather than through the SDK. The SDK's `screenshot()` and `screenshot_base64()` cannot read the stored-image path the API returns by default, and its `scroll()` takes no coordinates, so the computer scrolls at the top-left corner.

## Best practices

### 1. Clear instructions

```python theme={null}
# Good: specific and clear
task = "Go to amazon.com and find the top 3 rated laptops under $1000"

# Avoid: too vague
task = "Find some laptops"
```

### 2. Send a system instruction, not a user turn

```python theme={null}
interaction = client.interactions.create(
    model=MODEL,
    system_instruction=SYSTEM_PROMPT,  # OS-specific guidance
    input=[...],
    tools=TOOLS,
)
```

### 3. Convert coordinates

```python theme={null}
actual_x = denormalize_x(args["x"])
actual_y = denormalize_y(args["y"])
```

### 4. Handle image format

```python theme={null}
screenshot_png = get_screenshot_png()
```

### 5. Return one result per call

Every `function_call` needs a matching `function_result` with the same `call_id`. Attach the new screenshot to each of them.

### 6. Add delays

```python theme={null}
time.sleep(1)  # Wait for UI to update after actions
```

## Comparison with Claude and OpenAI

| Feature | Gemini computer use | Claude computer use | OpenAI computer use |
| - | - | - | - |
| API | Interactions API | Messages API | Responses API |
| Model | `gemini-3.7-flash` | `claude-sonnet-5-5` | `gpt-5.6` |
| Instructions | Supported | Supported | Supported |
| Coordinates | Normalized 0-999 | Screenshot pixels | Screenshot pixels |
| Image format | PNG required | JPEG or PNG | JPEG or PNG |
| Batched actions | Yes | Yes | Yes |

## Limitations

* **Image format**: PNG only. Orgo's default screenshot is already PNG.
* **Coordinates**: normalized, so a wrong `SCREEN_WIDTH` or `SCREEN_HEIGHT` silently misplaces every click.
* **Rate limits**: subject to Gemini API rate limits.

## Troubleshooting

### The model does not double-click desktop icons

Pass the Ubuntu system prompt as `system_instruction`. Without it, single clicks only select icons.

### INVALID\_ARGUMENT: Unable to process input image

Gemini received an image that is not PNG. Run the screenshot through `get_screenshot_png()` first.

### Clicks land in the wrong place

`SCREEN_WIDTH` and `SCREEN_HEIGHT` do not match the computer. Check them against `GET /computers/{id}/screens`.

### Missing API key

Ensure both environment variables are set in your `.env` file:

```bash theme={null}
ORGO_API_KEY=your_orgo_api_key
GEMINI_API_KEY=your_gemini_api_key
```

## Next steps

<CardGroup cols={2}>
  <Card title="Gemini Docs" icon="book" href="https://ai.google.dev/gemini-api/docs/computer-use">
    Official Gemini computer use documentation
  </Card>

  <Card title="Orgo Quickstart" icon="rocket" href="/quickstart">
    Learn more about Orgo computers
  </Card>

  <Card title="API Reference" icon="code" href="/api-reference/introduction">
    Complete Orgo API documentation
  </Card>
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.