> ## Documentation Index
> Fetch the complete documentation index at: https://docs.orgo.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Agent S3

> Let Agent S3 control an Orgo computer

Set up Agent S3, the open-source computer-use agent by Simular AI, against your own laptop or against an Orgo computer.

Agent S3 is the current release of the framework, distributed as the `gui-agents` package. It drops the manager and worker hierarchy Agent S2 used, adds a native coding agent that can write and run code alongside GUI actions, and adds Behavior Best-of-N, which runs several rollouts and picks the best one.

<Note>
  Agent S2 code imported from `gui_agents.s2`. Agent S3 lives under `gui_agents.s3` and renames the top-level class to `AgentS3`. Update both when you upgrade.
</Note>

## Coordinate space

Orgo computers boot at `1280x720x24`. Agent S3 never hands you a coordinate to translate. It emits `pyautogui` code, which `RemoteExecutor` runs **inside** the computer through `POST /computers/{id}/exec`. Both ends are already in the computer's native pixels, so there is nothing to scale.

Two things follow from that:

* Screenshots must reach the grounding model unresized. Downscaling one before Agent S3 sees it reintroduces exactly the offset this design avoids.
* Any resolution you mention in a prompt must match the computer's real one.

<Warning>
  `grounding_width` and `grounding_height` are not your computer's resolution. They declare the coordinate space the *grounding model* emits in, and the framework maps that space onto the screenshot. Set them to the values your grounding model's docs specify, not to `1280x720`. The Agent S README pairs UI-TARS-1.5-7B with `1920x1080`.
</Warning>

`GET /computers/{id}/screens` reports the current `width` and `height`.

## Setup

Install the required packages. `gui-agents` supports Python 3.9 through 3.12:

<CodeGroup>
  ```bash pip theme={null}
  pip install gui-agents pyautogui python-dotenv orgo
  ```

  ```bash requirements.txt theme={null}
  gui-agents
  pyautogui
  python-dotenv
  orgo
  pillow
  ```
</CodeGroup>

Set up your API keys. `HF_TOKEN` is for a grounding model hosted on Hugging Face Inference Endpoints:

<CodeGroup>
  ```bash terminal icon="terminal" theme={null}
  # Export as environment variables
  export OPENAI_API_KEY=your_openai_api_key
  export HF_TOKEN=your_huggingface_token
  export ORGO_API_KEY=your_orgo_api_key  # Remote mode only
  ```

  ```python setup.py icon="python" theme={null}
  import os
  os.environ["OPENAI_API_KEY"] = "your_openai_api_key"
  os.environ["HF_TOKEN"] = "your_huggingface_token"
  os.environ["ORGO_API_KEY"] = "your_orgo_api_key"  # Remote mode only
  ```

  ```bash .env icon="file" theme={null}
  OPENAI_API_KEY=your_openai_api_key
  HF_TOKEN=your_huggingface_token

  # Remote mode only
  ORGO_API_KEY=your_orgo_api_key
  USE_CLOUD_ENVIRONMENT=false
  ```
</CodeGroup>

## Simple usage

Run Agent S3 with natural language commands:

<CodeGroup>
  ```bash local icon="terminal" theme={null}
  # Local mode - controls your own laptop
  python agent_s3.py "Open Chrome and search for weather"
  ```

  ```bash remote icon="terminal" theme={null}
  # Remote mode - controls an Orgo computer
  USE_CLOUD_ENVIRONMENT=true python agent_s3.py "Open Chrome"
  ```

  ```bash interactive icon="terminal" theme={null}
  # Interactive mode
  python agent_s3.py
  ```
</CodeGroup>

The package also ships its own CLI, `agent_s`, which drives the laptop it runs on:

```bash icon="terminal" theme={null}
agent_s \
    --provider openai \
    --model gpt-5-2025-08-07 \
    --ground_provider huggingface \
    --ground_url http://localhost:8080 \
    --ground_model ui-tars-1.5-7b \
    --grounding_width 1920 \
    --grounding_height 1080
```

## Complete example

<CodeGroup>
  ```python agent_s3.py expandable icon="python" theme={null}
  #!/usr/bin/env python3

  import os
  import io
  import sys
  import time
  import requests
  from dotenv import load_dotenv
  from gui_agents.s3.agents.agent_s import AgentS3
  from gui_agents.s3.agents.grounding import OSWorldACI
  from orgo import Computer
  import pyautogui

  load_dotenv()

  CONFIG = {
      "model": os.getenv("AGENT_MODEL", "gpt-5-2025-08-07"),
      "model_type": os.getenv("AGENT_MODEL_TYPE", "openai"),
      "grounding_model": os.getenv("GROUNDING_MODEL", "ui-tars-1.5-7b"),
      "grounding_type": os.getenv("GROUNDING_MODEL_TYPE", "huggingface"),
      "grounding_url": os.getenv("GROUNDING_URL", "http://localhost:8080"),
      "grounding_width": int(os.getenv("GROUNDING_WIDTH", "1920")),
      "grounding_height": int(os.getenv("GROUNDING_HEIGHT", "1080")),
      "max_steps": int(os.getenv("MAX_STEPS", "10")),
      "step_delay": float(os.getenv("STEP_DELAY", "0.5")),
      "remote": os.getenv("USE_CLOUD_ENVIRONMENT", "false").lower() == "true"
  }


  class LocalExecutor:
      def __init__(self):
          self.pyautogui = pyautogui
          if sys.platform == "win32":
              self.platform = "windows"
          elif sys.platform == "darwin":
              self.platform = "darwin"
          else:
              self.platform = "linux"

      def screenshot(self):
          img = self.pyautogui.screenshot()
          buffer = io.BytesIO()
          img.save(buffer, format="PNG")
          buffer.seek(0)
          return buffer.getvalue()

      def exec(self, code):
          exec(code, {"pyautogui": self.pyautogui, "time": time})

      def destroy(self):
          # No cleanup needed for local executor
          pass


  class RemoteExecutor:
      def __init__(self):
          # Pass a size: SDKs before 0.0.47 default to 2 GB, which the API rejects.
          self.computer = Computer(ram=4, cpu=1)
          self.platform = "linux"
          # The Linux image ships python3 without pip or pyautogui, and the
          # actions Agent S3 emits import pyautogui. Install it once.
          self.computer.bash(
              "apt-get update -qq && apt-get install -y -qq python3-pip && "
              "pip3 install --break-system-packages -q pyautogui"
          )

      def screenshot(self):
          # Agent S3 wants raw image bytes, and response_format=binary returns
          # the PNG itself. The SDK's screenshot methods cannot read the
          # stored-image path the API returns by default.
          response = requests.get(
              f"https://www.orgo.ai/api/computers/{self.computer.computer_id}/screenshot",
              params={"response_format": "binary"},
              headers={"Authorization": f"Bearer {os.environ['ORGO_API_KEY']}"},
          )
          response.raise_for_status()
          return response.content

      def exec(self, code):
          result = self.computer.exec(code)
          if not result['success']:
              raise Exception(result.get('error', 'Execution failed'))
          if result['output']:
              print(f"Output: {result['output']}")

      def destroy(self):
          self.computer.destroy()


  def create_agent(executor):
      engine_params = {
          "engine_type": CONFIG["model_type"],
          "model": CONFIG["model"],
      }

      # grounding_width/height describe the grounding model's output space,
      # not the computer's screen. Leave them as that model's docs specify.
      engine_params_for_grounding = {
          "engine_type": CONFIG["grounding_type"],
          "model": CONFIG["grounding_model"],
          "base_url": CONFIG["grounding_url"],
          "grounding_width": CONFIG["grounding_width"],
          "grounding_height": CONFIG["grounding_height"],
      }

      grounding_agent = OSWorldACI(
          env=None,
          platform=executor.platform,
          engine_params_for_generation=engine_params,
          engine_params_for_grounding=engine_params_for_grounding,
      )

      return AgentS3(
          engine_params,
          grounding_agent,
          platform=executor.platform,
      )


  def run_task(agent, executor, instruction):
      print(f"\nTask: {instruction}")
      print(f"Mode: {'Remote' if CONFIG['remote'] else 'Local'}\n")

      for step in range(CONFIG["max_steps"]):
          print(f"Step {step + 1}/{CONFIG['max_steps']}")

          obs = {"screenshot": executor.screenshot()}
          info, action = agent.predict(instruction=instruction, observation=obs)

          if info:
              print(f"[thinking] {info}")

          if not action or not action[0]:
              print("Complete")
              return True

          try:
              print(f"[action] {action[0]}")
              executor.exec(action[0])
          except Exception as e:
              print(f"Error: {e}")
              instruction = "The previous action failed. Try a different approach."

          time.sleep(CONFIG["step_delay"])

      print("Max steps reached")
      return False


  def main():
      executor = RemoteExecutor() if CONFIG["remote"] else LocalExecutor()
      try:
          agent = create_agent(executor)

          if len(sys.argv) > 1:
              run_task(agent, executor, " ".join(sys.argv[1:]))
          else:
              print("Interactive mode (type 'exit' to quit)\n")
              while True:
                  task = input("Task: ").strip()
                  if task == "exit":
                      break
                  elif task:
                      run_task(agent, executor, task)
      finally:
          # Clean up
          executor.destroy()


  if __name__ == "__main__":
      main()
  ```
</CodeGroup>

<Note>
  Pass a `LocalEnv` from `gui_agents.s3.utils.local_env` as `OSWorldACI(env=...)` to turn on the coding agent. It executes code on the host running the script, which for `RemoteExecutor` is your laptop rather than the Orgo computer. Leave it `None` unless that is what you want.
</Note>

## Platform requirements

These apply to local mode. Remote mode runs the actions on the Orgo computer, and `RemoteExecutor` installs `pyautogui` there when it starts.

### macOS

Grant Terminal access: System Settings, then Privacy & Security, then Accessibility

### Windows

May require running Terminal as Administrator

### Linux

Install dependencies:

```bash icon="terminal" theme={null}
sudo apt-get install python3-tk python3-dev
```

## Environment variables

| Variable | Default | Description |
| - | - | - |
| `OPENAI_API_KEY` | - | Main model API key |
| `HF_TOKEN` | - | Hugging Face token for a hosted grounding model |
| `ORGO_API_KEY` | - | Orgo API key (remote mode) |
| `USE_CLOUD_ENVIRONMENT` | `false` | Set to `true` for remote execution |
| `AGENT_MODEL` | `gpt-5-2025-08-07` | Main reasoning model |
| `GROUNDING_MODEL` | `ui-tars-1.5-7b` | Visual grounding model |
| `GROUNDING_URL` | `http://localhost:8080` | Grounding model endpoint |
| `GROUNDING_WIDTH` | `1920` | Grounding model output width |
| `GROUNDING_HEIGHT` | `1080` | Grounding model output height |
| `MAX_STEPS` | `10` | Maximum steps per task |
| `STEP_DELAY` | `0.5` | Seconds between actions |

## Architecture

Agent S3 replaced Agent S2's hierarchical planner with a flatter design:

**Native coding agent**: writes and runs code alongside GUI actions, so tasks better solved in a shell stop being click sequences

**Behavior Best-of-N**: runs several rollouts of a task, summarizes each as a behavior narrative, and has a judge pick the one that completed it

**Mixture of grounding**: routes actions to specialized visual grounding models for precise UI localization

**Cross-platform support**: works on macOS, Windows, and Linux

For current benchmark results, see the [Agent S3 writeup](https://www.simular.ai/articles/agent-s3) and the [OSWorld leaderboard](https://os-world.github.io/).

## Resources

* [GitHub Repository](https://github.com/simular-ai/Agent-S)
* [Agent S3 writeup](https://www.simular.ai/articles/agent-s3)
* [Agent S2 whitepaper](https://arxiv.org/abs/2504.00906)
* [OSWorld Benchmark](https://os-world.github.io/)

## Video tutorial

<iframe width="100%" height="400" src="https://www.youtube.com/embed/GgUC4q7MTaw" title="Agent S2 Setup Tutorial" frameBorder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture" allowFullScreen />

The video was recorded against Agent S2. The imports and class names above are the Agent S3 ones.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.