gui-agents package. It drops the manager and worker hierarchy Agent S2 used, adds a native coding agent that can write and run code alongside GUI actions, and adds Behavior Best-of-N, which runs several rollouts and picks the best one.
Agent S2 code imported from
gui_agents.s2. Agent S3 lives under gui_agents.s3 and renames the top-level class to AgentS3. Update both when you upgrade.Coordinate space
Orgo computers boot at1280x720x24. Agent S3 never hands you a coordinate to translate. It emits pyautogui code, which RemoteExecutor runs inside the computer through POST /computers/{id}/exec. Both ends are already in the computer’s native pixels, so there is nothing to scale.
Two things follow from that:
- Screenshots must reach the grounding model unresized. Downscaling one before Agent S3 sees it reintroduces exactly the offset this design avoids.
- Any resolution you mention in a prompt must match the computer’s real one.
GET /computers/{id}/screens reports the current width and height.
Setup
Install the required packages.gui-agents supports Python 3.9 through 3.12:
HF_TOKEN is for a grounding model hosted on Hugging Face Inference Endpoints:
Simple usage
Run Agent S3 with natural language commands:agent_s, which drives the laptop it runs on:
Complete example
Pass a
LocalEnv from gui_agents.s3.utils.local_env as OSWorldACI(env=...) to turn on the coding agent. It executes code on the host running the script, which for RemoteExecutor is your laptop rather than the Orgo computer. Leave it None unless that is what you want.Platform requirements
These apply to local mode. Remote mode runs the actions on the Orgo computer, andRemoteExecutor installs pyautogui there when it starts.
macOS
Grant Terminal access: System Settings, then Privacy & Security, then AccessibilityWindows
May require running Terminal as AdministratorLinux
Install dependencies:Environment variables
Architecture
Agent S3 replaced Agent S2’s hierarchical planner with a flatter design: Native coding agent: writes and runs code alongside GUI actions, so tasks better solved in a shell stop being click sequences Behavior Best-of-N: runs several rollouts of a task, summarizes each as a behavior narrative, and has a judge pick the one that completed it Mixture of grounding: routes actions to specialized visual grounding models for precise UI localization Cross-platform support: works on macOS, Windows, and Linux For current benchmark results, see the Agent S3 writeup and the OSWorld leaderboard.Resources
Video tutorial
The video was recorded against Agent S2. The imports and class names above are the Agent S3 ones.