Building Blocks for Computer-Use Agents Research
🏆 ICLR Oral, Top 1%
Evaluate and train multimodal agents to use apps like humans do (by clicking, typing, and scrolling):
✅ Unlimited data (for evaluating and training UI-agents): Configurable state and design to generate thousands of versions of each app
✅ Lightweight: runs on a single CPU (and Python-based); no Docker or OS emulators needed
✅ Ground truth rewards: task rewards are based on the underlying state and all app logic is transparent in Python
- Clone
git clone https://github.com/facebookresearch/OpenApps.git
- Install
uv sync
see docs for details.
Simply run:
uv run launch.py
A browser opens on the apps once they are up, sized to the configured device.
Pass headless=True to just serve them.
Each app can be modified with variables available in config/apps. You can override any of these via command line:
uv run launch.py app.todo.title='Super Todo'Hydra syntax, and the one thing worth memorising up front:
| Form | Means | Example |
|---|---|---|
key=value |
Change something that already exists | uv run launch.py device=phone |
+key=value |
Add something not in the defaults list | uv run launch.py +experiment=phone |
++key=value |
Add or change, whichever applies | uv run launch.py ++apps.todo.title=Tasks |
A plain = fails on a key that does not exist yet, and a + fails on one that
does — the error tells you which you needed. Groups already in the defaults
list (device, agent, tasks, apps/theme, each app's layout and
content) take =. Only experiment needs +, because it is deliberately
not a default: an experiment config overrides other groups, so Hydra has to
compose it last, and appending it is what + does.
uv run launch.py device=phone # existing group
uv run launch.py apps/theme=dark # existing group
uv run launch.py +experiment=phone # preset bundle, not a default
uv run launch.py +experiment=phone device=tablet # bundle, then override one partAn experiment is a named bundle that sets several groups at once, for the
cases where the halves have to agree. +experiment=phone is device=phone
plus the home-screen layout plus a larger step budget; device=phone alone
gives you a phone-sized window still rendering the desktop layout. See
config/experiment/.
Learn more about to customize the content and appearance of apps in the docs.
For a hot reloading dev server (live changes in browser):
scripts/dev.sh
The shop is a Python rewrite of WebShop, on by default with a 999-product catalog from WebShop's item dump. See Online Shop for its variations, catalog and data.
For agents to directly interact with apps, install: playwright install chromium.
Launch an agent to perform a task of adding a meeting with Dennis to the calendar:
# export GPT55_API_KEY=""
uv run launch_agent.py agent=GPT-5.5-computer-use task_name=add_meeting_with_dennis
To see the agent solving the task live, add the headless argument:
uv run launch_agent.py ... browsergym_env_args.headless=False
You can specify the agent of your choice with the agent= argument. For example agent=dummy is a simple agent that clicks randomly on any buttons, great for exploration!
Learn more about launching with OpenAI, Claude, and VLLM models such as UI-Tars in our docs.
Copy .env.example and fill in what you need — launch_agent.py and
launch_parallel_agents.py call load_dotenv(), so a .env at the repo root is picked up
automatically, and .env is git-ignored so keys stay out of the configs:
cp .env.example .env| Variable | Read by | Purpose |
|---|---|---|
USER |
config/config*.yaml, config/mode/* |
W&B entity and the logs_dir path |
GPT55_API_KEY |
config/agent/GPT-5.5-*.yaml |
key for the OpenAI-compatible endpoint |
AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY, AWS_SESSION_TOKEN |
config/agent/claude_4_sonnet.yaml (client_type: aws) |
Bedrock credentials, when left null in the config |
WANDB_API_KEY, WANDB_BASE_URL, WANDB_MODE |
wandb |
auth, self-hosted server, and WANDB_MODE=offline to skip online logging |
EXPERIMENT_CONFIG_PATH |
src/open_apps/configs.py |
optional path loaded by load_config() instead of the default config |
Agent API keys are read through Hydra interpolation, so any variable name works — point the
agent's api_key at the one you use:
uv run launch_agent.py agent=GPT-5.5-computer-use 'agent.api_key=${oc.env:OPENAI_API_KEY}'The batch scripts take environment variables too (AGENTS, COUNT, MAX_PARALLEL,
VLLM_HOST, …), but they are read by the shell, not through .env — export them at the
call site, or set -a; source .env; set +a first. They are listed in the
agents docs; the MCP server's
variables are in src/open_apps/mcp/README.md.
config/mode/slurm_cluster.yaml and the #SBATCH lines in scripts/conduct_slurm.sh ship
with placeholder accounts and paths that sbatch will reject. Copy the mode to an
internal- twin — .gitignore keeps any internal-* file untracked, so your site's paths
and account names can't be committed by accident:
cp config/mode/slurm_cluster.yaml config/mode/internal-slurm_cluster.yaml
uv run launch_parallel_agents.py mode=internal-slurm_cluster agent=dummy \
tasks=longer_horizon parallel_tasks.task_names=all use_wandb=TrueSee the agents docs for the full SLURM + vLLM + W&B walkthrough.
demo.mp4
main moves. The v1.0-paper tag pins the paper-era config surface used for the variation grid in arXiv:2511.20766 — check it out if you are reproducing or comparing against our setup:
git clone https://github.com/facebookresearch/OpenApps.git
cd OpenApps
git checkout v1.0-paper
uv syncEverything after that tag is free to diverge. The first such change is the appearance refactor: the per-app appearance config group is replaced by a shared apps/theme (look) plus a per-app apps/<app>/layout (structure), so overrides written as apps/todo/appearance=dark_theme no longer resolve on main. See App variations in the docs for the current axes, which include a table mapping every old appearance value onto its replacement.
The tag is the config surface the paper used, not a byte-exact snapshot of the runs: it carries the app, task and harness fixes landed since publication, some of which move rewards (map tasks now match coordinates by ground distance, for example). Expect small differences from the published tables.
We welcome pull requests with new features or issues via GitHub.
uv sync --extra dev
To build docs:
mkdocs build
mkdocs serve
this will launch docs available at https://facebookresearch.github.io/OpenApps/
Run all tests via:
uv run -m pytest tests/Our apps are built on top of several excellent frameworks:
- FastHTML framework and examples which allowed us to build fully functional apps in Python, the language most familiar to AI researchers.
- Browser Gym and AgentLab:
- Open Street Maps: https://www.openstreetmap.org/copyright for our Maps apps.
- (for the online shop) WebShop, developed by Princeton University: our shop is a rewrite, and its catalog is converted from WebShop's item dump.
Some icons are have been designed using resources from Flaticon.com
Our work is licensed under CC-BY-NC, please refer to the LICENSE file in the top level directory.
Copyright © Meta Platforms, Inc. See the Terms of Use and Privacy Policy for this project.
- A big thank you to Taj Gillin for implementing an MCP interface for OpenApps and Jiayu Wang for suggestions to improve our harness!
- Another thank you to Yuval Kansal for fixing inconsistencies in task phrasings!
- BrowserGym: https://github.com/servicenow/browsergym
- NE Agents Day (🏆Oral Award) : https://ne-agents-day.github.io/
- OpenEnv (HugginFace and PyTorch RL environment): https://huggingface.co/docs/openenv/environments/openapp
@article{ullrich2025openapps0,
title = {OpenApps: Simulating Environment Variations to Measure UI-Agent Reliability},
author = {Karen Ullrich and Jingtong Su and Claudia Shi and Arjun Subramonian and Amir Bar and Ivan Evtimov and Nikolaos Tsilivis and Randall Balestriero and Julia Kempe and Mark Ibrahim},
year = {2025},
journal = {arXiv preprint arXiv: 2511.20766}
}
