Skip to content
facebookresearchPublic

About

An open source environment for digital agents.

Resources

Code of conduct

Contributing

Security policy

Stars

118 stars

Watchers

2 watching

Forks

Repository files navigation

image OpenApps

Building Blocks for Computer-Use Agents Research

🏆 ICLR Oral, Top 1%

📒 docs | 📑 ArXiV | 🎬 Video Tutorial

Evaluate and train multimodal agents to use apps like humans do (by clicking, typing, and scrolling):

✅ Unlimited data (for evaluating and training UI-agents): Configurable state and design to generate thousands of versions of each app

✅ Lightweight: runs on a single CPU (and Python-based); no Docker or OS emulators needed

✅ Ground truth rewards: task rewards are based on the underlying state and all app logic is transparent in Python

Install

  1. Clone
git clone https://github.com/facebookresearch/OpenApps.git
  1. Install
uv sync

see docs for details.

Run OpenApps

Simply run:

uv run launch.py
image

A browser opens on the apps once they are up, sized to the configured device. Pass headless=True to just serve them.

Each app can be modified with variables available in config/apps. You can override any of these via command line:

uv run launch.py app.todo.title='Super Todo'

Overrides: when to use =, += and +

Hydra syntax, and the one thing worth memorising up front:

Form Means Example
key=value Change something that already exists uv run launch.py device=phone
+key=value Add something not in the defaults list uv run launch.py +experiment=phone
++key=value Add or change, whichever applies uv run launch.py ++apps.todo.title=Tasks

A plain = fails on a key that does not exist yet, and a + fails on one that does — the error tells you which you needed. Groups already in the defaults list (device, agent, tasks, apps/theme, each app's layout and content) take =. Only experiment needs +, because it is deliberately not a default: an experiment config overrides other groups, so Hydra has to compose it last, and appending it is what + does.

uv run launch.py device=phone                    # existing group
uv run launch.py apps/theme=dark                 # existing group
uv run launch.py +experiment=phone               # preset bundle, not a default
uv run launch.py +experiment=phone device=tablet # bundle, then override one part

An experiment is a named bundle that sets several groups at once, for the cases where the halves have to agree. +experiment=phone is device=phone plus the home-screen layout plus a larger step budget; device=phone alone gives you a phone-sized window still rendering the desktop layout. See config/experiment/.

Learn more about to customize the content and appearance of apps in the docs.

For a hot reloading dev server (live changes in browser): scripts/dev.sh

The online shop

The shop is a Python rewrite of WebShop, on by default with a 999-product catalog from WebShop's item dump. See Online Shop for its variations, catalog and data.

Launch an Agent

For agents to directly interact with apps, install: playwright install chromium.

Launch an agent to perform a task of adding a meeting with Dennis to the calendar:

# export GPT55_API_KEY=""
uv run launch_agent.py agent=GPT-5.5-computer-use task_name=add_meeting_with_dennis

To see the agent solving the task live, add the headless argument:

uv run launch_agent.py ... browsergym_env_args.headless=False

gif (1)

You can specify the agent of your choice with the agent= argument. For example agent=dummy is a simple agent that clicks randomly on any buttons, great for exploration!

Learn more about launching with OpenAI, Claude, and VLLM models such as UI-Tars in our docs.

Environment variables

Copy .env.example and fill in what you need — launch_agent.py and launch_parallel_agents.py call load_dotenv(), so a .env at the repo root is picked up automatically, and .env is git-ignored so keys stay out of the configs:

cp .env.example .env
Variable Read by Purpose
USER config/config*.yaml, config/mode/* W&B entity and the logs_dir path
GPT55_API_KEY config/agent/GPT-5.5-*.yaml key for the OpenAI-compatible endpoint
AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY, AWS_SESSION_TOKEN config/agent/claude_4_sonnet.yaml (client_type: aws) Bedrock credentials, when left null in the config
WANDB_API_KEY, WANDB_BASE_URL, WANDB_MODE wandb auth, self-hosted server, and WANDB_MODE=offline to skip online logging
EXPERIMENT_CONFIG_PATH src/open_apps/configs.py optional path loaded by load_config() instead of the default config

Agent API keys are read through Hydra interpolation, so any variable name works — point the agent's api_key at the one you use:

uv run launch_agent.py agent=GPT-5.5-computer-use 'agent.api_key=${oc.env:OPENAI_API_KEY}'

The batch scripts take environment variables too (AGENTS, COUNT, MAX_PARALLEL, VLLM_HOST, …), but they are read by the shell, not through .env — export them at the call site, or set -a; source .env; set +a first. They are listed in the agents docs; the MCP server's variables are in src/open_apps/mcp/README.md.

Running on a cluster

config/mode/slurm_cluster.yaml and the #SBATCH lines in scripts/conduct_slurm.sh ship with placeholder accounts and paths that sbatch will reject. Copy the mode to an internal- twin — .gitignore keeps any internal-* file untracked, so your site's paths and account names can't be committed by accident:

cp config/mode/slurm_cluster.yaml config/mode/internal-slurm_cluster.yaml
uv run launch_parallel_agents.py mode=internal-slurm_cluster agent=dummy \
    tasks=longer_horizon parallel_tasks.task_names=all use_wandb=True

See the agents docs for the full SLURM + vLLM + W&B walkthrough.

OpenApps in action

demo.mp4

Reproducing the paper

main moves. The v1.0-paper tag pins the paper-era config surface used for the variation grid in arXiv:2511.20766 — check it out if you are reproducing or comparing against our setup:

git clone https://github.com/facebookresearch/OpenApps.git
cd OpenApps
git checkout v1.0-paper
uv sync

Everything after that tag is free to diverge. The first such change is the appearance refactor: the per-app appearance config group is replaced by a shared apps/theme (look) plus a per-app apps/<app>/layout (structure), so overrides written as apps/todo/appearance=dark_theme no longer resolve on main. See App variations in the docs for the current axes, which include a table mapping every old appearance value onto its replacement.

The tag is the config surface the paper used, not a byte-exact snapshot of the runs: it carries the app, task and harness fixes landed since publication, some of which move rewards (map tasks now match coordinates by ground distance, for example). Expect small differences from the published tables.

Contributing

We welcome pull requests with new features or issues via GitHub.

Development

uv sync --extra dev

To build docs:

mkdocs build
mkdocs serve

this will launch docs available at https://facebookresearch.github.io/OpenApps/

Testing

Run all tests via:

uv run -m pytest tests/

Attribution

Our apps are built on top of several excellent frameworks:

Some icons are have been designed using resources from Flaticon.com

Our work is licensed under CC-BY-NC, please refer to the LICENSE file in the top level directory.

Copyright © Meta Platforms, Inc. See the Terms of Use and Privacy Policy for this project.

Acknowledgements

  • A big thank you to Taj Gillin for implementing an MCP interface for OpenApps and Jiayu Wang for suggestions to improve our harness!
  • Another thank you to Yuval Kansal for fixing inconsistencies in task phrasings!

Featured In

Cite

@article{ullrich2025openapps0,
  title   = {OpenApps: Simulating Environment Variations to Measure UI-Agent Reliability},
  author  = {Karen Ullrich and Jingtong Su and Claudia Shi and Arjun Subramonian and Amir Bar and Ivan Evtimov and Nikolaos Tsilivis and Randall Balestriero and Julia Kempe and Mark Ibrahim},
  year    = {2025},
  journal = {arXiv preprint arXiv: 2511.20766}
}

About

An open source environment for digital agents.

Resources

Code of conduct

Contributing

Security policy

Stars

118 stars

Watchers

2 watching

Forks

Releases

Packages

Used by

Contributors

Languages