Skip to content

Latest commit

 

History

History

Folders and files

NameName
Last commit message
Last commit date

parent directory

..
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

README.md

Browser Commander (Python)

A universal browser automation library for Python that supports both Playwright and Selenium with a unified API. The key focus is on stoppable page triggers - ensuring automation logic is properly mounted/unmounted during page navigation.

Installation

The first PyPI release is pending trusted publisher registration. Until the PyPI project is available, install directly from this repository:

pip install "git+https://github.com/link-foundation/browser-commander.git#subdirectory=python"

You'll also need either Playwright or Selenium:

# With Playwright
pip install playwright
playwright install chromium

# Or with Selenium
pip install selenium

Core Concept: Page State Machine

Browser Commander manages the browser as a state machine with two states:

+------------------+                      +------------------+
|                  |   navigation start   |                  |
|  WORKING STATE   | -------------------> |  LOADING STATE   |
|  (action runs)   |                      |  (wait only)     |
|                  |   <-----------------  |                  |
+------------------+     page ready       +------------------+

LOADING STATE: Page is loading. Only waiting/tracking operations are allowed. No automation logic runs.

WORKING STATE: Page is fully loaded (30 seconds of network idle). Page triggers can safely interact with DOM.

Quick Start

import asyncio
from browser_commander import (
    launch_browser,
    make_browser_commander,
    make_url_condition,
    LaunchOptions,
)


async def main():
    # 1. Launch browser
    options = LaunchOptions(engine="playwright")
    result = await launch_browser(options)
    browser, page = result.browser, result.page

    # 2. Create commander
    commander = make_browser_commander(page=page, verbose=True)

    # 3. Register page trigger with condition and action
    async def example_action(ctx):
        print(f"Processing: {ctx['url']}")
        # Perform automation tasks
        await ctx["commander"].click_button(selector="button.submit")

    commander.page_trigger(
        {
            "name": "example-trigger",
            "condition": make_url_condition("*example.com*"),
            "action": example_action,
        }
    )

    # 4. Navigate - action auto-starts when page is ready
    await commander.goto(url="https://example.com")

    # 5. Cleanup
    await commander.destroy()
    await browser.close()


if __name__ == "__main__":
    asyncio.run(main())

URL Condition Helpers

The make_url_condition helper makes it easy to create URL matching conditions:

import re
from browser_commander import (
    make_url_condition,
    all_conditions,
    any_condition,
    not_condition,
)

# Exact URL match
make_url_condition("https://example.com/page")

# Contains substring (use * wildcards)
make_url_condition("*checkout*")  # URL contains 'checkout'
make_url_condition("*example.com*")  # URL contains 'example.com'

# Starts with / ends with
make_url_condition("/api/*")  # starts with '/api/'
make_url_condition("*.json")  # ends with '.json'

# Express-style route patterns
make_url_condition("/vacancy/:id")  # matches /vacancy/123
make_url_condition("https://hh.ru/vacancy/:vacancyId")

# RegExp
make_url_condition(re.compile(r"/product/\d+"))

# Custom function
make_url_condition(lambda url: url.startswith("https://"))

# Combine conditions
all_conditions(
    make_url_condition("*example.com*"),
    make_url_condition("*/checkout*"),
)  # Both must match

any_condition(
    make_url_condition("*/cart*"),
    make_url_condition("*/checkout*"),
)  # Either matches

not_condition(make_url_condition("*/admin*"))  # Negation

API Reference

launch_browser(options)

from browser_commander import launch_browser, LaunchOptions

options = LaunchOptions(
    engine="playwright",  # 'playwright' or 'selenium'
    headless=False,  # Run in headless mode
    user_data_dir="~/.browser-commander/playwright-data",
    slow_mo=150,  # Slow down operations (ms)
    verbose=False,  # Enable debug logging
    args=["--no-sandbox"],  # Custom Chrome args
    extra_args=["--lang=en-US"],  # Appended after legacy args
    ignore_default_args=["--disable-infobars"],  # Per-default opt-out
)
result = await launch_browser(options)
browser, page = result.browser, result.page

By default both launch APIs start the installed browser the way a person would: --user-data-dir=<fresh temporary profile> --remote-debugging-port=<reserved port> about:blank and nothing else (see Launch Command Line and Opt-In Restrictions). Switches the library used to add, such as --password-store=basic, are opt-in restrictions (restrictions=["legacy-defaults"] restores the old set). launch="engine" keeps the engine launcher, and ignore_default_args=True omits that launcher's own defaults.

Pass Playwright-compatible storage_state as a JSON file path or object to restore cookies and origin-scoped localStorage before navigation. The option works with launch_browser(), launch_real_browser(), and connect_browser() for both Playwright and Selenium:

from browser_commander import LaunchOptions, launch_browser, save_storage_state

result = await launch_browser(
    LaunchOptions(engine="playwright", storage_state="./session-state.json")
)
# After navigating and signing in, save the session for a later launch.
await save_storage_state(
    "playwright", result.browser, result.page, "./session-state.json"
)

Selenium saves cookies and localStorage for the current page origin. Playwright saves all origins available to its browser context. Treat saved state as a secret because cookies may contain login credentials.

connect_browser(options)

Attach Playwright or Selenium to an already-running Chrome-family browser over CDP. Exactly one HTTP cdp_endpoint or browser ws_endpoint is required:

from browser_commander import ConnectOptions, connect_browser

result = await connect_browser(
    ConnectOptions(
        engine="playwright",  # or "selenium"
        cdp_endpoint="http://127.0.0.1:9222",
        seed_cookies=[
            {"name": "session", "value": "saved", "url": "https://example.com"}
        ],
    )
)
browser, page = result.browser, result.page

The returned page is the raw engine page/driver and can be passed directly to make_browser_commander(). For Chrome 136 and newer, start the browser with a non-default --user-data-dir; remote debugging is intentionally disabled for the default Chrome profile. Cookie seeding uses only values supplied by the caller and does not read the default profile. Because connection is attach-only, it cannot retrofit launch flags. Start the external process with the documented defaults and a dedicated debugging profile, or use launch_real_browser(). The installed-browser import below provides an explicit local way to obtain those values.

Installed browser cookies

Discover profiles and read cookies in the exact Playwright/Puppeteer cookie shape:

from browser_commander import (
    BrowserCookieCacheOptions,
    BrowserCookieReadOptions,
    list_browser_profiles,
    read_browser_cookies,
)

print(list_browser_profiles("chrome"))

cookies = read_browser_cookies(
    BrowserCookieReadOptions(
        browser="chrome",  # chrome, edge, brave, chromium, or firefox
        profile="Default",  # optional; defaults to the selected browser profile
        domain_filter="example.com",
        cache=BrowserCookieCacheOptions(ttl_minutes=60),
    )
)

Each cookie contains name, value, domain, path, expires, httpOnly, secure, and sameSite. The helper runs only when explicitly called and never sends cookie data anywhere. Decrypted result and derived-key files default to ~/.browser-commander/cookie-cache/ with owner-only permissions. A process lock and the default 60-minute TTL keep Keychain, libsecret/KWallet, or DPAPI access to at most one read across concurrent and repeated processes. Set refresh=True to force a new credential read, customize dir/ttl_minutes, or set cache=False to opt out.

Browser family macOS Linux Windows
Chrome, Edge, Brave, Chromium Keychain + AES-128-CBC (v10/v11) libsecret/KWallet + AES-128-CBC (v11), or the Chromium v10 fallback key DPAPI-protected AES-256-GCM key (v10/v11)
Firefox cookies.sqlite cookies.sqlite cookies.sqlite

Chromium database-version-24 domain hashes and its 1601-based timestamps are handled automatically. Current Windows Chromium may use app-bound v20 encryption, which intentionally requires the browser's privileged service and cannot be decrypted by an ordinary external process. The helper reports this boundary; use a browser-supported export or saved Browser Commander storage state for those cookies. ignore_decryption_errors=True returns any remaining decryptable cookies. Treat imported cookies like passwords: keep cache paths private, use a short TTL, never commit them, and seed only a dedicated profile.

launch_real_browser(options)

Discover and start a genuine installed Chrome, Edge, Brave, or Chromium with a dedicated profile, wait for its loopback CDP endpoint, and attach with Playwright or Selenium:

from browser_commander import RealBrowserOptions, launch_real_browser

result = await launch_real_browser(
    RealBrowserOptions(
        engine="playwright",  # or "selenium"
        channel="chrome",  # chrome, msedge, brave, or chromium
        user_data_dir="/tmp/browser-commander-profile",
        extra_args=["--lang=en-US"],
        ignore_default_args=["--disable-infobars"],
        seed_cookies=[
            {"name": "session", "value": "saved", "url": "https://example.com"}
        ],
    )
)

browser, page = result.browser, result.page
print(result.cdp_endpoint, result.executable_path)

The helper also supports beta/dev/canary channels and an explicit executable_path. It rejects known default browser profiles and prevents custom arguments from overriding its loopback address, debugging port, or profile. The returned browser_process can be terminated explicitly after closing the browser. launch_and_connect_real_browser() is an alias. The remote-debugging address, port, and profile remain managed; headless mode is opt-in and uses --headless=new. The older args field remains an append-only compatibility alias.

The color_scheme option emulates prefers-color-scheme at launch time:

options = LaunchOptions(
    engine="playwright",
    color_scheme="dark",  # 'light', 'dark', or 'no-preference'
)
result = await launch_browser(options)

emulate_media(page, engine, color_scheme)

Emulate media features (e.g. prefers-color-scheme) for a page after launch:

from browser_commander import emulate_media

# Set dark mode
await emulate_media(page, engine="playwright", color_scheme="dark")

# Set light mode
await emulate_media(page, engine="playwright", color_scheme="light")

# Reset to system default
await emulate_media(page, engine="playwright", color_scheme=None)

Works with both Playwright (page.emulate_media) and Selenium (via CDP Emulation.setEmulatedMedia).

make_browser_commander(page, options)

from browser_commander import make_browser_commander

commander = make_browser_commander(
    page=page,  # Required: Playwright/Selenium page
    verbose=False,  # Enable debug logging
    enable_network_tracking=True,  # Track HTTP requests
    enable_navigation_manager=True,  # Enable navigation events
)

commander.goto(url, options)

result = await commander.goto(
    url="https://example.com",
    wait_until="domcontentloaded",
    timeout=60000,
)
print(f"Navigated: {result['navigated']}, URL: {result['actual_url']}")

commander.click_button(selector, options)

result = await commander.click_button(
    selector="button.submit",
    scroll_into_view=True,
    wait_after_click=1000,
)
print(f"Clicked: {result['clicked']}, Navigated: {result['navigated']}")

commander.fill_text_area(selector, text, options)

result = await commander.fill_text_area(
    selector="textarea.message",
    text="Hello world",
    check_empty=True,
)
print(f"Filled: {result['filled']}, Value: {result['actual_value']}")

Keyboard Interactions

from browser_commander import press_key, type_text, key_down, key_up

# Press a single key
await press_key(page=page, key="Escape", engine="playwright")
await press_key(page=page, key="Enter", engine="playwright")
await press_key(page=page, key="Tab", engine="playwright")

# Type text
await type_text(page=page, text="Hello World", engine="playwright")

# Hold and release modifier keys
await key_down(page=page, key="Control", engine="playwright")
await key_up(page=page, key="Control", engine="playwright")

Element Selection Methods

# Query single element
element = await commander.query_selector("button.submit")

# Query all matching elements
elements = await commander.query_selector_all(".list-item")

# Wait for selector
found = await commander.wait_for_selector("button.submit", visible=True, timeout=5000)

# Find by text content
selector = commander.find_by_text("Click me", selector="button", exact=False)

Element Inspection Methods

# Check visibility
is_vis = await commander.is_visible("button.submit")

# Check if enabled
is_en = await commander.is_enabled("button.submit")

# Count matching elements
count = await commander.count(".list-item")

# Get text content
text = await commander.text_content(".heading")

# Get input value
value = await commander.input_value("input.email")

# Get attribute
href = await commander.get_attribute("a.link", "href")

Wait and Evaluate Methods

# Wait for time
result = await commander.wait(ms=1000, reason="waiting for animation")
print(f"Completed: {result['completed']}, Aborted: {result['aborted']}")

# Evaluate JavaScript
result = await commander.evaluate("() => document.title")

# Safe evaluate (doesn't throw on navigation)
result = await commander.safe_evaluate(
    fn="() => document.title",
    default_value="Unknown",
)
print(f"Success: {result['success']}, Value: {result['value']}")

Managed Downloads

A download that only exists while the browser is open is not a download. Ask for downloads at any entry point - launch_browser(), connect_browser(), launch_real_browser() or commander.configure_downloads() - and the manager owns the file from then on:

from browser_commander import LaunchOptions, launch_browser, make_browser_commander

result = await launch_browser(
    LaunchOptions(
        engine="playwright",
        downloads={"directory": "/tmp/reports", "conflict": "rename"},
    )
)
commander = make_browser_commander(page=result.page)

# capture() starts listening before the action runs, so a download that
# finishes in 5ms cannot slip past the registration.
artifact = await result.downloads.capture(
    action=lambda: commander.click_button(selector="#export"),
    filename="q3-report.pdf",  # the page's UUID name gets the caller's name
    timeout=30000,
)
print(artifact.path, artifact.bytes, artifact.checksum)

await commander.destroy()
await result.browser.close()
# The file is still there: it outlives the page, the context and the browser.

downloads.on(DownloadEvent.COMPLETED, ...) observes every download in the session, including one a person started by hand in a visible browser, and each download is reported exactly once whether it was captured or merely observed. A failed or cancelled download raises the failure rather than returning a path. Playwright listens on a browser-wide CDP session; Selenium, which has no download events, watches the staging directory instead.

Native Extension Relay

attach_via_extension(RelayOptions(...)) hosts the companion extension's relay directly in Python, without the JavaScript CLI. It exposes typed tab results, CDP sessions and events. Install the extension extra and select the bundled extension_directory() in Chrome's Load unpacked dialog. Configure allowed_extension_ids to restrict the accepted installed extension. Native extension relay has startup, cancellation, resource limits and complete Python/Rust examples.

Live Profile Snapshots

Copy a selected Chromium profile while its source browser stays open, then launch the copy through Playwright or Selenium:

from browser_commander import RealBrowserOptions, SnapshotOptions, launch_snapshot

copy = await launch_snapshot(
    SnapshotOptions(browser="chrome", profile="Profile 1"),
    RealBrowserOptions(engine="playwright"),
)
try:
    await copy.page.goto("https://example.com")
    print(copy.snapshot["copied"])
finally:
    await copy.close()

user_data_dir in SnapshotOptions selects an explicit source root. SQLite backups retain committed WAL data; caches, locks and open-tab sessions are excluded and reported. Closing the browser deletes its copy. The original profile remains untouched. snapshot_user_data_dir(browser="chrome", ...) returns the same report without launching; callers own that returned directory.

Portable Traces

A trace is one versioned directory - manifest, ordered NDJSON timeline, per-checkpoint DOM snapshots and the mutation batches between them - and Python writes the same bundle JavaScript does:

from browser_commander.traces import write_trace_viewer

trace = await commander.start_trace(
    output="/tmp/traces/checkout",
    mode="continuous",
    links={"output": "/tmp/traces/checkout.lino"},
)
await commander.goto("https://example.com/cart")
await trace.checkpoint("cart")
stopped = await trace.stop()
write_trace_viewer(stopped["path"])  # viewer.html, readable offline

Navigations, interactions, console messages, page errors, failed requests, dialogs and downloads share one ordered timeline; password fields, [data-private] controls, credential headers and token query parameters are redacted before anything is written. To keep a trace only when something went wrong, wrap the work in traced():

from browser_commander.traces import traced

async with traced(commander, output="artifacts/checkout.bc-trace"):
    await commander.click_button("#pay")  # kept, with the viewer, only if this raises

A bundle recorded here or by a JavaScript run reads back the same way:

from browser_commander import diff_control_state, read_trace

trace = read_trace("/tmp/traces/checkout")
print(trace.manifest["schemaVersion"], trace.truncated)

for event in trace.events:
    print(event["at"], event["kind"])

for change in diff_control_state(trace.state(1), trace.state(2)):
    print(change.path, change.change, change.before, "->", change.after)

Typed Puppeteer

Puppeteer exists only for Node.js, so browser_commander.puppeteer drives it through the JavaScript CLI's serve --stdio bridge. Every Puppeteer class and interface has a Python class in the same hierarchy, with an async def for every method and getter, generated by scripts/generate-puppeteer-bindings.mjs from the lib/types.d.ts that puppeteer-core ships:

from browser_commander.puppeteer import JsFunction, PuppeteerBridge

async with await PuppeteerBridge.launch() as bridge:
    browser = await (await bridge.puppeteer()).launch({"headless": True})
    page = await browser.new_page()
    await page.goto("https://example.com")
    print(await page.title())
    print(await page.evaluate(JsFunction("(a, b) => a + b"), 1, 2))
    console = await page.subscribe("console")
    await browser.close()

Results come back as their Python types: handles as Page, ElementHandle and so on, binary data as bytes. Errors are BridgeError, whose is_timeout matches Puppeteer's TimeoutError. The bridge needs Node.js, the JavaScript CLI (BROWSER_COMMANDER_JS_CLI, or the browser-commander npm package in node_modules) and puppeteer-core or puppeteer where Node resolves it.

Truthful Click Results

click_element() reports what was observed, not what was attempted:

result = await click_element(page, engine, log, "#submit")

result.status  # 'succeeded' | 'failed' | 'timed_out' | 'interrupted' | 'unverified'
result.effect  # 'confirmed' | 'not-observed' | 'contradicted'
result.evidence  # why status and effect say what they say
result.clicked  # still here: whether the click reached the element
result.verified  # still here, now derived from effect == 'confirmed'

The scroll axis ('auto', 'preserve', 'none') replaces the deprecated no_auto_scroll flag. scroll='none' never scrolls: on an engine that cannot deliver a click without scrolling it raises ScrollConstraintError naming the alternatives, rather than scrolling the page and reporting success.

commander.destroy()

await commander.destroy()  # Stop actions, cleanup

Best Practices

1. Always Cleanup Resources

async def main():
    result = await launch_browser(options)
    browser, page = result.browser, result.page
    commander = make_browser_commander(page=page)

    try:
        # Your automation code
        await commander.goto(url="https://example.com")
    finally:
        await commander.destroy()
        await browser.close()

2. Use Verbose Mode for Debugging

commander = make_browser_commander(page=page, verbose=True)

3. Handle Navigation-Aware Operations

# Wait for page to be fully ready after navigation
await commander.wait_for_page_ready(timeout=30000)

# Check if should abort current operation
if commander.should_abort():
    return  # Navigation detected, stop current action

Extensibility / Escape Hatch

browser-commander cannot anticipate every browser API. When you need an API that is not yet supported, you can access the raw underlying engine objects directly as an official extensibility escape hatch.

Using commander.page for engine-specific APIs

make_browser_commander exposes commander.page — this is the raw Playwright or Selenium page object, not a wrapper. Use it directly for APIs browser-commander doesn't yet support:

from browser_commander import launch_browser, make_browser_commander, LaunchOptions

options = LaunchOptions(engine="playwright")
result = await launch_browser(options)
browser, page = result.browser, result.page

commander = make_browser_commander(page=page)

# Access engine-specific API via commander.page
# Example: PDF generation (issue #35)
pdf_buffer = await commander.page.pdf(
    format="A4",
    print_background=True,
)

# Example: Color scheme emulation (issue #36)
await commander.page.emulate_media(color_scheme="dark")

# Example: Keyboard interactions (issue #37)
await commander.page.keyboard.press("Escape")

# Example: Dialog handling (issue #38)


async def handle_dialog(dialog):
    await dialog.dismiss()


commander.page.on("dialog", handle_dialog)

Using launch_browser raw return values

launch_browser() returns the raw browser and page objects from the underlying engine. You can use these directly:

result = await launch_browser(options)
browser = result.browser
page = result.page

# Use raw page directly for engine-specific APIs
await page.pdf(format="A4")

# Or create a commander for the unified API
commander = make_browser_commander(page=page)

No more _page hacks

If you previously used page._page or page to access the raw page, replace it with commander.page:

# BEFORE (fragile hack):
raw_page = getattr(page, "_page", page)
await raw_page.pdf(format="A4")

# AFTER (official API):
await commander.page.pdf(format="A4")

This is the official extensibility mechanism while awaiting browser-commander to add first-class support for these APIs. Please report missing APIs so they can be added.

Architecture

The Python implementation follows the same architecture as the JavaScript version:

  • Core Module: Constants, logger, engine detection, navigation safety
  • Browser Module: Launcher, navigation management
  • Elements Module: Selectors, visibility, content extraction
  • Interactions Module: Click, fill, scroll operations
  • Utilities Module: Wait, URL helpers
  • High-Level Module: Universal logic, page triggers

License

UNLICENSE