diff --git a/.github/workflows/python-reference.yml b/.github/workflows/python-reference.yml index c8bb612..f31b492 100644 --- a/.github/workflows/python-reference.yml +++ b/.github/workflows/python-reference.yml @@ -2,10 +2,10 @@ name: python-reference # Regenerate the Python SDK API reference page from the lightpanda-python main # branch and open a pull request when the output changed. The page is -# src/content/reference/python-api.mdx, written by +# src/content/reference/python.mdx, written by # scripts/generate-python-reference.py with pdoc's Python API, so it renders # like any other docs page and goes live at -# https://lightpanda.io/docs/reference/python-api once the website bumps its +# https://lightpanda.io/docs/reference/python once the website bumps its # docs submodule like any other docs change. The package imports without a # browser binary, so none is needed here. Runs daily to pick up new package # changes, by hand, or as a smoke run when the generator itself changes. @@ -44,7 +44,7 @@ jobs: --with "git+https://github.com/lightpanda-io/lightpanda-python@$sha" \ python scripts/generate-python-reference.py echo "sha=${sha::7}" >> "$GITHUB_OUTPUT" - git status --short src/content/reference/python-api.mdx + git status --short src/content/reference/python.mdx - name: Open a pull request if the reference changed if: github.event_name != 'pull_request' @@ -52,18 +52,18 @@ jobs: GH_TOKEN: ${{ github.token }} SHA: ${{ steps.generate.outputs.sha }} run: | - if [ -z "$(git status --porcelain src/content/reference/python-api.mdx)" ]; then + if [ -z "$(git status --porcelain src/content/reference/python.mdx)" ]; then echo "reference already matches lightpanda-python $SHA" exit 0 fi git config user.name "github-actions[bot]" git config user.email "41898282+github-actions[bot]@users.noreply.github.com" git checkout -B python-reference - git add src/content/reference/python-api.mdx + git add src/content/reference/python.mdx git commit -m "Update the Python SDK API reference to lightpanda-python $SHA" git push --force origin python-reference if [ -z "$(gh pr list --head python-reference --state open --json number -q '.[].number')" ]; then gh pr create --base main --head python-reference \ --title "Update the Python SDK API reference to lightpanda-python $SHA" \ - --body "Regenerated \`src/content/reference/python-api.mdx\` from lightpanda-python $SHA. Served at https://lightpanda.io/docs/reference/python-api after the website's submodule bump." + --body "Regenerated \`src/content/reference/python.mdx\` from lightpanda-python $SHA. Served at https://lightpanda.io/docs/reference/python after the website's submodule bump." fi diff --git a/redirects.mjs b/redirects.mjs index 0594d4e..92fefc9 100644 --- a/redirects.mjs +++ b/redirects.mjs @@ -10,7 +10,8 @@ export const basePath = '/docs' export const redirects = { - '/python': '/reference/python-api', + '/python': '/reference/python', + '/reference/python-api': '/reference/python', '/quickstart/installation-and-setup': '/quickstart', '/quickstart/your-first-test': '/quickstart', '/quickstart/build-your-first-extraction-script': '/quickstart', diff --git a/scripts/generate-python-reference.py b/scripts/generate-python-reference.py index 7554688..b013927 100644 --- a/scripts/generate-python-reference.py +++ b/scripts/generate-python-reference.py @@ -1,11 +1,9 @@ -"""Generate the Python SDK API reference page from the lightpanda package. - -The hand-written src/content/reference/python.mdx explains how the package fits -together; this script writes the exhaustive companion page, -src/content/reference/python-api.mdx, by walking the installed `lightpanda` -package with pdoc's Python API and emitting one MDX section per public class, -method, property, function and exception, with the signatures and docstrings -shipped in the code. Emitting MDX instead of pdoc's own HTML keeps the page +"""Generate the Python SDK reference page from the lightpanda package. + +This script writes src/content/reference/python.mdx by walking the installed +`lightpanda` package with pdoc's Python API and emitting one MDX section per +public class, method, property, function and exception, with the signatures +and docstrings shipped in the code. Emitting MDX instead of pdoc's own HTML keeps the page inside the Nextra site: sidebar, search, dark mode and deep links all work as on any other page. @@ -34,11 +32,11 @@ import lightpanda ROOT = Path(__file__).resolve().parent.parent -DEFAULT_OUT = ROOT / "src" / "content" / "reference" / "python-api.mdx" +DEFAULT_OUT = ROOT / "src" / "content" / "reference" / "python.mdx" FRONTMATTER = """--- -title: Python API -description: Generated reference of every public class, method, property and exception in the lightpanda Python package, with the signatures and docstrings shipped in the code. +title: Python SDK +description: Reference of every public class, method, property and exception in the lightpanda Python package, generated from the signatures and docstrings shipped in the code. --- """ @@ -50,12 +48,31 @@ INTRO = ( "Every public class, method, property and exception of the " "[`lightpanda` package](https://pypi.org/project/lightpanda/), with the signatures and " - "docstrings shipped in the code. See [Python SDK](/reference/python) for a curated " - "overview of the same API and [Use the Python SDK](/guides/use-python) for practical " - "documentation. Every sync class has an asyncio twin with the same methods, " + "docstrings shipped in the code. See [Use the Python SDK](/guides/use-python) for a " + "practical walkthrough. Every sync class has an asyncio twin with the same methods, " "awaitable; the async sections below list only what the twin adds." ) +CONVENTIONS = [ + "Browser actions are keyword-only methods on [`Session`](#session) and " + "[`AsyncSession`](#asyncsession), named in snake_case after the browser's own action " + "names: the `waitForSelector` action is `wait_for_selector`, and its `backendNodeId` " + "argument is `backend_node_id`.", + "Where a method accepts both `selector` and `backend_node_id`, pass one of the two. " + "`selector` is preferred for reproducibility and wins when both are given; " + "`backend_node_id` takes the values returned by [`tree`](#session-tree), " + "[`links`](#session-links) or [`find_element`](#session-find-element).", +] + +# Fixed paragraphs shown under a class heading, after its docstring. +CLASS_NOTES = { + "Session": ( + "[`call`](#session-call) is the escape hatch that takes the action and argument " + "names exactly as the browser declares them. A failed action raises " + "[`ToolError`](#toolerror)." + ), +} + FENCE_RE = re.compile(r"^\s*```") CODE_SPAN_RE = re.compile(r"(`+)(.+?)\1", re.DOTALL) MODULE_PREFIX_RE = re.compile(r"\blightpanda\.\w+\.") @@ -289,6 +306,7 @@ def class_code(cls: pdoc.doc.Class) -> str: def emit_class(page: Page, module: pdoc.doc.Module, cls: pdoc.doc.Class, links: dict[str, str]) -> None: page.heading(2, cls.name, slug(cls.name)) page.para(render_docstring(cls, links)) + page.para(CLASS_NOTES.get(cls.name, "")) page.fence(class_code(cls)) init = cls.members.get("__init__") if isinstance(init, pdoc.doc.Function) and "__init__" in vars(cls.obj): @@ -382,10 +400,11 @@ def generate() -> str: page.lines.append(FRONTMATTER.rstrip()) page.lines.append(BANNER) page.lines.append("") - page.lines.append("# Python API") + page.lines.append("# Python SDK") page.lines.append("") page.para(INTRO) - page.para(render_docstring(module, links)) + for paragraph in CONVENTIONS: + page.para(paragraph) exceptions: list[pdoc.doc.Class] = [] for name in names: diff --git a/src/content/guides/use-python.mdx b/src/content/guides/use-python.mdx index b48af54..19396a7 100644 --- a/src/content/guides/use-python.mdx +++ b/src/content/guides/use-python.mdx @@ -164,7 +164,7 @@ asyncio.run(main()) Every browser action is a `Session` method, typed and documented in your IDE, with the action and its arguments in snake_case (`wait_for_selector`, `backend_node_id`). -Find every method's arguments in the [Python SDK reference](/reference/python), or browse the generated [Python API](/reference/python-api) reference. +Find every method's signature and docstring in the [Python SDK reference](/reference/python). ## Replay a saved script diff --git a/src/content/reference/_meta.ts b/src/content/reference/_meta.ts index 7f6d572..ea6c7c1 100644 --- a/src/content/reference/_meta.ts +++ b/src/content/reference/_meta.ts @@ -11,7 +11,6 @@ const meta: MetaRecord = { 'mcp-tools': 'MCP tools', pandascript: 'PandaScript', python: 'Python SDK', - 'python-api': 'Python API', } export default meta diff --git a/src/content/reference/python-api.mdx b/src/content/reference/python-api.mdx deleted file mode 100644 index cb3ca3f..0000000 --- a/src/content/reference/python-api.mdx +++ /dev/null @@ -1,848 +0,0 @@ ---- -title: Python API -description: Generated reference of every public class, method, property and exception in the lightpanda Python package, with the signatures and docstrings shipped in the code. ---- -{/* Generated by scripts/generate-python-reference.py from lightpanda-python main. Do not edit; rerun the script or wait for the python-reference workflow. */} - -# Python API - -Every public class, method, property and exception of the [`lightpanda` package](https://pypi.org/project/lightpanda/), with the signatures and docstrings shipped in the code. See [Python SDK](/reference/python) for a curated overview of the same API and [Use the Python SDK](/guides/use-python) for practical documentation. Every sync class has an asyncio twin with the same methods, awaitable; the async sections below list only what the twin adds. - -Lightpanda for Python: a lightweight headless browser. - -```python -from lightpanda import Browser - -with Browser() as b: - page = b.new_session() - page.goto(url="https://example.com") - data = page.extract(schema={"title": "h1"}) -``` - -The same API is available for asyncio: - -```python -from lightpanda import AsyncBrowser - -async with AsyncBrowser() as b: - page = await b.new_session() - await page.goto(url="https://example.com") - data = await page.extract(schema={"title": "h1"}) -``` - -For Playwright or Puppeteer code, [``CDPServer``](#cdpserver) runs the browser's own -Chrome DevTools Protocol server and hands you the endpoint to connect to -(see its docs for an example). For Selenium, [``BiDiServer``](#bidiserver) serves WebDriver -BiDi the same way and hands you the ``command_executor`` URL. - -## Browser [#browser] - -A lightpanda browser process. Spawns the bundled binary on first use. - -Not fork-inheritable: after ``os.fork()``/``multiprocessing``, create a -fresh Browser in the child. - -```python -class Browser( - binary: str | os.PathLike | None = None, - env: dict[str, str] | None = None, - timeout: float = 300.0, - verbose: bool = False, - args: Sequence[str] = () -) -``` - -``args`` are extra CLI flags for the spawned browser process -(e.g. ``["--http-cache-dir", path]`` or cookie flags). - -Usable as a context manager (`with`). - -### Browser.tools [#browser-tools] - -```python -tools: dict[str, dict] -``` - -*Property.* Tool name → \{description, schema\}, as reported by the browser. - -### Browser.new_session [#browser-new-session] - -```python -def new_session(self) -> Session -``` - -### Browser.close [#browser-close] - -```python -def close(self) -> None -``` - -## Session [#session] - -One isolated browsing context (own page, cookies, memory). - -Do not construct directly — use [`Browser.new_session()`](#browser-new-session). - -```python -class Session(browser: Browser, session_id: str) -``` - -Usable as a context manager (`with`). - -### Session.id [#session-id] - -```python -id: str -``` - -*Property.* - -### Session.call [#session-call] - -```python -def call(self, tool: str, **kwargs) -``` - -Invoke a browser tool by name. The generated methods route here. - -Returns parsed JSON for JSON-carrying tools, ``bytes`` for image -results ([``screenshot``](#session-screenshot) without ``path``), otherwise the result text. - -### Session.close [#session-close] - -```python -def close(self) -> None -``` - -### Session.click [#session-click] - -```python -def click( - self, - *, - selector: str | None = None, - backend_node_id: int | None = None -) -> Any -``` - -Click on an interactive element. Provide either a CSS selector (preferred for reproducibility) or a backendNodeId. Returns the current page URL and title after the click. - -### Session.console_logs [#session-console-logs] - -```python -def console_logs(self) -> Any -``` - -Get buffered console.log/warn/error messages from the current page. Returns all messages since last call and clears the buffer. - -### Session.detect_forms [#session-detect-forms] - -```python -def detect_forms(self, *, url: str | None = None, timeout: int | None = None) -> Any -``` - -Detect all forms on the page and return their structure including fields, types, and required status. If a url is provided, it navigates to that url first. - -### Session.evaluate [#session-evaluate] - -```python -def evaluate( - self, - *, - script: str, - url: str | None = None, - timeout: int | None = None, - save: str | None = None -) -> Any -``` - -Evaluate JavaScript in the current page context — an escape hatch for page-side logic the dedicated tools can't express; prefer [`extract`](#session-extract) for data and click/fill/etc. for actions. It runs in the page, so it cannot see the agent script's variables or builtins — interpolate any value into the `script` string. A bare trailing expression yields its value; top-level `await` and `return` are supported (the body then runs as an async function, so use `return` to produce a value). Objects and arrays return as JSON, so no `JSON.stringify` is needed. If a url is provided, it navigates there first. The `globalThis.lp` object exposes a Session-scoped bridge store: values written via `lp.foo = ...` auto-sync at end of evaluate, surviving navigation; values previously set via `/extract save=` or `/evaluate save=` appear as `lp.`. - -### Session.extract [#session-extract] - -```python -def extract(self, *, schema: str | dict | list, save: str | None = None) -> Any -``` - -Extract structured data from the current page (navigate first). `schema` is a JSON object (passed as a string) mapping output field names to CSS-selector specs. It is NOT a JSON Schema — no "type"/"properties" wrappers; the keys ARE your output fields. Value shapes: - -```text -"" → first match's text (trimmed; null if no match) -[""] → every match's text (string[]) -{"selector":"","attr":""} → first match's attribute value (href/src resolved to absolute URLs) -[{"selector":"","attr":""}] → every match's attribute (string[]) -[{"selector":"","fields":{…}}] → one object per match; field selectors resolve relative to that match and accept any shape above ("" = the match's own text; nest arrays for per-item sub-lists) -``` - -Add "limit": N inside any array's object spec to cap matches. -Every extracted value is a string or null — parse numbers downstream. An empty array is a valid result, but if ALL top-level keys miss, the call errors: inspect the page (tree/markdown) and retry with corrected selectors. -Finish data tasks with extract — it is the only read recorded as a replayable `extract(...)` script call; answers lifted from [`markdown`](#session-markdown) text in chat are not. - -Examples (schema → result): - -```text -{"karma": "#karma"} → {"karma":"42"} -{"items": [".story .title"]} → {"items":["Title 1","Title 2"]} -{"top3": [{"selector":".story .title","limit":3}]} → {"top3":["A","B","C"]} -{"links": [{"selector":"a.title","attr":"href"}]} → {"links":["https://site/a","https://site/b"]} -{"stories": [{"selector":".athing","fields":{"title":".titleline","rank":".rank"}}]} → {"stories":[{"title":"Foo","rank":"1"}]} -``` - -### Session.fill [#session-fill] - -```python -def fill( - self, - *, - value: str, - selector: str | None = None, - backend_node_id: int | None = None -) -> Any -``` - -Fill text into an input element. Provide either a CSS selector (preferred for reproducibility) or a backendNodeId. - -### Session.find_element [#session-find-element] - -```python -def find_element(self, *, role: str | None = None, name: str | None = None) -> Any -``` - -Find interactive elements by role and/or accessible name. Returns matching elements with their backend node IDs. Useful for locating specific elements without parsing the full semantic tree. - -### Session.get_cookies [#session-get-cookies] - -```python -def get_cookies(self, *, url: str | None = None, all: bool | None = None) -> Any -``` - -Cookies stored in the browser. Defaults to cookies whose domain matches the current page's host. Pass `url=` to filter for another host, or `all=true` to dump every cookie regardless of host. Useful for debugging authentication and session state. - -### Session.get_env [#session-get-env] - -```python -def get_env(self, *, name: str | None = None) -> Any -``` - -With `name`: read an LP_* env var (other namespaces report as not set) — for non-secret config only (base URLs, flags). Without `name`: list LP_* names that are set (no values) — safe credential discovery. For secrets, pass `$LP_*` placeholders in tool args; never request a credential by name (the value would land in your context). - -### Session.get_url [#session-get-url] - -```python -def get_url(self) -> Any -``` - -Current page URL. The browser may already have a page loaded (command, replayed script) not visible in this conversation — call this before assuming nothing is loaded when the user references the current page/site. Also useful to verify a navigation or detect a redirect. - -### Session.goto [#session-goto] - -```python -def goto( - self, - *, - url: str, - timeout: int | None = None, - wait_until: str | None = None -) -> Any -``` - -Navigate to a specified URL and load the page in memory so it can be reused later for info extraction. - -### Session.hover [#session-hover] - -```python -def hover( - self, - *, - selector: str | None = None, - backend_node_id: int | None = None -) -> Any -``` - -Hover over an element, triggering mouseover and mouseenter events. Provide either a CSS selector (preferred for reproducibility) or a backendNodeId. Useful for menus, tooltips, and hover states. - -### Session.html [#session-html] - -```python -def html( - self, - *, - selector: str | None = None, - backend_node_id: int | None = None, - max_bytes: int | None = None, - strip: dict | None = None, - url: str | None = None, - timeout: int | None = None -) -> Any -``` - -Raw HTML for the document or, with `selector`/`backendNodeId`, a single node's outerHTML. Verbose; use only when you need attributes that markdown discards. - -### Session.interactive_elements [#session-interactive-elements] - -```python -def interactive_elements(self, *, url: str | None = None, timeout: int | None = None) -> Any -``` - -Extract interactive elements from the opened page. If a url is provided, it navigates to that url first. - -### Session.links [#session-links] - -```python -def links( - self, - *, - limit: int | None = None, - url: str | None = None, - timeout: int | None = None -) -> Any -``` - -Extract the visible links in the opened page as JSON objects with `text` (anchor text, falling back to aria-label/title/image alt), `href` (resolved URL), and `backendNodeId` (pass to click/nodeDetails). One entry per href; hidden links are omitted. If a url is provided, it navigates to that url first. - -### Session.markdown [#session-markdown] - -```python -def markdown( - self, - *, - selector: str | None = None, - backend_node_id: int | None = None, - max_bytes: int | None = None, - url: str | None = None, - timeout: int | None = None -) -> Any -``` - -Render the page (or a subtree) as markdown. Scope with `selector` or `backendNodeId` to read just the relevant region — full-page markdown is the last resort. Use `maxBytes` to cap long pages. - -### Session.node_details [#session-node-details] - -```python -def node_details(self, *, backend_node_id: int) -> Any -``` - -Details for a node by backendNodeId: a ready-to-use CSS `selector` that resolves to the node (the first match, as click/fill resolve it), plus tag, role, name, interactivity, disabled, value, input type, placeholder, href, id, class, checked, select options. The canonical way to turn a tree backendNodeId into a CSS selector. - -### Session.press [#session-press] - -```python -def press( - self, - *, - key: str, - selector: str | None = None, - backend_node_id: int | None = None -) -> Any -``` - -Press a keyboard key, dispatching keydown and keyup events. Use key names like 'Enter', 'Tab', 'Escape', 'ArrowDown', 'Backspace', or single characters like 'a', '1'. Common shorthand is normalized: 'enter'/'return' → 'Enter', 'esc' → 'Escape', 'up'/'down'/'left'/'right' → 'Arrow*', 'space' → ' '. Pressing 'Enter' on a form input or submit button triggers implicit form submission. - -### Session.screenshot [#session-screenshot] - -```python -def screenshot( - self, - *, - path: str | None = None, - selector: str | None = None, - backend_node_id: int | None = None, - full_page: bool | None = None, - url: str | None = None, - timeout: int | None = None -) -> Any -``` - -Render the page, or one node, as a PNG: the text layout Lightpanda computes, not a pixel-accurate browser rendering (no images, fonts or CSS colours). With `path`, writes the file at full size and returns its location; without it, returns the image inline where the client can display one, at most 1280px wide and 4096px tall. Use it to see spatial layout; read content with [`markdown`](#session-markdown)/[`tree`](#session-tree). - -### Session.scroll [#session-scroll] - -```python -def scroll( - self, - *, - backend_node_id: int | None = None, - x: int | None = None, - y: int | None = None -) -> Any -``` - -Scroll the page or a specific element. Returns the scroll position and current page URL and title. - -### Session.search [#session-search] - -```python -def search(self, *, query: str, timeout: int | None = None) -> Any -``` - -Run a web search and return results as markdown: a numbered list of \{title, url, snippet\}. Search tries brave, tavily, exa, then keenable in order, each when its API key (BRAVE_API_KEY, TAVILY_API_KEY, EXA_API_KEY or KEENABLE_API_KEY) is set; keenable also works without a key through its public endpoint (rate-limited per client IP). Prefer this over goto-ing google.com/search directly (Google blocks the browser on User-Agent/TLS). The browser does not navigate — to open a result, use [`goto`](#session-goto) with its URL. - -### Session.select_option [#session-select-option] - -```python -def select_option( - self, - *, - value: str, - selector: str | None = None, - backend_node_id: int | None = None -) -> Any -``` - -Select an option in a <select> dropdown element by its value. Provide either a CSS selector (preferred for reproducibility) or a backendNodeId. Dispatches input and change events. - -### Session.set_checked [#session-set-checked] - -```python -def set_checked( - self, - *, - checked: bool, - selector: str | None = None, - backend_node_id: int | None = None -) -> Any -``` - -Check or uncheck a checkbox or radio button. Provide either a CSS selector (preferred for reproducibility) or a backendNodeId. Dispatches input, change, and click events. - -### Session.structured_data [#session-structured-data] - -```python -def structured_data(self, *, url: str | None = None, timeout: int | None = None) -> Any -``` - -Extract structured data (like JSON-LD, OpenGraph, etc) from the opened page. If a url is provided, it navigates to that url first. - -### Session.tree [#session-tree] - -```python -def tree( - self, - *, - url: str | None = None, - timeout: int | None = None, - backend_node_id: int | None = None, - max_depth: int | None = None -) -> Any -``` - -Simplified semantic DOM tree (role, name, value, backendNodeId per node). Pass `backendNodeId` to scope, `maxDepth` to limit depth. - -### Session.wait_for_script [#session-wait-for-script] - -```python -def wait_for_script(self, *, script: str, timeout: int | None = None) -> Any -``` - -Wait until a JS expression returns truthy. Re-evaluates on each tick of the event loop. Use for synchronization beyond what CSS selectors can express — e.g. `window.dataLoaded === true`, `document.readyState === 'complete'`, `document.querySelectorAll('.row').length >= 5`. - -### Session.wait_for_selector [#session-wait-for-selector] - -```python -def wait_for_selector(self, *, selector: str, timeout: int | None = None) -> Any -``` - -Wait for an element matching a CSS selector to appear in the page. Returns the backend node ID of the matched element. - -### Session.wait_for_state [#session-wait-for-state] - -```python -def wait_for_state(self, *, state: str, timeout: int | None = None) -> Any -``` - -Wait for the CURRENT page to reach a load state (no navigation). After a [`goto`](#session-goto), the page is returned at the fast `load` snapshot, so content rendered by post-load JS (XHR-loaded lists, feeds, search results) may still be missing. When a read looks incomplete — empty lists, spinners, skeletons — call this with 'networkidle' and re-read. Prefer 'networkidle'; 'done' can be slow on sites with constant background activity (ads, polling). - -## run_script [#run-script] - -```python -def run_script( - script: str | os.PathLike, - env: dict[str, str] | None = None, - binary: str | os.PathLike | None = None, - timeout: float | None = None -) -> str -``` - -Replay a saved lightpanda script (no LLM) and return its stdout. - -``env`` entries (e.g. ``LP_*`` placeholder values) are added to the -child's environment. Raises [`ScriptError`](#scripterror) on a non-zero exit. - -## AsyncBrowser [#asyncbrowser] - -A lightpanda browser process, driven from asyncio. - -The subprocess is spawned by [`start()`](#asyncbrowser-start) — called automatically on -``async with`` entry and by [`new_session()`](#browser-new-session). Not fork-inheritable, -same as [`Browser`](#browser). - -```python -class AsyncBrowser( - binary: str | os.PathLike | None = None, - env: dict[str, str] | None = None, - timeout: float = 300.0, - verbose: bool = False, - args: Sequence[str] = (), - max_concurrency: int = 32 -) -``` - -``binary``/``env``/``timeout``/``verbose``/``args`` are forwarded -to [`Browser`](#browser). ``max_concurrency`` caps concurrently executing -tool calls across this browser's sessions (worker threads are -created lazily). - -Usable as a context manager (`async with`). - -Every public method and property of [`Browser`](#browser) exists on `AsyncBrowser` with the same name and signature; methods are coroutines to `await`. Only the members `AsyncBrowser` adds are listed below. - -### AsyncBrowser.wrap [#asyncbrowser-wrap] - -```python -@classmethod -def wrap( - cls, - browser: Browser, - max_concurrency: int = 32 -) -> AsyncBrowser -``` - -Adopt an already-running [`Browser`](#browser) — the migration path for -driving existing sync setup from asyncio. [`close()`](#browser-close) shuts down -the facade but leaves the wrapped browser running. - -### AsyncBrowser.start [#asyncbrowser-start] - -```python -async def start(self) -> AsyncBrowser -``` - -Spawn the browser process and fetch its tool list. Idempotent. - -### AsyncBrowser.session [#asyncbrowser-session] - -```python -@contextlib.asynccontextmanager -async def session(self) -``` - -``async with browser.session() as page:`` — a session scoped to -the block and closed on exit. Use [`new_session()`](#browser-new-session) for the -unscoped form. - -## AsyncSession [#asyncsession] - -One isolated browsing context (own page, cookies, memory), async. - -Do not construct directly — use [`AsyncBrowser.new_session()`](#browser-new-session). - -```python -class AsyncSession( - session: Session, - executor: concurrent.futures.thread.ThreadPoolExecutor -) -``` - -Usable as a context manager (`async with`). - -Every public method and property of [`Session`](#session) exists on `AsyncSession` with the same name and signature; methods are coroutines to `await`. `AsyncSession` adds no members of its own. - -## run_script_async [#run-script-async] - -```python -async def run_script_async( - script: str | os.PathLike, - env: dict[str, str] | None = None, - binary: str | os.PathLike | None = None, - timeout: float | None = None -) -> str -``` - -Async variant of `lightpanda.run_script()` (runs in a worker thread). - -## CDPServer [#cdpserver] - -A lightpanda process serving the Chrome DevTools Protocol on 127.0.0.1. - -```python -from lightpanda import CDPServer -from playwright.sync_api import sync_playwright - -with CDPServer() as server, sync_playwright() as p: - browser = p.chromium.connect_over_cdp(server.ws_endpoint) - page = browser.new_context().new_page() - page.goto("https://example.com") -``` - -Every connected client gets its own browser; up to 16 connect at once -by default (``args=["--cdp-max-connections", "N"]`` to change). The -process is stopped by [`close()`](#cdpserver-close) / leaving the ``with`` block, and on -Linux also when the interpreter dies. - -```python -class CDPServer( - binary: str | os.PathLike | None = None, - env: dict[str, str] | None = None, - verbose: bool = False, - args: Sequence[str] = (), - port: int | None = None -) -``` - -[``port``](#cdpserver-port) pins the listening port (default: a free one). ``args`` -are extra ``lightpanda serve`` flags; pass ``port=`` rather than -``--port``. ``verbose`` lets the browser log through to stderr. - -Usable as a context manager (`with`). - -### CDPServer.ws_endpoint [#cdpserver-ws-endpoint] - -```python -ws_endpoint: str -``` - -*Property.* The CDP WebSocket URL, ``ws://127.0.0.1:/``. - -Keep it as is: the server only upgrades on path ``/`` and only -accepts an IP-literal or ``localhost`` host. - -### CDPServer.version [#cdpserver-version] - -```python -def version(self) -> dict -``` - -The ``/json/version`` document (browser, protocol version, -``webSocketDebuggerUrl``). - -### CDPServer.port [#cdpserver-port] - -```python -port: int -``` - -*Property.* - -### CDPServer.http_endpoint [#cdpserver-http-endpoint] - -```python -http_endpoint: str -``` - -*Property.* ``http://127.0.0.1:``, the server's HTTP root: what Puppeteer -(``browserURL``) and Playwright (``connect_over_cdp`` with an http URL) -discover the CDP WebSocket from, and Selenium's ``command_executor``. - -### CDPServer.close [#cdpserver-close] - -```python -def close(self) -> None -``` - -## AsyncCDPServer [#asynccdpserver] - -[`CDPServer`](#cdpserver) for asyncio: the process is spawned by -[`start()`](#asynccdpserver-start), called automatically on ``async with`` entry. - -```python -async with AsyncCDPServer() as server, async_playwright() as p: - browser = await p.chromium.connect_over_cdp(server.ws_endpoint) -``` - -```python -class AsyncCDPServer( - binary: str | os.PathLike | None = None, - env: dict[str, str] | None = None, - verbose: bool = False, - args: Sequence[str] = (), - port: int | None = None -) -``` - -Arguments are forwarded to the sync class. - -Usable as a context manager (`async with`). - -Every public method and property of [`CDPServer`](#cdpserver) exists on `AsyncCDPServer` with the same name and signature; methods are coroutines to `await`. Only the members `AsyncCDPServer` adds are listed below. - -### AsyncCDPServer.start [#asynccdpserver-start] - -```python -async def start(self) -``` - -Spawn the server process. Idempotent. - -## BiDiServer [#bidiserver] - -A lightpanda process serving WebDriver BiDi on 127.0.0.1. - -```python -from lightpanda import BiDiServer -from selenium import webdriver -from selenium.webdriver.common.options import ArgOptions - -options = ArgOptions() -options.web_socket_url = True # ask for a WebDriver BiDi session - -with BiDiServer() as server: - driver = webdriver.Remote(command_executor=server.http_endpoint, options=options) - context = driver.browsing_context.create(type="tab") - driver.browsing_context.navigate(context=context, url="https://example.com", wait="complete") - print(driver.script.execute("() => document.title", context_id=context)["value"]) - driver.quit() -``` - -[`http_endpoint`](#bidiserver-http-endpoint) is Selenium's ``command_executor``. The browser -serves the BiDi modules (``session``, ``browser``, ``browsingContext``, -``script``, ``input``) over the WebSocket plus the classic session -bootstrap (``GET /status``, ``POST /session`` with the ``webSocketUrl`` -capability, ``DELETE /session/``); other classic WebDriver commands -such as Selenium's ``driver.get`` or ``find_element`` are not served, so -drive the page through ``driver.browsing_context`` and ``driver.script`` -with an explicit context, created first as above. Pass -``args=["--protocol", "cdp"]`` to serve CDP on the same port as well -(``--protocol`` is additive). The process is stopped by [`close()`](#bidiserver-close) / -leaving the ``with`` block, and on Linux also when the interpreter dies. - -```python -class BiDiServer( - binary: str | os.PathLike | None = None, - env: dict[str, str] | None = None, - verbose: bool = False, - args: Sequence[str] = (), - port: int | None = None -) -``` - -[``port``](#bidiserver-port) pins the listening port (default: a free one). ``args`` -are extra ``lightpanda serve`` flags; pass ``port=`` rather than -``--port``. ``verbose`` lets the browser log through to stderr. - -Usable as a context manager (`with`). - -### BiDiServer.bidi_endpoint [#bidiserver-bidi-endpoint] - -```python -bidi_endpoint: str -``` - -*Property.* The session-less BiDi WebSocket URL, ``ws://127.0.0.1:/session``, -for clients that speak BiDi directly (``session.new`` over the socket). -A session bootstrapped through ``POST /session`` gets its own socket at -``/``, returned as the ``webSocketUrl`` -capability. - -Keep the IP literal: the WebSocket upgrade rejects any ``Origin`` -header and only accepts an IP-literal or ``localhost`` host. - -### BiDiServer.status [#bidiserver-status] - -```python -def status(self) -> dict -``` - -The ``GET /status`` value, ``{"ready": True, "message": ""}``. - -### BiDiServer.port [#bidiserver-port] - -```python -port: int -``` - -*Property.* - -### BiDiServer.http_endpoint [#bidiserver-http-endpoint] - -```python -http_endpoint: str -``` - -*Property.* ``http://127.0.0.1:``, the server's HTTP root: what Puppeteer -(``browserURL``) and Playwright (``connect_over_cdp`` with an http URL) -discover the CDP WebSocket from, and Selenium's ``command_executor``. - -### BiDiServer.close [#bidiserver-close] - -```python -def close(self) -> None -``` - -## AsyncBiDiServer [#asyncbidiserver] - -[`BiDiServer`](#bidiserver) for asyncio: the process is spawned by -[`start()`](#asyncbidiserver-start), called automatically on ``async with`` entry. - -```python -class AsyncBiDiServer( - binary: str | os.PathLike | None = None, - env: dict[str, str] | None = None, - verbose: bool = False, - args: Sequence[str] = (), - port: int | None = None -) -``` - -Arguments are forwarded to the sync class. - -Usable as a context manager (`async with`). - -Every public method and property of [`BiDiServer`](#bidiserver) exists on `AsyncBiDiServer` with the same name and signature; methods are coroutines to `await`. Only the members `AsyncBiDiServer` adds are listed below. - -### AsyncBiDiServer.start [#asyncbidiserver-start] - -```python -async def start(self) -``` - -Spawn the server process. Idempotent. - -## Exceptions [#exceptions] - -### LightpandaError [#lightpandaerror] - -```python -class LightpandaError(Exception) -``` - -Base error for the lightpanda package. - -### ProcessError [#processerror] - -```python -class ProcessError(LightpandaError) -``` - -The browser binary could not be found, started, or reached. - -### ProtocolError [#protocolerror] - -```python -class ProtocolError(LightpandaError) -ProtocolError(message: str, code: int | None = None) -``` - -JSON-RPC level failure (invalid request, timeout, internal error). - -- `code` - -### ScriptError [#scripterror] - -```python -class ScriptError(LightpandaError) -ScriptError(message: str, returncode: int, stdout: str = '', stderr: str = '') -``` - -A script replay ([`run_script`](#run-script)) exited with a failure. - -- `returncode` -- `stdout` -- `stderr` - -### ToolError [#toolerror] - -```python -class ToolError(LightpandaError) -``` - -A browser tool reported failure (bad selector, JS exception, ...). diff --git a/src/content/reference/python.mdx b/src/content/reference/python.mdx index c60d4b3..2620b2e 100644 --- a/src/content/reference/python.mdx +++ b/src/content/reference/python.mdx @@ -1,176 +1,827 @@ --- title: Python SDK -description: Reference of the classes and methods in the Lightpanda Python package, covering every browser action and script replay. +description: Reference of every public class, method, property and exception in the lightpanda Python package, generated from the signatures and docstrings shipped in the code. --- +{/* Generated by scripts/generate-python-reference.py from lightpanda-python main. Do not edit; rerun the script or wait for the python-reference workflow. */} # Python SDK -The [`lightpanda` package](https://pypi.org/project/lightpanda/) exposes `Browser`/`AsyncBrowser`, which spawn and manage the bundled binary, and `Session`/`AsyncSession`, with one method per browser action. See [Use the Python SDK](/guides/use-python) for practical documentation, and the [Python API](/reference/python-api) page for every signature and docstring as shipped in the package. +Every public class, method, property and exception of the [`lightpanda` package](https://pypi.org/project/lightpanda/), with the signatures and docstrings shipped in the code. See [Use the Python SDK](/guides/use-python) for a practical walkthrough. Every sync class has an asyncio twin with the same methods, awaitable; the async sections below list only what the twin adds. -## Browser +Browser actions are keyword-only methods on [`Session`](#session) and [`AsyncSession`](#asyncsession), named in snake_case after the browser's own action names: the `waitForSelector` action is `wait_for_selector`, and its `backendNodeId` argument is `backend_node_id`. -[`Browser()`](/reference/python-api#browser) spawns the bundled binary when constructed. It is not fork-inheritable: create a fresh instance in a forked child. +Where a method accepts both `selector` and `backend_node_id`, pass one of the two. `selector` is preferred for reproducibility and wins when both are given; `backend_node_id` takes the values returned by [`tree`](#session-tree), [`links`](#session-links) or [`find_element`](#session-find-element). -| Argument | Default | Description | -|---|---|---| -| `binary` | `None` | Path to a specific lightpanda binary. When omitted, resolved from the `LIGHTPANDA_BIN` environment variable, then the binary bundled in the package, then `PATH`. | -| `env` | `None` | Extra environment variables for the spawned process. | -| `timeout` | `300.0` | Seconds to wait for a response before raising `ProtocolError`. | -| `verbose` | `False` | Print the spawned process's own logging. | -| `args` | `()` | Extra CLI flags for the spawned process, for example `["--http-cache-dir", path]`. | +## Browser [#browser] -| Method | Returns | Description | -|---|---|---| -| `new_session()` | `Session` | Open a new isolated browsing context: its own page, cookies, and memory. | -| `tools` (property) | `dict[str, dict]` | Every available action, as `name → {description, schema}`, reported live by the running browser. | -| `close()` | `None` | Stop the browser process. | +A lightpanda browser process. Spawns the bundled binary on first use. -`with Browser() as b:` calls `close()` on exit. +Not fork-inheritable: after ``os.fork()``/``multiprocessing``, create a +fresh Browser in the child. -## AsyncBrowser +```python +class Browser( + binary: str | os.PathLike | None = None, + env: dict[str, str] | None = None, + timeout: float = 300.0, + verbose: bool = False, + args: Sequence[str] = () +) +``` + +``args`` are extra CLI flags for the spawned browser process +(e.g. ``["--http-cache-dir", path]`` or cookie flags). + +Usable as a context manager (`with`). + +### Browser.tools [#browser-tools] + +```python +tools: dict[str, dict] +``` + +*Property.* Tool name → \{description, schema\}, as reported by the browser. + +### Browser.new_session [#browser-new-session] + +```python +def new_session(self) -> Session +``` + +### Browser.close [#browser-close] + +```python +def close(self) -> None +``` + +## Session [#session] + +One isolated browsing context (own page, cookies, memory). + +Do not construct directly — use [`Browser.new_session()`](#browser-new-session). + +[`call`](#session-call) is the escape hatch that takes the action and argument names exactly as the browser declares them. A failed action raises [`ToolError`](#toolerror). + +```python +class Session(browser: Browser, session_id: str) +``` + +Usable as a context manager (`with`). + +### Session.id [#session-id] + +```python +id: str +``` + +*Property.* + +### Session.call [#session-call] + +```python +def call(self, tool: str, **kwargs) +``` + +Invoke a browser tool by name. The generated methods route here. + +Returns parsed JSON for JSON-carrying tools, ``bytes`` for image +results ([``screenshot``](#session-screenshot) without ``path``), otherwise the result text. + +### Session.close [#session-close] + +```python +def close(self) -> None +``` + +### Session.click [#session-click] + +```python +def click( + self, + *, + selector: str | None = None, + backend_node_id: int | None = None +) -> Any +``` + +Click on an interactive element. Provide either a CSS selector (preferred for reproducibility) or a backendNodeId. Returns the current page URL and title after the click. + +### Session.console_logs [#session-console-logs] + +```python +def console_logs(self) -> Any +``` + +Get buffered console.log/warn/error messages from the current page. Returns all messages since last call and clears the buffer. + +### Session.detect_forms [#session-detect-forms] + +```python +def detect_forms(self, *, url: str | None = None, timeout: int | None = None) -> Any +``` + +Detect all forms on the page and return their structure including fields, types, and required status. If a url is provided, it navigates to that url first. + +### Session.evaluate [#session-evaluate] + +```python +def evaluate( + self, + *, + script: str, + url: str | None = None, + timeout: int | None = None, + save: str | None = None +) -> Any +``` + +Evaluate JavaScript in the current page context — an escape hatch for page-side logic the dedicated tools can't express; prefer [`extract`](#session-extract) for data and click/fill/etc. for actions. It runs in the page, so it cannot see the agent script's variables or builtins — interpolate any value into the `script` string. A bare trailing expression yields its value; top-level `await` and `return` are supported (the body then runs as an async function, so use `return` to produce a value). Objects and arrays return as JSON, so no `JSON.stringify` is needed. If a url is provided, it navigates there first. The `globalThis.lp` object exposes a Session-scoped bridge store: values written via `lp.foo = ...` auto-sync at end of evaluate, surviving navigation; values previously set via `/extract save=` or `/evaluate save=` appear as `lp.`. + +### Session.extract [#session-extract] + +```python +def extract(self, *, schema: str | dict | list, save: str | None = None) -> Any +``` + +Extract structured data from the current page (navigate first). `schema` is a JSON object (passed as a string) mapping output field names to CSS-selector specs. It is NOT a JSON Schema — no "type"/"properties" wrappers; the keys ARE your output fields. Value shapes: + +```text +"" → first match's text (trimmed; null if no match) +[""] → every match's text (string[]) +{"selector":"","attr":""} → first match's attribute value (href/src resolved to absolute URLs) +[{"selector":"","attr":""}] → every match's attribute (string[]) +[{"selector":"","fields":{…}}] → one object per match; field selectors resolve relative to that match and accept any shape above ("" = the match's own text; nest arrays for per-item sub-lists) +``` + +Add "limit": N inside any array's object spec to cap matches. +Every extracted value is a string or null — parse numbers downstream. An empty array is a valid result, but if ALL top-level keys miss, the call errors: inspect the page (tree/markdown) and retry with corrected selectors. +Finish data tasks with extract — it is the only read recorded as a replayable `extract(...)` script call; answers lifted from [`markdown`](#session-markdown) text in chat are not. + +Examples (schema → result): + +```text +{"karma": "#karma"} → {"karma":"42"} +{"items": [".story .title"]} → {"items":["Title 1","Title 2"]} +{"top3": [{"selector":".story .title","limit":3}]} → {"top3":["A","B","C"]} +{"links": [{"selector":"a.title","attr":"href"}]} → {"links":["https://site/a","https://site/b"]} +{"stories": [{"selector":".athing","fields":{"title":".titleline","rank":".rank"}}]} → {"stories":[{"title":"Foo","rank":"1"}]} +``` + +### Session.fill [#session-fill] + +```python +def fill( + self, + *, + value: str, + selector: str | None = None, + backend_node_id: int | None = None +) -> Any +``` + +Fill text into an input element. Provide either a CSS selector (preferred for reproducibility) or a backendNodeId. -[`AsyncBrowser`](/reference/python-api#asyncbrowser) mirrors `Browser` for asyncio: every call runs on a browser-owned thread pool, so the event loop is never blocked. +### Session.find_element [#session-find-element] -| Argument | Default | Description | -|---|---|---| -| `binary`, `env`, `timeout`, `verbose`, `args` | same as `Browser` | Forwarded to the underlying `Browser`. | -| `max_concurrency` | `32` | Caps method calls executing concurrently across this browser's sessions. Worker threads are created lazily. | +```python +def find_element(self, *, role: str | None = None, name: str | None = None) -> Any +``` + +Find interactive elements by role and/or accessible name. Returns matching elements with their backend node IDs. Useful for locating specific elements without parsing the full semantic tree. + +### Session.get_cookies [#session-get-cookies] + +```python +def get_cookies(self, *, url: str | None = None, all: bool | None = None) -> Any +``` + +Cookies stored in the browser. Defaults to cookies whose domain matches the current page's host. Pass `url=` to filter for another host, or `all=true` to dump every cookie regardless of host. Useful for debugging authentication and session state. + +### Session.get_env [#session-get-env] + +```python +def get_env(self, *, name: str | None = None) -> Any +``` + +With `name`: read an LP_* env var (other namespaces report as not set) — for non-secret config only (base URLs, flags). Without `name`: list LP_* names that are set (no values) — safe credential discovery. For secrets, pass `$LP_*` placeholders in tool args; never request a credential by name (the value would land in your context). + +### Session.get_url [#session-get-url] + +```python +def get_url(self) -> Any +``` + +Current page URL. The browser may already have a page loaded (command, replayed script) not visible in this conversation — call this before assuming nothing is loaded when the user references the current page/site. Also useful to verify a navigation or detect a redirect. + +### Session.goto [#session-goto] + +```python +def goto( + self, + *, + url: str, + timeout: int | None = None, + wait_until: str | None = None +) -> Any +``` + +Navigate to a specified URL and load the page in memory so it can be reused later for info extraction. + +### Session.hover [#session-hover] + +```python +def hover( + self, + *, + selector: str | None = None, + backend_node_id: int | None = None +) -> Any +``` + +Hover over an element, triggering mouseover and mouseenter events. Provide either a CSS selector (preferred for reproducibility) or a backendNodeId. Useful for menus, tooltips, and hover states. + +### Session.html [#session-html] + +```python +def html( + self, + *, + selector: str | None = None, + backend_node_id: int | None = None, + max_bytes: int | None = None, + strip: dict | None = None, + url: str | None = None, + timeout: int | None = None +) -> Any +``` + +Raw HTML for the document or, with `selector`/`backendNodeId`, a single node's outerHTML. Verbose; use only when you need attributes that markdown discards. + +### Session.interactive_elements [#session-interactive-elements] + +```python +def interactive_elements(self, *, url: str | None = None, timeout: int | None = None) -> Any +``` + +Extract interactive elements from the opened page. If a url is provided, it navigates to that url first. + +### Session.links [#session-links] + +```python +def links( + self, + *, + limit: int | None = None, + url: str | None = None, + timeout: int | None = None +) -> Any +``` -| Method | Returns | Description | -|---|---|---| -| `start()` | `AsyncBrowser` | Spawn the process and fetch its action list. Idempotent; called automatically on `async with` entry and by `new_session()`. | -| `new_session()` | `AsyncSession` | Start the browser if needed, then open a new session. | -| `session()` | async context manager | `async with browser.session() as page:` opens a session scoped to the block and closes it on exit. | -| `tools` (property) | `dict[str, dict]` | Same as `Browser.tools`. | -| `close()` | `None` | Stop the browser process, unless it was adopted with `wrap` (see below). | -| `AsyncBrowser.wrap(browser, max_concurrency=32)` (classmethod) | `AsyncBrowser` | Adopt an already-running `Browser` for use from asyncio. `close()` then shuts down only the async facade, leaving the wrapped browser running. | +Extract the visible links in the opened page as JSON objects with `text` (anchor text, falling back to aria-label/title/image alt), `href` (resolved URL), and `backendNodeId` (pass to click/nodeDetails). One entry per href; hidden links are omitted. If a url is provided, it navigates to that url first. -`async with AsyncBrowser() as b:` calls `start()` on entry and `close()` on exit. +### Session.markdown [#session-markdown] -## Session and AsyncSession +```python +def markdown( + self, + *, + selector: str | None = None, + backend_node_id: int | None = None, + max_bytes: int | None = None, + url: str | None = None, + timeout: int | None = None +) -> Any +``` + +Render the page (or a subtree) as markdown. Scope with `selector` or `backendNodeId` to read just the relevant region — full-page markdown is the last resort. Use `maxBytes` to cap long pages. + +### Session.node_details [#session-node-details] + +```python +def node_details(self, *, backend_node_id: int) -> Any +``` -`Browser.new_session()` and `AsyncBrowser.new_session()` are the only way to obtain a [`Session`](/reference/python-api#session) or [`AsyncSession`](/reference/python-api#asyncsession); do not construct one directly. +Details for a node by backendNodeId: a ready-to-use CSS `selector` that resolves to the node (the first match, as click/fill resolve it), plus tag, role, name, interactivity, disabled, value, input type, placeholder, href, id, class, checked, select options. The canonical way to turn a tree backendNodeId into a CSS selector. -| Member | Description | -|---|---| -| `id` (property) | The session's id. | -| `close()` | Close the session. Calls made after `close()` raise `ToolError`. | -| `call(action, **kwargs)` | Invoke any action by name. The methods documented below route through this; it also accepts the action and argument names exactly as the browser declares them, for example `page.call("tree", maxDepth=1)`. | +### Session.press [#session-press] -Sessions are context managers too: `with browser.new_session() as page:` closes the session on exit. Closing the browser ends every session anyway. +```python +def press( + self, + *, + key: str, + selector: str | None = None, + backend_node_id: int | None = None +) -> Any +``` -## Calling an action +Press a keyboard key, dispatching keydown and keyup events. Use key names like 'Enter', 'Tab', 'Escape', 'ArrowDown', 'Backspace', or single characters like 'a', '1'. Common shorthand is normalized: 'enter'/'return' → 'Enter', 'esc' → 'Escape', 'up'/'down'/'left'/'right' → 'Arrow*', 'space' → ' '. Pressing 'Enter' on a form input or submit button triggers implicit form submission. + +### Session.screenshot [#session-screenshot] + +```python +def screenshot( + self, + *, + path: str | None = None, + selector: str | None = None, + backend_node_id: int | None = None, + full_page: bool | None = None, + url: str | None = None, + timeout: int | None = None +) -> Any +``` -Every browser action is a method on `Session`/`AsyncSession`, keyword-only, with the action and its arguments in snake_case: the `waitForSelector` action is `wait_for_selector`, and its `backendNodeId` argument is `backend_node_id`. The methods are generated from the bundled browser's action schemas, so the signatures and docstrings your IDE shows come straight from the binary. The [Python API](/reference/python-api#session) page lists every method with its exact signature and docstring. +Render the page, or one node, as a PNG: the text layout Lightpanda computes, not a pixel-accurate browser rendering (no images, fonts or CSS colours). With `path`, writes the file at full size and returns its location; without it, returns the image inline where the client can display one, at most 1280px wide and 4096px tall. Use it to see spatial layout; read content with [`markdown`](#session-markdown)/[`tree`](#session-tree). -A failed action raises `ToolError`. +### Session.scroll [#session-scroll] -In the Arguments column below, `?` marks an optional keyword argument, and `selector` / `backend_node_id` marks a pair where one of the two is required. Prefer `selector` for reproducibility; it also wins when you pass both. `backend_node_id` takes the `backendNodeId` values returned by a prior `tree`, `links`, or `find_element` call. In the Returns column, `JSON` is a parsed Python `dict` or `list`, `text` is a plain string. +```python +def scroll( + self, + *, + backend_node_id: int | None = None, + x: int | None = None, + y: int | None = None +) -> Any +``` -### Navigation and search +Scroll the page or a specific element. Returns the scroll position and current page URL and title. -These methods bring a page into the browser: +### Session.search [#session-search] -| Method | Arguments | Returns | Description | -|---|---|---|---| -| `goto` | `url`, `timeout?`, `wait_until?` | text | Navigate to a URL and load the page in memory so it can be reused later for info extraction. `wait_until` accepts the same states as `wait_for_state` and defaults to `load`. | -| `search` | `query`, `timeout?` | text | Run a web search and return results as markdown: a numbered list of `{title, url, snippet}`. The browser does not navigate; to open a result, call `goto` with its URL. | +```python +def search(self, *, query: str, timeout: int | None = None) -> Any +``` -### Reading the page +Run a web search and return results as markdown: a numbered list of \{title, url, snippet\}. Search tries brave, tavily, exa, then keenable in order, each when its API key (BRAVE_API_KEY, TAVILY_API_KEY, EXA_API_KEY or KEENABLE_API_KEY) is set; keenable also works without a key through its public endpoint (rate-limited per client IP). Prefer this over goto-ing google.com/search directly (Google blocks the browser on User-Agent/TLS). The browser does not navigate — to open a result, use [`goto`](#session-goto) with its URL. -These methods read the loaded page without modifying it: +### Session.select_option [#session-select-option] -| Method | Arguments | Returns | Description | -|---|---|---|---| -| `markdown` | `selector?`, `backend_node_id?`, `max_bytes?`, `url?`, `timeout?` | text | Render the page, or a subtree, as markdown. Scope with `selector` or `backend_node_id` to read just the relevant region; use `max_bytes` to cap long pages. | -| `html` | `selector?`, `backend_node_id?`, `max_bytes?`, `strip?`, `url?`, `timeout?` | text | Raw HTML for the document, or a single node's outerHTML when scoped. Verbose; use only when you need attributes that markdown discards. Use `max_bytes` to cap long pages. `strip` is an object of element groups to omit: `js` (script, noscript, script preloads), `css` (style, stylesheet links), `ui` (css plus img, picture, video, audio, svg, canvas, iframe) and `invisible` (elements set to display:none). `{"js": True, "css": True}` keeps a page dump small. | -| `screenshot` | `path?`, `selector?`, `backend_node_id?`, `full_page?`, `url?`, `timeout?` | text or bytes | Render the page, or one node, as a PNG: the text layout Lightpanda computes, not a pixel-accurate rendering (no images, fonts, or CSS colors). `path` must be a relative path. Without `path`, the PNG is returned as `bytes`. | -| `tree` | `url?`, `timeout?`, `backend_node_id?`, `max_depth?` | text | Simplified semantic DOM tree: role, name, value, and `backendNodeId` per node. | -| `links` | `limit?`, `url?`, `timeout?` | JSON | Extract all links as `text` (visible anchor text, falling back to aria-label/title/image alt), `href` (resolved URL), and `backendNodeId` (pass to `node_details`). One entry per href; hidden links are omitted. `limit` returns at most that many links, in document order. | -| `node_details` | `backend_node_id` | JSON | Tag, role, name, value, and other state for a node, plus a ready-to-use CSS `selector` that resolves to it. The way to turn a `backendNodeId` into a selector. | -| `find_element` | `role?`, `name?` | JSON | Find interactive elements by role and/or accessible name, with their `backendNodeId`. | -| `interactive_elements` | `url?`, `timeout?` | JSON | Every interactive element on the page. | -| `structured_data` | `url?`, `timeout?` | JSON | Structured data on the page, such as JSON-LD or OpenGraph tags. | -| `detect_forms` | `url?`, `timeout?` | JSON | Forms on the page: fields, types, and required status. | +```python +def select_option( + self, + *, + value: str, + selector: str | None = None, + backend_node_id: int | None = None +) -> Any +``` -### Data extraction and scripting +Select an option in a <select> dropdown element by its value. Provide either a CSS selector (preferred for reproducibility) or a backendNodeId. Dispatches input and change events. -| Method | Arguments | Returns | Description | -|---|---|---|---| -| `extract` | `schema`, `save?` | JSON | Extract structured data from the current page using a schema mapping output field names to CSS-selector specs. | -| `evaluate` | `script`, `url?`, `timeout?`, `save?` | typed | Evaluate a JavaScript string in the page context and return its value. Runs in the page, so it cannot see your Python variables. | +### Session.set_checked [#session-set-checked] -**`evaluate`'s return** is typed like the JavaScript result: for example `1+1` comes back as the `int` `2`, and `({a:1})` comes back as the `dict` `{"a": 1}`. +```python +def set_checked( + self, + *, + checked: bool, + selector: str | None = None, + backend_node_id: int | None = None +) -> Any +``` -**`extract`'s `schema`** maps output field names to CSS-selector specs. Pass it as a Python `dict` or `list`, it's encoded for you; a JSON string also works: +Check or uncheck a checkbox or radio button. Provide either a CSS selector (preferred for reproducibility) or a backendNodeId. Dispatches input, change, and click events. -| Schema value | Result | -|---|---| -| `""` | First match's text, or `None` | -| `[""]` | Every match's text | -| `{"selector": "", "attr": ""}` | First match's attribute (`href`/`src` resolve to absolute URLs) | -| `[{"selector": "", "attr": ""}]` | Every match's attribute | -| `[{"selector": "", "fields": {...}}]` | One dict per match, with fields resolved relative to each match | +### Session.structured_data [#session-structured-data] -Add `"limit": N` inside any array spec to cap matches. Every extracted value is a string or `None`; parse numbers yourself. +```python +def structured_data(self, *, url: str | None = None, timeout: int | None = None) -> Any +``` -### Interacting with the page +Extract structured data (like JSON-LD, OpenGraph, etc) from the opened page. If a url is provided, it navigates to that url first. -These methods dispatch real DOM events on the page: +### Session.tree [#session-tree] -| Method | Arguments | Returns | Description | -|---|---|---|---| -| `click` | `selector` / `backend_node_id` | text | Click an interactive element. | -| `fill` | `selector` / `backend_node_id`, `value` | text | Fill text into an input element. | -| `scroll` | `backend_node_id?`, `x?`, `y?` | text | Scroll the page, or a specific element if `backend_node_id` is given. | -| `hover` | `selector` / `backend_node_id` | text | Hover over an element, triggering `mouseover` and `mouseenter`. | -| `press` | `key`, `selector?`, `backend_node_id?` | text | Press a keyboard key, dispatching `keydown` and `keyup`. Targets the document if no element is given. | -| `select_option` | `selector` / `backend_node_id`, `value` | text | Select an option in a `