Browser automation that doesn’t guess.

Real Chromium in the cloud, controlled from inside the engine. From Go, TypeScript, MCP or BrowserVM.

session 7f3a · eu-central
Cart › Shipping › Payment€ 129.00Emailiframe · js.pay.example · own processPlace orderTotal € 129.00 · pay nowmatched · visible · steadyana@example.de4242 4242 4242 4242Get 10% off your first orderSubscribediv[role=dialog] .newsletterYour privacyAcceptdiv#cookie-wall · fixedcovered 50%shop.example/checkout
  1. 00
    b.AddReactionWith(ctx, bs.CSS(".newsletter"), bs.ReactionOpts{On: bs.CSS(".close")})
    registeringarmed · watched in the browserfired once · closed .newsletter
  2. 01
    b.Wait(ctx, bs.CSS("#email"), bs.Timeout(15000))
    waiting for the formready in 1.6 s
  3. 02
    b.Fill(ctx, bs.CSS("#email"), "ana@example.de")
    typing14 keys · de-DE layout
  4. 03
    b.Fill(ctx, bs.CSS("[name=card]").InAllFrames(), card.Number)
    typing19 keys · frame 2
    paused for reaction 00 · resumed at key 15
  5. 04
    b.Click(ctx, bs.CSS("#place-order"))
    workinglanded on button#place-order
    re-aimed past div#cookie-wall

Five calls in your code. The loading, the pointer paths, the typing rhythm, the popup, the pixel check and the re-aim all happen inside them.

Try it now

One package and an API key. Or let the CLI scaffold a module your coding agent builds in.

TypeScript SDKnpm ↗

Node only. Drive a session from any script or service.

The same browser API, idiomatic Go.

CLI · for your agentGitHub ↗

init scaffolds a module that runs; dev rebuilds and streams logs.

02Waiting

One wait. Every way the page can go.

One call names every way a step can end and how long it may take. Until one holds, visible, reachable and still, nothing polls and nothing crosses the wire. The page reports its own changes, the check runs inside the browser, and only the answer comes back.

checkout.goone deadline, four outcomes
1_, _ = b.Navigate(ctx, "https://shop.example/checkout", 0)2 3// A step can take a while. Say how long, and every way it can go.4res, err := b.Wait(ctx,5    bs.CSS("#order-confirmation"),6    bs.CSS("#address-form"),7    bs.CSS("iframe[title*=verification]").InAllFrames(),8    bs.JS("!!document.querySelector('[data-error]')"),9    bs.Timeout(25000),10)11 12switch res.Index {13case 0: // done; res.FrameId and res.Bounds say where14case 2: // the verification frame, and nothing here named it15}

And while it waits, it does nothing

A condition that does not hold yet is not hammered on a timer. The documents it concerns report when they change, and the check runs then, inside the browser, so breadth costs a registration, not a loop.

ten seconds of one wait4 conditions × 12 frames
The page
holds
Polling
browserscale
answer
0s2s4s6s8s10s
2,160 checks

every 200 ms, whether the page moved or not, and one tick late

5 wake-ups · 1 message

the page reports each change, a slow sweep covers the rest, all inside the browser; one message back

And almost none of it crosses the wire

Checked from outside, every question about the layout is a message to the browser and one back: where the element is, what sits on top of it, whether it still moves. Here those checks run where the page is, and only the answer travels.

one check, from outside× 2,160
your codebrowser
find #pay-button
node 412
box of node 412
342, 618 · 96×40
what is at 390, 638?
div#promo-modal
box again, next frame
342, 618 · 96×40
17,280 messages

8 per check, for one condition in one frame, × 2,160 checks. Visible, covered, still moving: each is a question over the protocol and an answer back.

the whole wait, here× 1
your codebrowser
wait · 4 conditions · 25 s
nothing on the wirere-checked in the browser
when the page reports a change
holds · #order-confirmation
1 message back

The answer, with the element, its frame and its bounds. The ten seconds before it cost the wire nothing.

03Clicking

A click that lands — or says what stopped it.

What usually becomes a few hundred lines of retry logic and a dozen round trips is one call, done where the layout lives.

b.Click(ctx, bs.CSS("#place-order"))one command · one answer
  1. 01locate

    any frame, for up to 5 s

  2. 02scroll

    through nested scrollers and frames

  3. 03settle

    until the bounds hold still

  4. 04move

    a hand’s path, not a jump

  5. 05verify

    which element owns that pixel

  6. 06re-aim

    exposed patch, or step out of the overlay

Pressed

frame · node · bounds · where it landed

Refused, with the blocker

never presses what lies on top

*bs.ClickErrorrefused
code
occluded_after_evade
occluder
div#cookie-wall.banner
text
“We value your privacy”
bounds
x 0 · y 612 · 1280×188
pointerEvents
auto
position
fixed
zIndex
2147483647
hittableWhileInvisible
false

And typing that stays in the field

Keys on the layout of the session’s region, timed the way a hand types. If something steals focus half way, it stops and names the element instead of typing the rest of a password into it.

fill · input[name=email]de-DE layout
a92msn118a84.141k76@AltGr+Qs97h88o112p79.134d101e93
13 keys · all on targetfocus lost → stop, name the element
04Frames

The iframe that is not there yet.

The flake is not naming an iframe. It is that the frame does not exist yet. Here the frame set is whatever the page has right now, including frames written in while you wait.

bs.CSS("[name=otp]").InAllFrames()frame set, re-derived live
  • ▪shop.example/checkout
  • └js.pay.example
  • └3ds.bank.example
  • └ads.example
  • └ads.example/slot/4
  • └consent.example
matched in frame 3 · depth 2node handle → Fill()
01

Frames that appear mid-wait

Looked at the moment they exist. No wait-for-the-iframe before the wait for what is inside it.

02

Cross-process is not a special case

No session per frame, no injected world, no depth limit. A grandchild costs what a child costs.

03

The answer names the frame

A match carries its frame and a node handle. Pass the handle on and the next action goes to the right document.

05Reactions

A cookie bar is not a step in your flow.

Consent bars, newsletter modals, a survey two minutes in. No linear script can say when they appear, so you register a standing instruction and the browser carries it out.

one run · 14 s
Flow
gotoemailcard↻clickwait
Page
consentnewsletter
Fired
acceptclose
0s2s4s6s8s10s12s14s
click retrying under the modalfired in the browser, no round trip
reactions.goregistered once
1// The consent bar, wherever in the page it turns up.2_, _ = b.AddReaction(ctx,3    bs.CSS("#accept-all").InAllFrames())4 5// A modal we do not want to accept: click its close button instead.6_, _ = b.AddReactionWith(ctx,7    bs.CSS("div[role=dialog] .newsletter"),8    bs.ReactionOpts{On: bs.CSS("button[aria-label=Close]")})9 10// The rest of the run never mentions either of them again.
01

Fires in the gaps

Armed and carried out in the browser, only while the pointer is idle. A click already retrying under a modal lands.

02

Once, then it retires

A banner that comes back is not clicked twice by accident. List what is still armed, or remove it by id.

03

Every frame, every navigation

Including frames created later. Click the match itself, or something beside it such as its close button.

06Bookkeeping

Same job. One side is mostly bookkeeping.

A CDP client can only reach a frame it has attached to and run code in a context it has been told about. So before the checkout even starts, it is keeping a copy of the browser current, fed by events it has to subscribe to. browserscale leaves that state in the browser, where it already is.

what a CDP client keeps in memoryrebuilt on every navigation
  • ▪targetpage · session 1Target.attachedToTarget
  • └frameshop.example/checkoutPage.frameNavigated
  • └contextmain world · id 7Runtime.executionContextCreated
  • └contextutility world · id 8Runtime.executionContextCreated
  • └frameconsent.examplePage.frameAttached
  • └contextmain world · id 9Runtime.executionContextCreated
  • └targetjs.pay.example · session 2Target.attachedToTarget
  • └framejs.pay.example/v3Page.frameNavigated
  • └contextmain world · id 3Runtime.executionContextCreated
  • └target3ds.bank.example · session 3attached mid-checkout
  • ▪handles18 remote objectsRuntime.releaseObject, by you
Error: Execution context was destroyed, most likely because of a navigation.
what a browserscale client keepsone id
page
7f3a
  • Frames, sessions and contexts stay in the browser, where they already are.
  • Every call names the page and a selector; the browser finds the frame.
  • An answer carries the frame, node and bounds it is about.
Nothing to rebuild when the page navigates.

And the copy is kept current over the wire

Every request arrives as five or six network events, every frame reports its lifecycle, every context announces its creation and its end, whether your code cares or not. Here the connection carries what you asked and what came back.

the client’s connection during one checkout12 frames · ~150 requests · typical counts
CDP client · events it has to readsubscribed to keep the copy
  • Network.*≈ 900
    5–6 per request
  • Page.lifecycleEvent≈ 80
    every frame, every load
  • Page.frame*≈ 45
    attached, navigated, loading
  • Runtime.executionContext*≈ 30
    created, destroyed, per world
  • Target.*≈ 8
    a session per cross-origin frame
request and reply0 events
  1. navigate→ ←
  2. addReaction→ ←
  3. wait→ ←
  4. fill→ ←
  5. click→ ←
  6. wait→ ←
≈ 1,060 events

plus every command and its reply, several per action. The client reads all of it, including what it never asked about.

6 requests · 6 replies

Events arrive only when you ask for them: a network capture, a DOM mirror.

07Under the hood

Undetectable is a consequence of where it runs.

Detection reads the page, the input and the connection, and compares what it finds. When control lives in the engine instead of on a debugging protocol, every answer comes from the browser itself and agrees with the others, with nothing patched over.

what a page can ask about its browserhandled in the engine
01

Nothing to find

Control never touches the page.

  • no debugging protocolthe engine is driven directly, so there is no session for a probe to trip
  • nothing in the pageno injected script, no patched prototype, toString stays native
  • branded ChromeGoogle Chrome in every brand list, no headless marker anywhere
  • a real GPUWebGL and canvas come from the hardware that draws them
02

Input like a hand

Delivered the way hardware delivers it.

  • pointer pathscurved and uneven, never a jump to the centre
  • frame cadencemoves arrive on the display’s rhythm, jitter included
  • coalesced samplesseveral samples per frame, like a real mouse
  • native keystrokesthe region’s layout, dead keys and AltGr included
  • IME typingJapanese goes through the IME as romaji, key by key
03

One profile, no contradictions

One region, every frame and worker.

  • languagenavigator.languages, Accept-Language and Intl agree
  • time and placetimezone and locale match the exit IP
  • keyboard mapthe layout map matches the keys actually typed
  • speech voicesthe voice list a machine in that region has
  • WebRTCICE candidates show the exit IP, not the host
  • screenone monitor geometry and scrollbar width throughout
04

Chrome on the wire

Exactly what a stock install sends.

  • network stackChrome’s own TLS and HTTP/2, because it is Chrome’s
  • Client Hintsbrands and full version agree with the User-Agent
  • request headersmatching a regular Chrome installation
  • background requestscarry the same values as the page’s own

No stealth plugin and no patch list to keep current. A detector compares what it can read, and none of it was ever made to disagree.

08BrowserVM

Or put your code inside the browser.

Automation has lived in two places: in the page, fast but visible, or outside the browser, invisible but a round trip per question. BrowserVM is a third, an isolate beside the page.

In the page

Fast — and the site can see it.

Outside the browser

Invisible — and a round trip per question.

BrowserVM

Beside the page. Neither.

checkout.js — runs in the browserno frame ids
1// The processor's own cross-origin iframe, in its own process.2const pay = window.document3  .querySelector('iframe[name=pay]').contentDocument;4 5// A live element from that other document, handed to a real click.6await browser.fill(pay.querySelector('[name=number]'), card.number);7 8// Armed before the click that causes it, awaited after. No race:9// the script never left the browser in between.10const token = browser.waitForResponse('*/api/challenge*');11await browser.click(css('#verify'));12const body = (await token).response.body;
01

Cross-origin is a dot

Read contentDocument of the processor’s frame and hand its elements straight to a click. No frame ids.

02

Microseconds per step

Loops over tables and pagination are affordable, and arm-then-trigger has no race in it.

03

Nothing in the page

No global, no patched prototype, no evaluated source. Code in the page stays a separate, explicit call.

09Network & identity

Every request through you. Every login portable.

Watching is attached to the session, not a tab, so every iframe is in the log before its first request. And the browser’s identity, login included, is one object you take to the next run.

networkCapture · session-widecapturing
  • GET/checkout
    main200network
  • GET/static/app.9f2c.js
    main200static cache
  • POST/api/cart
    main200network
  • GETjs.pay.example/v3
    frame 2200http cache
  • GETjs.pay.example/inner
    frame 2200network
  • GET/sw/precache.json
    worker200service worker
  • POST3ds.bank.example/auth
    frame 3302network
  • GET3ds.bank.example/challenge
    frame 3200network
01

On before the first request

Armed on the session, not a tab. A new iframe is in the log before it loads anything.

02

A full log that never pauses the page

Every frame, worker and redirect hop, with the headers that went on the wire. Bodies are copied off to the side.

03

Or catch one call

Wait for a request, then read, rewrite or swallow it, or answer a whole navigation yourself.

04

Static paths, paid for once

Repeat JS, CSS and images come from a server-side cache that skips the proxy and its bill.

Who the browser is, kept between runs

01

A login you can carry

Import it into a fresh context and it comes up signed in, without a profile directory to copy around.

02

The same machine, returning

Pin it once and the next run is the same computer coming back, not a new install holding the same cookies.

03

A shipped Chrome, not a build of one

The state a consumer browser carries is filled in, consistent with the region the session exits from.

persona · de-shopper-014one object
region
DE · language de-DE · Europe/Berlin
keyboard
German, QWERTZ
machine
8 cores · 16 GB · same in every frame and worker
renderer
real GPU, pinned to this persona
webrtc
public IP matches the exit
state
214 cookies · storage for 9 origins
account
signed in · refresh token · device-bound session
GetAuthSession→persona.json→SetAuthSession in a fresh context→signed in
10Scale

Thousands of browsers, ready in milliseconds.

No VM to boot, no container to warm. Isolated contexts on real GPUs, thousands in parallel, and nothing dies with the worker that started it.

one host · pooled GPUsrunningwarmwatched live
<250ms
until a session is ready
1,000+
isolated sessions per job
0
machines to boot or warm
1id
is all another worker needs to attach
01

Contexts, not machines

Isolated browser contexts on shared hosts with real GPUs. Nothing to boot, nothing to warm.

02

Isolated every time

Own cookies, storage, cache, proxy and persona. Thousands of jobs side by side, nothing shared.

03

Outlives its worker

Attach from any process by session id. A redeploy or a crash doesn’t take the browser with it.

11For agents

Explore with MCP. Ship with the CLI.

A model needs a page it can reason about without a megabyte of HTML, and failures specific enough to repair. Then somewhere to put the logic, so it is not re-derived every run.

b.GetObservation(ctx)token-budgeted
# frame 8F03BF2E… https://shop.example/checkout
# title "Checkout"
h1[32] "Checkout"
input#email[44] type="email" name="email" value="ada@…" [required click] "E-Mail"
iframe[51] src="js.stripe.com" frameId=A91C…
input[7] name="cardnumber" value="" [required click] "Card number"
button[9] type="submit" [click] "Pay"
a[63] href="/help" [click] "Need help?"
01

Prompt-sized

One line per element: role, name, live value, label and flags, bounded by a token budget.

02

Every frame in one round trip

Iframes, cross-origin frames and closed shadow roots, each tagged with the frame it lives in.

03

Failures it can repair

The element that blocked a click, the condition that got how far. Input for a fix, not a retry.

Explore: the hosted MCP server

One URL for any MCP client: Cursor, Claude, Codex or your own. The key is a header, never a tool argument.

mcp.json
{  "mcpServers": {    "browserscale": {      "url": "https://mcp.browserscale.cloud/mcp",      "headers": {        "Authorization": "Bearer sk_…"      }    }  }}
23 toolssame verbs as the SDK
rentnavigateobservewaitclickfilltypeselectpress_keyinsert_textmove_toscroll_todragevaluatescreenshotsolve_captchaget_cookiesset_cookiesclear_cookiesget_storageset_storageclear_storagestop
01

The same verbs as the SDK

What the agent proves out by hand becomes the flow it writes, line for line.

02

Sessions outlive the chat

rent returns an id. An agent can attach to what a failed run left behind and look at the real page.

Ship: a project it can build in

One command writes a Go module that already runs. Your agent fills in the flow.

$ browserscale init -name orders-botcompiles and runs
  • orders-bot/
  • main.gohands off to the harness
  • module.goconfig schema + workers
  • flow.gothin dispatcher
  • register.goa worked browser flow
  • run.goqueue loop, 8 workers
  • rent.goproxy pool + backoff
  • AGENTS.mdthe agent reads this first
  • docs/the whole SDK, offline
  • data/proxies.txt · accounts.csv
01

Signatures, not vibes

Every scaffold ships an AGENTS.md and the whole SDK reference offline, so method names are real.

02

It runs and fixes its own work

browserscale dev rebuilds, restarts and streams the logs. The agent reads what happened and iterates.

Why have it write code at all

For exploring a page, clicking through with the MCP tools is exactly right. For work you repeat, it bills per step on every run, and when the chat ends nothing is left.

In parallel
one session per conversationas many workers as you configure
When it breaks
re-prompt and hopelogs, artifacts, a store you can query
What you keep
a chat transcripta program in your repo
model cost over a thousand runslog scale · shape, not a quote
run 1101001,000costAgent clicks through · every step, every runAgent writes the code · oncebreak-even
  • Agent writes the code · once
  • Agent clicks through · every step, every run
  • break-even around run 10
12In production

The half that decides whether it survives Monday.

A flow that works on your machine is the easy part. The rest is a thousand of them a day on shared hardware, and knowing afterwards what each one did.

runs · orders-bot · today8 workers
  • r_8f21w3 · register14.2 s
    artifacts · network log
  • r_8f20w1 · checkout12.9 s
    otp fetched via IMAP
  • r_8f1fw7 · checkout18.0 s
    occluded_after_evade · div#promo-modal
  • ↻r_8f1ew2 · register0.3 s
    rent refused at once · host full
  • r_8f1dw5 · login6.1 s
    persona restored · signed in
  • r_8f1cw4 · checkout9.8 s
    2 static paths from cache
r_8f1f · screenshot · network log · harness logread it, don’t reconstruct it

On shared hardware

01

Fair under contention

Time is divided between sessions and nothing is hard-capped, so a heavy neighbour cannot starve yours.

02

Warm before you ask

Capacity is held ahead of demand. A full host refuses at once instead of leaving the rent hanging.

03

Rotated before it wears

Sessions move onto fresh processes with the replacement ready first. A wedged one is drained, not shared.

In the harness

01

The infrastructure is written

Worker loop, config, queue, proxy pool, rent with backoff and a store per run. You add the flow.

02

Ordinary code in your repo

No workflow format, no designer. Review, diff, test and deploy it like the rest of your software.

03

Nothing to install where it runs

No browser binary on the worker, no Chromium download in CI. You scale sessions, not containers.

In every scaffold
Worker harnessConfig TUIWork queuesProxy poolsSQLite storeIMAP OTP fetchTree loggerPer-run artifacts
example.com/checkout
LIVE
Checkout
Step 2 of 3 · Payment
Email
alice@example.com
Card number
4242 4242 4242 4242
Expiry
09 / 28
CVC
•••
Order summary
Pro plan · annual$192.00
Tax$36.48
Total$228.48
Place order
WebRTC
MouseKeyboardTake over

And a person can step into any run

01

Streamed from inside the browser

A WebRTC peer encoded on the GPU that paints the page, not screenshots. Relay-only by default.

02

Mouse, keyboard and clipboard

Take over from the dashboard or with browserscale view, finish the step by hand, hand it back.

13How it compares

Not another Playwright wrapper in the cloud.

Everyone else drives Chromium from outside over a debugging protocol. browserscale replaced the control plane and rebuilt the parts that kept failing.

  • A click that is blocked

    browserscale
    Re-aims, evades, then refuses and names the blocker
    Playwright / Puppeteer
    Dispatched anyway, or a timeout
    Other cloud browsers
    Same, remotely
  • A wait that times out

    browserscale
    Per-condition diagnosis, with the occluder
    Playwright / Puppeteer
    A deadline and a screenshot
    Other cloud browsers
    A deadline
  • Interruptions (banners, modals)

    browserscale
    Standing reactions the browser carries out
    Playwright / Puppeteer
    Your own retry logic, in every flow
    Other cloud browsers
    Your own retry logic
  • Cross-origin iframes

    browserscale
    Any depth, no frame bookkeeping
    Playwright / Puppeteer
    ~A session and an injected world per frame
    Other cloud browsers
    ~Same, plus the network
  • What a wait costs while waiting

    browserscale
    Nothing until the page changes
    Playwright / Puppeteer
    A poll loop per condition, per frame
    Other cloud browsers
    Same, over the wire
  • Where control lives

    browserscale
    Inside the browser engine
    Playwright / Puppeteer
    Outside, over CDP
    Other cloud browsers
    Playwright/CDP, hosted
  • Fingerprint & GPU

    browserscale
    Real consumer GPUs, no spoof layer
    Playwright / Puppeteer
    ~Whatever your machine is
    Other cloud browsers
    Spoofing layer / hash database
  • Spin-up & scale

    browserscale
    Contexts on pooled GPUs, < 250 ms, thousands
    Playwright / Puppeteer
    Browsers you host and scale
    Other cloud browsers
    ~VMs/containers, cold starts
  • Watch & take over live

    browserscale
    In-browser WebRTC, GPU-encoded, plus input
    Playwright / Puppeteer
    None
    Other cloud browsers
    ~View-only, if any
14The surface

Around fifty commands, one session behind them.

The same commands from both SDKs, the CLI, MCP and inside BrowserVM. Targets are interchangeable: a selector, an expression, a node handle or coordinates.

Pointer & keyboard

10
  • click
  • moveTo
  • drag
  • scrollTo
  • fill
  • type
  • insertText
  • pressKey
  • selectOption
  • getSelection

Waiting & reacting

5
  • wait
  • waitAny
  • addReaction
  • listReactions
  • removeReaction

Network

7
  • waitForRequest
  • waitForResponse
  • modifyRequest
  • setBlockList
  • setStaticPaths
  • loadHTML
  • networkCapture

Identity & state

6
  • getAuthSession
  • setAuthSession
  • cookies
  • storage
  • setProxy
  • fingerprints

Seeing the page

7
  • getObservation
  • getDOM
  • domMirror
  • screenshot
  • readCanvas
  • inspectAtPosition
  • highlightNode

Session & runs

9
  • rent
  • connect
  • getPages
  • navigate
  • evaluate
  • runScript
  • startScript
  • stream
  • solveCaptcha
15FAQ

Questions, answered.

The ones that actually come up: about reliability first, and about everything around the browser after that.

What is browserscale?

browserscale runs real Chromium browsers in the cloud and rebuilds the parts of the browser that browser automation keeps failing on. You rent an isolated session and drive it from Go or TypeScript — or from a script that runs beside the page in BrowserVM — instead of installing, scaling and babysitting your own browser farm. Waits, clicks, typing, frame handling and a session-wide network log behave correctly rather than approximately.

What problem does it actually solve?

The three failures every browser bot dies of: the element is inside a cross-origin iframe, the click did not land and nobody knows why, and the wait was a poll racing the page. browserscale moves all three into the engine — a wait is reported by the document the instant it becomes true, a click verifies the pixel it is about to press and names the element that blocked it, and a frame at any depth is reachable without bookkeeping.

My clicks miss elements or hit the cookie banner instead. What does browserscale do differently?

A click scrolls the element into view through nested scrollers and frames, waits until its bounds stop moving, moves the pointer there along a human path, and then verifies the exact pixel the way a real mouse event is routed — including across process boundaries, so a page-level modal over a cross-origin iframe is caught. If something covers the point it re-aims at the largest exposed part of the element; failing that it steps clear of the whole overlay once, which collapses hover menus, and re-checks. If it still cannot land it refuses instead of clicking the wrong thing, and returns the intercepting element: tag, id, class, text, bounds, pointer-events, stacking order, whether it is pinned, and whether it swallows clicks while invisible.

My waits are flaky. How is waiting different here?

There is no polling loop that can race the page. Each document that a condition concerns watches for it itself and pushes a wake-up the moment it starts holding, so a match arrives within a frame or two — including from a frame in its own process, which reports for itself. The default meaning of a condition is visible, reachable and holding still for half a second, not merely present in the DOM, and reachability is decided by the same hit test a click would do. A wait also survives a navigation and re-attaches to the new document.

What do I get when a wait times out?

A per-condition breakdown instead of a bare deadline: for each condition, whether it never appeared, appeared but was hidden, appeared but was covered — with the element that covered it — or was found but still moving when time ran out. Those four cases need four different fixes, which is why they are reported separately.

Can I wait for several possible outcomes at once?

Yes, and it is the normal way to write a flow here. Pass the success state, the error state, the challenge and the queue page as conditions of one wait; the first to match wins and the result tells you which one it was and in which frame. One call, one deadline, one branch — instead of a nest of short-timeout probes racing each other.

How do I deal with cookie banners, modals and interstitials?

You register a standing reaction: if this ever appears anywhere in the page, click it — or click something beside it, like its close button. The browser carries it out on its own, only while the pointer is idle, so it slots into the gaps of an action that is already retrying. A click blocked by a modal therefore lands, because the reaction dismissed the modal while the click was still working its way in. Reactions are one-shot, watch every frame including frames created later, and survive navigation.

How do I work with iframes and cross-origin frames?

A condition or an action can search one frame or every frame in the page at any depth, same-origin or cross-origin, including frames that appear after the call started. What comes back carries the frame it was found in plus a node handle for the element, so the next action goes to the right place without you switching frames. Inside BrowserVM it goes one step further: a cross-origin child document is reached by reading contentDocument, and an element out of it can be handed straight to a click.

What is BrowserVM?

BrowserVM runs your own script inside the browser, in an isolate of its own beside the page rather than in it. Cross-origin frames become property access, values come back as live objects you can assign to, a step into the document costs microseconds instead of a network round trip — so loops are affordable — and nothing is injected into the page. Every command the session has is callable from in there, on the same session, in the same run.

How is browserscale different from Playwright or Puppeteer?

Playwright and Puppeteer drive the browser from outside over the DevTools protocol. browserscale replaces the control plane: commands are carried out by the browser itself, nothing is injected into the page and there is no DevTools handshake on the connection, which removes the traces sites look for and makes fine-grained work fast. On top of that you get the reliability machinery — occlusion-aware clicks, pushed waits with diagnostics, standing reactions — plus real GPUs, a persistable persona, a session-wide network log that does not pause the page, captcha solving, managed proxies and live remote control. BrowserVM runs your own script beside the page rather than in it.

Can I migrate my Playwright or Puppeteer scripts?

It is not a drop-in replacement — you drive it through the Go or TypeScript SDK rather than the Playwright API. But the concepts map closely: click, fill, wait, evaluate and the network primitives are all there, so porting is mostly mechanical. What changes is that a lot of the retry and frame-handling code around your steps stops being necessary.

What happens to our automation if we stop using browserscale?

Your flow is ordinary Go or TypeScript in your own repository. There is no proprietary workflow format, no designer file, nothing that only opens in our tool — you review, diff, test and deploy it like the rest of your software. What is browserscale-specific is the client calls, so moving means replacing those and putting back the retry and frame-handling code the engine was doing for you. That is a real cost and worth knowing up front; it is not a hostage situation.

Will websites detect that I am automating the browser?

That is a core design goal. Control happens below the page with no injected JavaScript and no DevTools control-plane signals, so page scripts have nothing to observe. Pointer and key events travel the path real hardware takes and carry none of the markers that identify remote-controlled input. Combined with real GPUs, a persona you can pin and bring back, and a browser that presents itself on the wire the way a shipped Chrome does rather than a bare automation build, sessions look like a genuine user's browser rather than a headless bot.

Does browserscale fake WebGL and canvas fingerprints?

No — and that is the point. Sessions run hardware-accelerated on real consumer GPUs that browserscale owns and operates, so canvas and WebGL readbacks return genuinely rendered pixels. There is no spoofing layer and no fingerprint hash database that a new probe could unmask, which is also why sessions hold up as passive checks get deeper.

Will my typing look human, and will it end up in the right field?

Both. Text is typed character by character on the keyboard layout of the session's region, with per-character timing that varies the way a hand does, and composition for layouts that need it. Filling is strictly bound to its target: the field is verified as still holding focus as it types, and if something steals focus half way through it stops and names the element that took it rather than typing the rest of a password into it. There is also a deliberately loose mode for one-time-code boxes that move focus themselves on every digit.

How do I stay logged in across separate runs?

Export the signed-in persona as one object — the account, its refresh token and the device-bound sessions that keep it alive long term — and import it into a fresh context, where it comes up signed in. Cookies and local storage can also be read and written as plain data. A country pins language, locale, timezone and keyboard together, the machine stays the same in every frame and every worker, and the rest of what a shipped Chrome carries comes back with it — so a later run is the same computer returning rather than a new install that happens to hold the same cookies.

Can I inspect, modify, mock or block network traffic?

Yes. Watching is attached to the session before anything loads — not to a tab you later debug — so a new iframe is already in the log before it fires its first request. A capture reports every finished request without pausing the page: every frame, every worker, every hop of a redirect, with the headers that actually went on the wire. Separately you can arm a wait for one call, rewrite it, swallow it, answer a navigation yourself, or drop matching requests by pattern.

Can browserscale solve captchas?

Yes. One call detects the challenge, solves it with browserscale's own solver and wires the token back into the page within the same run — no third-party account or per-solve contract. No token is ever synthesized: the challenge is completed in the live browser and the provider's own JavaScript issues the token exactly as it would for a real user. Passive checks usually pass on their own thanks to real hardware and control the page cannot observe, so an interactive challenge tends to be the exception rather than the rule.

How do I cut proxy bandwidth on repeat runs?

Mark the repeated JS, CSS and image paths as static once. Matching GETs are then served from a server-side cache that is reached directly — not through the session’s proxy — so repeat runs skip both the download and the proxy bill. A miss goes to the network as usual and is stored for next time.

Do I have to bring my own proxies?

Either way works. Pass your own host, port and credentials, or leave them empty and browserscale allocates a managed proxy server-side. You can also switch the proxy mid-session without relaunching the browser.

Can I run many browsers in parallel, and how fast do they start?

Yes — that is what it is built for. Each session is an isolated browser context rather than a fresh VM, sharing a host with its neighbours and rendering on a real GPU. There is no machine to boot: a browser is ready in under 250 ms. Every context has its own cookies, storage, cache, proxy and persona, so large queues never collide on state, and you can fan out from one job to thousands of concurrent sessions without managing browsers yourself.

Does someone else's heavy session slow mine down?

Sessions share a host, so it is a fair question. Under contention the host divides time between sessions rather than letting whichever one is burning CPU take the machine, so a session grinding through a CPU-heavy challenge cannot starve its neighbours. Nothing is hard-capped either: on a quiet host a session uses what is there, which also means no run carries the flat, benchmarkable slowness of a capped browser.

Can another worker pick up a running session?

Yes. The browser lives server-side, so any other worker, process or machine can attach to a running session by its id. The session survives a worker restart or a redeploy instead of dying with the process that launched it.

We already run an RPA platform. Where does browserscale fit?

As the browser tier underneath it, not as a replacement for your orchestrator. Scheduling, approvals and reporting stay where they are; what changes is that the browser step stops being the flaky one. Sessions are addressed by id and outlive the worker that rented them, so they attach cleanly to whatever already owns the queue — and a person can take over a live session by hand when a run needs supervising. If you want the worker tier as well, the CLI scaffolds one as ordinary code you own.

Can I watch a session live and take control?

Every session opens a real peer connection inside the browser and encodes the page on the same GPU that paints it — not a sidecar grabbing screenshots. You watch from the dashboard or the CLI. Mouse and keyboard ride dedicated data channels; paste from your machine inserts into the session, copy brings the remote selection back. The viewport size arrives with the stream so you are never aiming at a stale coordinate space. By default only a relay is used, so the session’s address never leaves the host.

Can I use browserscale with AI browser agents?

Yes. One call returns a compact, token-budgeted view of the interactive elements across every frame — role, type, name, live value, label and flags — each tagged with a node handle you can hand straight to an action. Invisible-but-still-interactive elements are marked rather than dropped. And because failures are structured, a model gets something it can repair instead of a stack trace it can only retry.

Is there an MCP server?

Yes, a hosted one at https://mcp.browserscale.cloud/mcp. It exposes the browser as tools that map one-to-one onto the SDK — rent, navigate, observe, wait, click, fill, type, evaluate, screenshot, solve_captcha and the cookie and storage tools. Add the URL to any MCP client with an Authorization: Bearer header; the key stays in the client config and never becomes a tool argument. The server is stateless, so an agent can also attach to a session a failed run left behind and inspect the real page.

Can an AI coding agent build my automation for me?

That is what the CLI is for. browserscale init scaffolds a Go module that already compiles and runs — worker harness, config schema, lifecycle loop, proxy pool and rent-with-backoff — plus a worked browser flow to pattern-match against. Open the folder in Cursor, Claude, Codex or any other coding agent and describe the job: it fills in the flow against real API signatures, because every scaffold ships an AGENTS.md and the full SDK reference offline. browserscale dev rebuilds, restarts and streams the logs from one command, so the agent can run its own work and fix what broke.

Why have an agent write code instead of just clicking through the browser for me?

For exploring an unfamiliar page, clicking through with the MCP tools is exactly right. For work you repeat it is not: every step costs tokens on every run, you get one session per conversation, and when the chat ends nothing is left. Having the agent write a program moves the logic into code once. After that it runs with as many parallel workers as you configure, with no model in the loop, and you keep something you can review, diff and re-run.

Do I need to install Chromium or run my own browser farm?

No. There is no browser binary on your machine and no infrastructure to manage. You rent a session over the API and drive it remotely; browserscale handles the browsers, proxies, scaling and lifecycle server-side.

Which languages have an SDK?

Official SDKs for Go and TypeScript, both driving the exact same browser API, plus a CLI and a hosted MCP server. If you drive the browser yourself the TypeScript SDK needs nothing but Node; the agent scaffold is Go, so that path wants the Go toolchain installed.

How much does it cost?

Pay-as-you-go and credit-based: you top up, earn bonus credits at higher tiers, and scale without per-seat fees or a monthly minimum. See the pricing page for the current tiers.

How do I get started?

Install the Go or TypeScript SDK, grab an API key from your dashboard, rent a browser and drive it — the first script runs in under five minutes. If you would rather have a coding agent build the automation, install the CLI and run browserscale init to get a project that runs before you write a line.

Rent a browser and drive it.

A first script in five minutes, or a module your agent scaffolds that already runs. Pay as you go, no browser farm on your side.