Browser automation that doesn’t guess.
Real Chromium in the cloud, controlled from inside the engine. From Go, TypeScript, MCP or BrowserVM.
- 00
b.AddReactionWith(ctx, bs.CSS(".newsletter"), bs.ReactionOpts{On: bs.CSS(".close")})
registeringarmed · watched in the browserfired once · closed .newsletter - 01
b.Wait(ctx, bs.CSS("#email"), bs.Timeout(15000))
waiting for the formready in 1.6 s - 02
b.Fill(ctx, bs.CSS("#email"), "ana@example.de")
typing14 keys · de-DE layout - 03
b.Fill(ctx, bs.CSS("[name=card]").InAllFrames(), card.Number)
typing19 keys · frame 2paused for reaction 00 · resumed at key 15 - 04
b.Click(ctx, bs.CSS("#place-order"))
workinglanded on button#place-orderre-aimed past div#cookie-wall
Five calls in your code. The loading, the pointer paths, the typing rhythm, the popup, the pixel check and the re-aim all happen inside them.
Try it now
One package and an API key. Or let the CLI scaffold a module your coding agent builds in.
Every browser bot dies of the same three things.
Nobody who runs automation in production loses time on the API. It goes on these three, the same in every tool, and on what the log says when they happen.
“It works, except sometimes.”
TimeoutError: Timeout 30000ms exceeded. waiting for element to be visible, enabled and stable
WaitError after 25 s #pay-button found_occluded occluder div#promo-modal
“The click didn’t do anything.”
✓ locator.click() resolved // the order was never placed
ClickError occluded_after_evade occluder div#cookie-wall position fixed · z 2147483647
“The element is in an iframe.”
TimeoutError: locator.fill:
waiting for locator('[name=otp]')
// it lives two frames downmatched [name=otp] frame 3 of 6 · depth 2 cross-origin, written in mid-wait
None of these are solved in a helper library. They are solved where the browser already knows the answer.
One wait. Every way the page can go.
One call names every way a step can end and how long it may take. Until one holds, visible, reachable and still, nothing polls and nothing crosses the wire. The page reports its own changes, the check runs inside the browser, and only the answer comes back.
1_, _ = b.Navigate(ctx, "https://shop.example/checkout", 0)2 3// A step can take a while. Say how long, and every way it can go.4res, err := b.Wait(ctx,5 bs.CSS("#order-confirmation"),6 bs.CSS("#address-form"),7 bs.CSS("iframe[title*=verification]").InAllFrames(),8 bs.JS("!!document.querySelector('[data-error]')"),9 bs.Timeout(25000),10)11 12switch res.Index {13case 0: // done; res.FrameId and res.Bounds say where14case 2: // the verification frame, and nothing here named it15}And while it waits, it does nothing
A condition that does not hold yet is not hammered on a timer. The documents it concerns report when they change, and the check runs then, inside the browser, so breadth costs a registration, not a loop.
every 200 ms, whether the page moved or not, and one tick late
the page reports each change, a slow sweep covers the rest, all inside the browser; one message back
And almost none of it crosses the wire
Checked from outside, every question about the layout is a message to the browser and one back: where the element is, what sits on top of it, whether it still moves. Here those checks run where the page is, and only the answer travels.
8 per check, for one condition in one frame, × 2,160 checks. Visible, covered, still moving: each is a question over the protocol and an answer back.
when the page reports a change
The answer, with the element, its frame and its bounds. The ten seconds before it cost the wire nothing.
A click that lands — or says what stopped it.
What usually becomes a few hundred lines of retry logic and a dozen round trips is one call, done where the layout lives.
- 01locate
any frame, for up to 5 s
- 02scroll
through nested scrollers and frames
- 03settle
until the bounds hold still
- 04move
a hand’s path, not a jump
- 05verify
which element owns that pixel
- 06re-aim
exposed patch, or step out of the overlay
frame · node · bounds · where it landed
never presses what lies on top
- code
- occluded_after_evade
- occluder
- div#cookie-wall.banner
- text
- “We value your privacy”
- bounds
- x 0 · y 612 · 1280×188
- pointerEvents
- auto
- position
- fixed
- zIndex
- 2147483647
- hittableWhileInvisible
- false
And typing that stays in the field
Keys on the layout of the session’s region, timed the way a hand types. If something steals focus half way, it stops and names the element instead of typing the rest of a password into it.
The iframe that is not there yet.
The flake is not naming an iframe. It is that the frame does not exist yet. Here the frame set is whatever the page has right now, including frames written in while you wait.
- ▪shop.example/checkout
- └js.pay.example
- └3ds.bank.example
- └ads.example
- └ads.example/slot/4
- └consent.example
Frames that appear mid-wait
Looked at the moment they exist. No wait-for-the-iframe before the wait for what is inside it.
Cross-process is not a special case
No session per frame, no injected world, no depth limit. A grandchild costs what a child costs.
The answer names the frame
A match carries its frame and a node handle. Pass the handle on and the next action goes to the right document.
A cookie bar is not a step in your flow.
Consent bars, newsletter modals, a survey two minutes in. No linear script can say when they appear, so you register a standing instruction and the browser carries it out.
1// The consent bar, wherever in the page it turns up.2_, _ = b.AddReaction(ctx,3 bs.CSS("#accept-all").InAllFrames())4 5// A modal we do not want to accept: click its close button instead.6_, _ = b.AddReactionWith(ctx,7 bs.CSS("div[role=dialog] .newsletter"),8 bs.ReactionOpts{On: bs.CSS("button[aria-label=Close]")})9 10// The rest of the run never mentions either of them again.Fires in the gaps
Armed and carried out in the browser, only while the pointer is idle. A click already retrying under a modal lands.
Once, then it retires
A banner that comes back is not clicked twice by accident. List what is still armed, or remove it by id.
Every frame, every navigation
Including frames created later. Click the match itself, or something beside it such as its close button.
Same job. One side is mostly bookkeeping.
A CDP client can only reach a frame it has attached to and run code in a context it has been told about. So before the checkout even starts, it is keeping a copy of the browser current, fed by events it has to subscribe to. browserscale leaves that state in the browser, where it already is.
- ▪targetpage · session 1Target.attachedToTarget
- └frameshop.example/checkoutPage.frameNavigated
- └contextmain world · id 7Runtime.executionContextCreated
- └contextutility world · id 8Runtime.executionContextCreated
- └frameconsent.examplePage.frameAttached
- └contextmain world · id 9Runtime.executionContextCreated
- └targetjs.pay.example · session 2Target.attachedToTarget
- └framejs.pay.example/v3Page.frameNavigated
- └contextmain world · id 3Runtime.executionContextCreated
- └target3ds.bank.example · session 3attached mid-checkout
- ▪handles18 remote objectsRuntime.releaseObject, by you
- Frames, sessions and contexts stay in the browser, where they already are.
- Every call names the page and a selector; the browser finds the frame.
- An answer carries the frame, node and bounds it is about.
And the copy is kept current over the wire
Every request arrives as five or six network events, every frame reports its lifecycle, every context announces its creation and its end, whether your code cares or not. Here the connection carries what you asked and what came back.
- Network.*≈ 9005–6 per request
- Page.lifecycleEvent≈ 80every frame, every load
- Page.frame*≈ 45attached, navigated, loading
- Runtime.executionContext*≈ 30created, destroyed, per world
- Target.*≈ 8a session per cross-origin frame
- navigate→ ←
- addReaction→ ←
- wait→ ←
- fill→ ←
- click→ ←
- wait→ ←
plus every command and its reply, several per action. The client reads all of it, including what it never asked about.
Events arrive only when you ask for them: a network capture, a DOM mirror.
Undetectable is a consequence of where it runs.
Detection reads the page, the input and the connection, and compares what it finds. When control lives in the engine instead of on a debugging protocol, every answer comes from the browser itself and agrees with the others, with nothing patched over.
Nothing to find
Control never touches the page.
- no debugging protocolthe engine is driven directly, so there is no session for a probe to trip
- nothing in the pageno injected script, no patched prototype, toString stays native
- branded ChromeGoogle Chrome in every brand list, no headless marker anywhere
- a real GPUWebGL and canvas come from the hardware that draws them
Input like a hand
Delivered the way hardware delivers it.
- pointer pathscurved and uneven, never a jump to the centre
- frame cadencemoves arrive on the display’s rhythm, jitter included
- coalesced samplesseveral samples per frame, like a real mouse
- native keystrokesthe region’s layout, dead keys and AltGr included
- IME typingJapanese goes through the IME as romaji, key by key
One profile, no contradictions
One region, every frame and worker.
- languagenavigator.languages, Accept-Language and Intl agree
- time and placetimezone and locale match the exit IP
- keyboard mapthe layout map matches the keys actually typed
- speech voicesthe voice list a machine in that region has
- WebRTCICE candidates show the exit IP, not the host
- screenone monitor geometry and scrollbar width throughout
Chrome on the wire
Exactly what a stock install sends.
- network stackChrome’s own TLS and HTTP/2, because it is Chrome’s
- Client Hintsbrands and full version agree with the User-Agent
- request headersmatching a regular Chrome installation
- background requestscarry the same values as the page’s own
No stealth plugin and no patch list to keep current. A detector compares what it can read, and none of it was ever made to disagree.
Or put your code inside the browser.
Automation has lived in two places: in the page, fast but visible, or outside the browser, invisible but a round trip per question. BrowserVM is a third, an isolate beside the page.
Fast — and the site can see it.
Invisible — and a round trip per question.
Beside the page. Neither.
1// The processor's own cross-origin iframe, in its own process.2const pay = window.document3 .querySelector('iframe[name=pay]').contentDocument;4 5// A live element from that other document, handed to a real click.6await browser.fill(pay.querySelector('[name=number]'), card.number);7 8// Armed before the click that causes it, awaited after. No race:9// the script never left the browser in between.10const token = browser.waitForResponse('*/api/challenge*');11await browser.click(css('#verify'));12const body = (await token).response.body;Cross-origin is a dot
Read contentDocument of the processor’s frame and hand its elements straight to a click. No frame ids.
Microseconds per step
Loops over tables and pagination are affordable, and arm-then-trigger has no race in it.
Nothing in the page
No global, no patched prototype, no evaluated source. Code in the page stays a separate, explicit call.
Every request through you. Every login portable.
Watching is attached to the session, not a tab, so every iframe is in the log before its first request. And the browser’s identity, login included, is one object you take to the next run.
- GET/checkoutmain200network
- GET/static/app.9f2c.jsmain200static cache
- POST/api/cartmain200network
- GETjs.pay.example/v3frame 2200http cache
- GETjs.pay.example/innerframe 2200network
- GET/sw/precache.jsonworker200service worker
- POST3ds.bank.example/authframe 3302network
- GET3ds.bank.example/challengeframe 3200network
On before the first request
Armed on the session, not a tab. A new iframe is in the log before it loads anything.
A full log that never pauses the page
Every frame, worker and redirect hop, with the headers that went on the wire. Bodies are copied off to the side.
Or catch one call
Wait for a request, then read, rewrite or swallow it, or answer a whole navigation yourself.
Static paths, paid for once
Repeat JS, CSS and images come from a server-side cache that skips the proxy and its bill.
Who the browser is, kept between runs
A login you can carry
Import it into a fresh context and it comes up signed in, without a profile directory to copy around.
The same machine, returning
Pin it once and the next run is the same computer coming back, not a new install holding the same cookies.
A shipped Chrome, not a build of one
The state a consumer browser carries is filled in, consistent with the region the session exits from.
- region
- DE · language de-DE · Europe/Berlin
- keyboard
- German, QWERTZ
- machine
- 8 cores · 16 GB · same in every frame and worker
- renderer
- real GPU, pinned to this persona
- webrtc
- public IP matches the exit
- state
- 214 cookies · storage for 9 origins
- account
- signed in · refresh token · device-bound session
Thousands of browsers, ready in milliseconds.
No VM to boot, no container to warm. Isolated contexts on real GPUs, thousands in parallel, and nothing dies with the worker that started it.
- <250ms
- until a session is ready
- 1,000+
- isolated sessions per job
- 0
- machines to boot or warm
- 1id
- is all another worker needs to attach
Contexts, not machines
Isolated browser contexts on shared hosts with real GPUs. Nothing to boot, nothing to warm.
Isolated every time
Own cookies, storage, cache, proxy and persona. Thousands of jobs side by side, nothing shared.
Outlives its worker
Attach from any process by session id. A redeploy or a crash doesn’t take the browser with it.
Explore with MCP. Ship with the CLI.
A model needs a page it can reason about without a megabyte of HTML, and failures specific enough to repair. Then somewhere to put the logic, so it is not re-derived every run.
Prompt-sized
One line per element: role, name, live value, label and flags, bounded by a token budget.
Every frame in one round trip
Iframes, cross-origin frames and closed shadow roots, each tagged with the frame it lives in.
Failures it can repair
The element that blocked a click, the condition that got how far. Input for a fix, not a retry.
Explore: the hosted MCP server
One URL for any MCP client: Cursor, Claude, Codex or your own. The key is a header, never a tool argument.
{ "mcpServers": { "browserscale": { "url": "https://mcp.browserscale.cloud/mcp", "headers": { "Authorization": "Bearer sk_…" } } }}The same verbs as the SDK
What the agent proves out by hand becomes the flow it writes, line for line.
Sessions outlive the chat
rent returns an id. An agent can attach to what a failed run left behind and look at the real page.
Ship: a project it can build in
One command writes a Go module that already runs. Your agent fills in the flow.
- orders-bot/
- main.gohands off to the harness
- module.goconfig schema + workers
- flow.gothin dispatcher
- register.goa worked browser flow
- run.goqueue loop, 8 workers
- rent.goproxy pool + backoff
- AGENTS.mdthe agent reads this first
- docs/the whole SDK, offline
- data/proxies.txt · accounts.csv
Signatures, not vibes
Every scaffold ships an AGENTS.md and the whole SDK reference offline, so method names are real.
It runs and fixes its own work
browserscale dev rebuilds, restarts and streams the logs. The agent reads what happened and iterates.
Why have it write code at all
For exploring a page, clicking through with the MCP tools is exactly right. For work you repeat, it bills per step on every run, and when the chat ends nothing is left.
- In parallel
- one session per conversationas many workers as you configure
- When it breaks
- re-prompt and hopelogs, artifacts, a store you can query
- What you keep
- a chat transcripta program in your repo
- Agent writes the code · once
- Agent clicks through · every step, every run
- break-even around run 10
The half that decides whether it survives Monday.
A flow that works on your machine is the easy part. The rest is a thousand of them a day on shared hardware, and knowing afterwards what each one did.
- r_8f21w3 · register14.2 sartifacts · network log
- r_8f20w1 · checkout12.9 sotp fetched via IMAP
- r_8f1fw7 · checkout18.0 soccluded_after_evade · div#promo-modal
- ↻r_8f1ew2 · register0.3 srent refused at once · host full
- r_8f1dw5 · login6.1 spersona restored · signed in
- r_8f1cw4 · checkout9.8 s2 static paths from cache
On shared hardware
Fair under contention
Time is divided between sessions and nothing is hard-capped, so a heavy neighbour cannot starve yours.
Warm before you ask
Capacity is held ahead of demand. A full host refuses at once instead of leaving the rent hanging.
Rotated before it wears
Sessions move onto fresh processes with the replacement ready first. A wedged one is drained, not shared.
In the harness
The infrastructure is written
Worker loop, config, queue, proxy pool, rent with backoff and a store per run. You add the flow.
Ordinary code in your repo
No workflow format, no designer. Review, diff, test and deploy it like the rest of your software.
Nothing to install where it runs
No browser binary on the worker, no Chromium download in CI. You scale sessions, not containers.
And a person can step into any run
Streamed from inside the browser
A WebRTC peer encoded on the GPU that paints the page, not screenshots. Relay-only by default.
Mouse, keyboard and clipboard
Take over from the dashboard or with browserscale view, finish the step by hand, hand it back.
Not another Playwright wrapper in the cloud.
Everyone else drives Chromium from outside over a debugging protocol. browserscale replaced the control plane and rebuilt the parts that kept failing.
A click that is blocked
- browserscale
- Re-aims, evades, then refuses and names the blocker
- Playwright / Puppeteer
- Dispatched anyway, or a timeout
- Other cloud browsers
- Same, remotely
A wait that times out
- browserscale
- Per-condition diagnosis, with the occluder
- Playwright / Puppeteer
- A deadline and a screenshot
- Other cloud browsers
- A deadline
Interruptions (banners, modals)
- browserscale
- Standing reactions the browser carries out
- Playwright / Puppeteer
- Your own retry logic, in every flow
- Other cloud browsers
- Your own retry logic
Cross-origin iframes
- browserscale
- Any depth, no frame bookkeeping
- Playwright / Puppeteer
- ~A session and an injected world per frame
- Other cloud browsers
- ~Same, plus the network
What a wait costs while waiting
- browserscale
- Nothing until the page changes
- Playwright / Puppeteer
- A poll loop per condition, per frame
- Other cloud browsers
- Same, over the wire
Where control lives
- browserscale
- Inside the browser engine
- Playwright / Puppeteer
- Outside, over CDP
- Other cloud browsers
- Playwright/CDP, hosted
Fingerprint & GPU
- browserscale
- Real consumer GPUs, no spoof layer
- Playwright / Puppeteer
- ~Whatever your machine is
- Other cloud browsers
- Spoofing layer / hash database
Spin-up & scale
- browserscale
- Contexts on pooled GPUs, < 250 ms, thousands
- Playwright / Puppeteer
- Browsers you host and scale
- Other cloud browsers
- ~VMs/containers, cold starts
Watch & take over live
- browserscale
- In-browser WebRTC, GPU-encoded, plus input
- Playwright / Puppeteer
- None
- Other cloud browsers
- ~View-only, if any
Around fifty commands, one session behind them.
The same commands from both SDKs, the CLI, MCP and inside BrowserVM. Targets are interchangeable: a selector, an expression, a node handle or coordinates.
Pointer & keyboard
10- click
- moveTo
- drag
- scrollTo
- fill
- type
- insertText
- pressKey
- selectOption
- getSelection
Waiting & reacting
5- wait
- waitAny
- addReaction
- listReactions
- removeReaction
Network
7- waitForRequest
- waitForResponse
- modifyRequest
- setBlockList
- setStaticPaths
- loadHTML
- networkCapture
Identity & state
6- getAuthSession
- setAuthSession
- cookies
- storage
- setProxy
- fingerprints
Seeing the page
7- getObservation
- getDOM
- domMirror
- screenshot
- readCanvas
- inspectAtPosition
- highlightNode
Session & runs
9- rent
- connect
- getPages
- navigate
- evaluate
- runScript
- startScript
- stream
- solveCaptcha
Questions, answered.
The ones that actually come up: about reliability first, and about everything around the browser after that.
What is browserscale?
browserscale runs real Chromium browsers in the cloud and rebuilds the parts of the browser that browser automation keeps failing on. You rent an isolated session and drive it from Go or TypeScript — or from a script that runs beside the page in BrowserVM — instead of installing, scaling and babysitting your own browser farm. Waits, clicks, typing, frame handling and a session-wide network log behave correctly rather than approximately.
What problem does it actually solve?
The three failures every browser bot dies of: the element is inside a cross-origin iframe, the click did not land and nobody knows why, and the wait was a poll racing the page. browserscale moves all three into the engine — a wait is reported by the document the instant it becomes true, a click verifies the pixel it is about to press and names the element that blocked it, and a frame at any depth is reachable without bookkeeping.
My clicks miss elements or hit the cookie banner instead. What does browserscale do differently?
A click scrolls the element into view through nested scrollers and frames, waits until its bounds stop moving, moves the pointer there along a human path, and then verifies the exact pixel the way a real mouse event is routed — including across process boundaries, so a page-level modal over a cross-origin iframe is caught. If something covers the point it re-aims at the largest exposed part of the element; failing that it steps clear of the whole overlay once, which collapses hover menus, and re-checks. If it still cannot land it refuses instead of clicking the wrong thing, and returns the intercepting element: tag, id, class, text, bounds, pointer-events, stacking order, whether it is pinned, and whether it swallows clicks while invisible.
My waits are flaky. How is waiting different here?
There is no polling loop that can race the page. Each document that a condition concerns watches for it itself and pushes a wake-up the moment it starts holding, so a match arrives within a frame or two — including from a frame in its own process, which reports for itself. The default meaning of a condition is visible, reachable and holding still for half a second, not merely present in the DOM, and reachability is decided by the same hit test a click would do. A wait also survives a navigation and re-attaches to the new document.
What do I get when a wait times out?
A per-condition breakdown instead of a bare deadline: for each condition, whether it never appeared, appeared but was hidden, appeared but was covered — with the element that covered it — or was found but still moving when time ran out. Those four cases need four different fixes, which is why they are reported separately.
Can I wait for several possible outcomes at once?
Yes, and it is the normal way to write a flow here. Pass the success state, the error state, the challenge and the queue page as conditions of one wait; the first to match wins and the result tells you which one it was and in which frame. One call, one deadline, one branch — instead of a nest of short-timeout probes racing each other.
How do I deal with cookie banners, modals and interstitials?
You register a standing reaction: if this ever appears anywhere in the page, click it — or click something beside it, like its close button. The browser carries it out on its own, only while the pointer is idle, so it slots into the gaps of an action that is already retrying. A click blocked by a modal therefore lands, because the reaction dismissed the modal while the click was still working its way in. Reactions are one-shot, watch every frame including frames created later, and survive navigation.
How do I work with iframes and cross-origin frames?
A condition or an action can search one frame or every frame in the page at any depth, same-origin or cross-origin, including frames that appear after the call started. What comes back carries the frame it was found in plus a node handle for the element, so the next action goes to the right place without you switching frames. Inside BrowserVM it goes one step further: a cross-origin child document is reached by reading contentDocument, and an element out of it can be handed straight to a click.
What is BrowserVM?
BrowserVM runs your own script inside the browser, in an isolate of its own beside the page rather than in it. Cross-origin frames become property access, values come back as live objects you can assign to, a step into the document costs microseconds instead of a network round trip — so loops are affordable — and nothing is injected into the page. Every command the session has is callable from in there, on the same session, in the same run.
How is browserscale different from Playwright or Puppeteer?
Playwright and Puppeteer drive the browser from outside over the DevTools protocol. browserscale replaces the control plane: commands are carried out by the browser itself, nothing is injected into the page and there is no DevTools handshake on the connection, which removes the traces sites look for and makes fine-grained work fast. On top of that you get the reliability machinery — occlusion-aware clicks, pushed waits with diagnostics, standing reactions — plus real GPUs, a persistable persona, a session-wide network log that does not pause the page, captcha solving, managed proxies and live remote control. BrowserVM runs your own script beside the page rather than in it.
Can I migrate my Playwright or Puppeteer scripts?
It is not a drop-in replacement — you drive it through the Go or TypeScript SDK rather than the Playwright API. But the concepts map closely: click, fill, wait, evaluate and the network primitives are all there, so porting is mostly mechanical. What changes is that a lot of the retry and frame-handling code around your steps stops being necessary.
What happens to our automation if we stop using browserscale?
Your flow is ordinary Go or TypeScript in your own repository. There is no proprietary workflow format, no designer file, nothing that only opens in our tool — you review, diff, test and deploy it like the rest of your software. What is browserscale-specific is the client calls, so moving means replacing those and putting back the retry and frame-handling code the engine was doing for you. That is a real cost and worth knowing up front; it is not a hostage situation.
Will websites detect that I am automating the browser?
That is a core design goal. Control happens below the page with no injected JavaScript and no DevTools control-plane signals, so page scripts have nothing to observe. Pointer and key events travel the path real hardware takes and carry none of the markers that identify remote-controlled input. Combined with real GPUs, a persona you can pin and bring back, and a browser that presents itself on the wire the way a shipped Chrome does rather than a bare automation build, sessions look like a genuine user's browser rather than a headless bot.
Does browserscale fake WebGL and canvas fingerprints?
No — and that is the point. Sessions run hardware-accelerated on real consumer GPUs that browserscale owns and operates, so canvas and WebGL readbacks return genuinely rendered pixels. There is no spoofing layer and no fingerprint hash database that a new probe could unmask, which is also why sessions hold up as passive checks get deeper.
Will my typing look human, and will it end up in the right field?
Both. Text is typed character by character on the keyboard layout of the session's region, with per-character timing that varies the way a hand does, and composition for layouts that need it. Filling is strictly bound to its target: the field is verified as still holding focus as it types, and if something steals focus half way through it stops and names the element that took it rather than typing the rest of a password into it. There is also a deliberately loose mode for one-time-code boxes that move focus themselves on every digit.
How do I stay logged in across separate runs?
Export the signed-in persona as one object — the account, its refresh token and the device-bound sessions that keep it alive long term — and import it into a fresh context, where it comes up signed in. Cookies and local storage can also be read and written as plain data. A country pins language, locale, timezone and keyboard together, the machine stays the same in every frame and every worker, and the rest of what a shipped Chrome carries comes back with it — so a later run is the same computer returning rather than a new install that happens to hold the same cookies.
Can I inspect, modify, mock or block network traffic?
Yes. Watching is attached to the session before anything loads — not to a tab you later debug — so a new iframe is already in the log before it fires its first request. A capture reports every finished request without pausing the page: every frame, every worker, every hop of a redirect, with the headers that actually went on the wire. Separately you can arm a wait for one call, rewrite it, swallow it, answer a navigation yourself, or drop matching requests by pattern.
Can browserscale solve captchas?
Yes. One call detects the challenge, solves it with browserscale's own solver and wires the token back into the page within the same run — no third-party account or per-solve contract. No token is ever synthesized: the challenge is completed in the live browser and the provider's own JavaScript issues the token exactly as it would for a real user. Passive checks usually pass on their own thanks to real hardware and control the page cannot observe, so an interactive challenge tends to be the exception rather than the rule.
How do I cut proxy bandwidth on repeat runs?
Mark the repeated JS, CSS and image paths as static once. Matching GETs are then served from a server-side cache that is reached directly — not through the session’s proxy — so repeat runs skip both the download and the proxy bill. A miss goes to the network as usual and is stored for next time.
Do I have to bring my own proxies?
Either way works. Pass your own host, port and credentials, or leave them empty and browserscale allocates a managed proxy server-side. You can also switch the proxy mid-session without relaunching the browser.
Can I run many browsers in parallel, and how fast do they start?
Yes — that is what it is built for. Each session is an isolated browser context rather than a fresh VM, sharing a host with its neighbours and rendering on a real GPU. There is no machine to boot: a browser is ready in under 250 ms. Every context has its own cookies, storage, cache, proxy and persona, so large queues never collide on state, and you can fan out from one job to thousands of concurrent sessions without managing browsers yourself.
Does someone else's heavy session slow mine down?
Sessions share a host, so it is a fair question. Under contention the host divides time between sessions rather than letting whichever one is burning CPU take the machine, so a session grinding through a CPU-heavy challenge cannot starve its neighbours. Nothing is hard-capped either: on a quiet host a session uses what is there, which also means no run carries the flat, benchmarkable slowness of a capped browser.
Can another worker pick up a running session?
Yes. The browser lives server-side, so any other worker, process or machine can attach to a running session by its id. The session survives a worker restart or a redeploy instead of dying with the process that launched it.
We already run an RPA platform. Where does browserscale fit?
As the browser tier underneath it, not as a replacement for your orchestrator. Scheduling, approvals and reporting stay where they are; what changes is that the browser step stops being the flaky one. Sessions are addressed by id and outlive the worker that rented them, so they attach cleanly to whatever already owns the queue — and a person can take over a live session by hand when a run needs supervising. If you want the worker tier as well, the CLI scaffolds one as ordinary code you own.
Can I watch a session live and take control?
Every session opens a real peer connection inside the browser and encodes the page on the same GPU that paints it — not a sidecar grabbing screenshots. You watch from the dashboard or the CLI. Mouse and keyboard ride dedicated data channels; paste from your machine inserts into the session, copy brings the remote selection back. The viewport size arrives with the stream so you are never aiming at a stale coordinate space. By default only a relay is used, so the session’s address never leaves the host.
Can I use browserscale with AI browser agents?
Yes. One call returns a compact, token-budgeted view of the interactive elements across every frame — role, type, name, live value, label and flags — each tagged with a node handle you can hand straight to an action. Invisible-but-still-interactive elements are marked rather than dropped. And because failures are structured, a model gets something it can repair instead of a stack trace it can only retry.
Is there an MCP server?
Yes, a hosted one at https://mcp.browserscale.cloud/mcp. It exposes the browser as tools that map one-to-one onto the SDK — rent, navigate, observe, wait, click, fill, type, evaluate, screenshot, solve_captcha and the cookie and storage tools. Add the URL to any MCP client with an Authorization: Bearer header; the key stays in the client config and never becomes a tool argument. The server is stateless, so an agent can also attach to a session a failed run left behind and inspect the real page.
Can an AI coding agent build my automation for me?
That is what the CLI is for. browserscale init scaffolds a Go module that already compiles and runs — worker harness, config schema, lifecycle loop, proxy pool and rent-with-backoff — plus a worked browser flow to pattern-match against. Open the folder in Cursor, Claude, Codex or any other coding agent and describe the job: it fills in the flow against real API signatures, because every scaffold ships an AGENTS.md and the full SDK reference offline. browserscale dev rebuilds, restarts and streams the logs from one command, so the agent can run its own work and fix what broke.
Why have an agent write code instead of just clicking through the browser for me?
For exploring an unfamiliar page, clicking through with the MCP tools is exactly right. For work you repeat it is not: every step costs tokens on every run, you get one session per conversation, and when the chat ends nothing is left. Having the agent write a program moves the logic into code once. After that it runs with as many parallel workers as you configure, with no model in the loop, and you keep something you can review, diff and re-run.
Do I need to install Chromium or run my own browser farm?
No. There is no browser binary on your machine and no infrastructure to manage. You rent a session over the API and drive it remotely; browserscale handles the browsers, proxies, scaling and lifecycle server-side.
Which languages have an SDK?
Official SDKs for Go and TypeScript, both driving the exact same browser API, plus a CLI and a hosted MCP server. If you drive the browser yourself the TypeScript SDK needs nothing but Node; the agent scaffold is Go, so that path wants the Go toolchain installed.
How much does it cost?
Pay-as-you-go and credit-based: you top up, earn bonus credits at higher tiers, and scale without per-seat fees or a monthly minimum. See the pricing page for the current tiers.
How do I get started?
Install the Go or TypeScript SDK, grab an API key from your dashboard, rent a browser and drive it — the first script runs in under five minutes. If you would rather have a coding agent build the automation, install the CLI and run browserscale init to get a project that runs before you write a line.
Rent a browser and drive it.
A first script in five minutes, or a module your agent scaffolds that already runs. Pay as you go, no browser farm on your side.