Skip to main content

Authenticated Web Scraping

Authenticated web scraping starts with the right browser context.

Return to an authorized account workflow with a saved browser profile. Inspect the login state, complete any required sign-in and stop the runtime when finished.

01 / OVERVIEW

What is Authenticated Web Scraping?

Authenticated web scraping retrieves content that requires a valid account session. AdsCrawl saved cloud browser profiles support persistent browser session workflows by retaining browser context between runs, while the target website remains responsible for whether a login is accepted.

02 / USE CASES

Where it fits

01

Resume an account workflow

Reuse a saved profile for an account you are authorized to access. Confirm a signed-in page is visible before collecting content.

02

Handle interactive sign-in

Open the authenticated live viewer to inspect login state and complete required manual steps. A saved profile does not bypass MFA or session expiry.

03

Keep account contexts separate

Use separate profiles for workflows that require separate cookies and proxy settings. Avoid copying credentials or session data into public code examples.

03 / QUICKSTART

Start with one complete workflow

  1. Create or launch a saved browser profile with a valid explicit proxy; the example demonstrates the launch and stop lifecycle.
  2. Open runtime.connectUrl while signed in to AdsCrawl and complete the target site's authorized login flow.
  3. Verify the target account and content before continuing your collection workflow; launch alone does not extract data.
  4. Stop the runtime, confirm it is stopped, and keep the profile id for a later start instead of creating another profile.

Requires Bash and cURL. Set the ADSCRAWL_API_KEY environment variable on your server first.

Full request and response reference →
cURL
# Replace the proxy address with your own proxy.
curl --fail-with-body -sS --max-time 200 \
  -X POST "https://api.adscrawl.net/cloud-browsers/launch" \
  -H "x-api-key: $ADSCRAWL_API_KEY" \
  -H "content-type: application/json" \
  -d '{
    "proxy": { "server": "http://proxy.example.com:8080" },
    "tabs": ["https://www.adscrawl.net"]
  }'

# Open runtime.connectUrl from the response.
# When finished, replace <browser-id> with the returned id.
curl --fail-with-body -sS -X POST \
  "https://api.adscrawl.net/cloud-browsers/<browser-id>/stop" \
  -H "x-api-key: $ADSCRAWL_API_KEY"

# Confirm runtime.status is "stopped"; otherwise query/retry stop.
# Closing the viewer does not stop billing.
# If launch fails, inspect its returned id; do not repeat launch.

04 / LIMITS & BILLING

Know the boundaries

Compare plans and allowances →
  • Saved browser context does not guarantee a permanent login. Sites can expire sessions or require another challenge.
  • The example manages a browser profile; you must implement the collection step using a supported workflow for your target.
  • Cloud Browser launch requires a paid plan and an explicit valid proxy for API-key requests.
  • A Cloud Browser profile is not automatically imported into a temporary Remote CDP session. Closing the viewer does not stop billing.

Keep API keys on the server. For HTTP 402 check your balance; for 429 reduce concurrency and retry with backoff.

05 / FAQ

Common questions

Does this bypass a login or MFA?

No. You need an authorized account and must satisfy the site's login requirements.

Is persistent context the same as a running browser?

No. A saved profile and a live runtime have different lifecycles and limits. Stop the runtime when finished.

What if the saved login has expired?

Open the profile, sign in again through the required flow and verify the account state before continuing.