Skip to main content

Web Scraping API

Web Scraping API. Get the fields your application needs.

Render JavaScript pages in a browser and extract a defined set of fields. Give your pipeline structured results instead of another page to parse.

01 / OVERVIEW

What is Web Scraping API?

AdsCrawl Web Scraping API extracts specified fields from browser-rendered websites through POST /spa-extract. Define fields using DOM selectors or supported network response sources. When you need rendered markup or readable page content instead, use POST /html with the appropriate output mode.

02 / USE CASES

Where it fits

01

Extract precise fields

Define CSS selectors for the names, headings or values your application needs. Give each field a stable name so downstream code does not depend on the complete page markup.

02

Read dynamic data

Use supported network fields when the page retrieves JSON while loading. Consult the API reference for matching and extraction rules.

03

Check completeness

Inspect missingFields alongside returned data. Detect changed selectors and incomplete loads before storing results in a monitoring or research pipeline.

03 / QUICKSTART

Start with one complete workflow

  1. Create an API key and store it in a server-side environment variable.
  2. Inspect your target page and identify a selector for each required field.
  3. Send POST /spa-extract with mode set to extract and the fields object.
  4. Check the returned values and missingFields before using the result.

Requires Bash and cURL. Set the ADSCRAWL_API_KEY environment variable on your server first.

Full request and response reference →
cURL
curl --fail-with-body -sS \
  -X POST "https://api.adscrawl.net/spa-extract" \
  -H "x-api-key: $ADSCRAWL_API_KEY" \
  -H "content-type: application/json" \
  -d '{
    "url": "https://www.adscrawl.net",
    "mode": "extract",
    "fields": {
      "title": {
        "source": "dom",
        "selector": "h1",
        "parse": "string"
      }
    }
  }'

04 / LIMITS & BILLING

Know the boundaries

Compare plans and allowances →
  • A target-site redesign can invalidate selectors. Validate the response rather than assuming every request returns every field.
  • Page readiness depends on the target site. Choose wait conditions within the endpoint's supported limits.
  • Use Remote CDP for multi-step navigation, clicks and forms. Field extraction is not a replacement for a browser script.
  • Requests use account credits and are subject to concurrency limits. Check your plan before scheduling a batch.

Keep API keys on the server. For HTTP 402 check your balance; for 429 reduce concurrency and retry with backoff.

05 / FAQ

Common questions

Should I use /html or /spa-extract?

Use /html for rendered HTML or readable article content. Use /spa-extract when you need a specified set of DOM or supported network fields.

What happens when a selector does not match?

Inspect missingFields and the returned data. Check the current selector and wait settings, then update your field definition if the target site's structure has changed.

Does a successful response mean my data is correct?

No. Validate required fields, types and any business rules before storing the result. Rendering and extraction do not establish that the source's claims are accurate.