Firecrawl
Service domainWEB SCRAPING
Arcade OptimizedBYOCPro
Arcade.dev LLM tools for reading the web via Firecrawl
Author:Arcade
Version:
4.0.1Auth:No authentication required
12tools
12require secrets
Firecrawl is a web-reading service; this toolkit lets LLMs scrape, crawl, extract, map, and search the web through Arcade by calling Firecrawl's API.
Capabilities
- Single-page and multi-page reading: Scrape a known URL or crawl an entire site, collecting each page's content; long-running crawls return a job ID for later polling or cancellation.
- Structured data extraction: Describe the fields you want (products, prices, contacts) in plain language and receive structured records; supports async job tracking.
- Site mapping: Enumerate all URLs on a site without downloading page content.
- Crawl lifecycle management: Check progress, retrieve collected pages incrementally, or cancel an active crawl by job ID.
- Web and specialized search: Search the open web (with optional inline page fetching), developer docs and repositories, GitHub issues, and a research-paper corpus (PubMed, bioRxiv, medRxiv, arXiv) — each endpoint targets a distinct index.
Secrets
FIRECRAWL_API_KEY — A Firecrawl API key that authenticates every request to the Firecrawl API. To obtain one, sign in to the Firecrawl dashboard, navigate to API Keys, and generate a new key. Free and paid tiers are available; higher-volume crawls and concurrent extractions may require a paid plan. Copy the key immediately — it is not shown again after creation.
Configure secrets in Arcade at https://docs.arcade.dev/en/guides/create-tools/tool-basics/create-tool-secrets or via the dashboard at https://api.arcade.dev/dashboard/auth/secrets.
Available tools(12)
12 of 12 tools
Operations
Behavior
| Tool name | Description | Secrets | |
|---|---|---|---|
Stop a crawl that is still running.
Pages the crawl already collected stay readable by its job id. | 1 | ||
Read many pages of one website and return each page's content.
A crawl that outruns wait_seconds keeps running: the result carries its
job_id for checking progress, reading the pages, or cancelling it later. | 1 | ||
Collect structured records from the web from a plain-language description.
Use this when the answer is fields such as products, prices, or contacts,
rather than the page text itself. An extraction that outruns wait_seconds
keeps running, and the result carries its job_id for collecting the results
later. | 1 | ||
Read the pages a crawl has collected so far. | 1 | ||
Check how far a crawl has progressed, without fetching its pages. | 1 | ||
Collect the results of an extraction that is already running. | 1 | ||
List the URLs on a website, without reading the pages. | 1 | ||
Read one web page whose URL is already known and return its content.
It reads only the page at the given URL. It does not search the web, crawl
other pages of the site, or pull structured fields out of the page text. | 1 | ||
Search the web and optionally read each result page in the same call.
Use this when no URL is known yet. It searches the open web rather than a
corpus of scientific papers or indexed code documentation. | 1 | ||
Search developer documentation and repositories, quoting the text that matched.
Use this to answer a question about how a library or API works. It searches
indexed repositories, not the open web. | 1 | ||
Search GitHub issues across indexed repositories.
Only issues are searched; documentation, READMEs, and pull requests are not. | 1 | ||
Search published scientific papers and return their abstracts.
This reads a corpus of paper records drawn from PubMed, bioRxiv, medRxiv,
and arXiv. To find ordinary web pages that happen to sit on academic sites,
use Search narrowed to research instead. | 1 |
Last updated on