logo
icon

Firecrawl

Open-source web crawler and scraper API for AI applications, knowledge bases, and data pipelines.

Open-source web crawler and scraper API for AI applications, knowledge bases, and data pipelines.

PlatformZeabur
Deployed12
PublisherzeaburZeabur

Firecrawl

🚀 Firecrawl is an open-source web crawler and scraper API for AI applications, knowledge bases, and data pipelines.

This template deploys the full self-hosted Firecrawl stack:

  • api — main Firecrawl process. Runs harness.js --start-docker, which manages the API server plus worker, extract-worker, and several NUQ queue worker subprocesses inside the same container.
  • playwright — browser automation microservice for JS-rendered pages.
  • redis — cache and rate-limit store.
  • rabbitmq — broker for the NUQ task queue.
  • postgres — Postgres 17 + pg_cron, pre-loaded with the NUQ queue schema.

Resource recommendation

Firecrawl is resource-intensive. The upstream docker-compose.yaml recommends at least 4 vCPU / 8 GB RAM for the api service and 2 vCPU / 4 GB RAM for playwright. We recommend deploying this template on a Pro plan or a dedicated server.

Configuration

Only the public domain is required. Core scraping and crawling work with no LLM configured.

LLM extract (optional, bring your own OpenAI-compatible provider): the deploy form offers two optional fields — OpenAI API Key and OpenAI Base URL. Fill in your provider key to enable Firecrawl's LLM-powered features (llm-extract, summary, /v2/extract, /v2/agent). Leave both blank to deploy without LLM.

  • OpenAI direct: set OPENAI_API_KEY and leave OPENAI_BASE_URL blank — Firecrawl falls back to the OpenAI default endpoint.
  • Other OpenAI-compatible providers (OpenRouter, xAI, Ollama, a self-hosted gateway, …): set both OPENAI_API_KEY and OPENAI_BASE_URL (e.g. https://openrouter.ai/api/v1). You can also set MODEL_NAME on the api service env tab. The full list of supported variables lives in apps/api/src/config.ts.

You can change these any time on the api service env tab and restart the service.

Caveats:

  • Some OpenAI-compatible providers do not expose embedding models. When embeddings are unavailable, crawl link relevance ranking falls back to no-ranking — the rest of the LLM features (extract, agent, summary) are unaffected.
  • A few Firecrawl features (interactive browser-agent, direct-quote handling) call Google Gemini directly via @ai-sdk/google rather than through OPENAI_BASE_URL. To enable them, also set GOOGLE_GENERATIVE_AI_API_KEY on the api service.

BULL_AUTH_KEY (which protects the internal Bull queue dashboard) is auto-generated; you can find it on the api service env tab.

Authentication is disabled by default (USE_DB_AUTHENTICATION=false). The API endpoint is public — protect your domain with Cloudflare, an auth proxy, or a network ACL if exposed to the internet.

Quick test

After deployment finishes, hit the scrape endpoint:

curl -X POST https://<your-domain>/v1/scrape \
  -H 'Content-Type: application/json' \
  -d '{"url": "https://docs.firecrawl.dev"}'

Queue dashboard

Firecrawl has no end-user UI — it is an API-only service. For operators, a built-in Bull Dashboard is mounted at:

https://<your-domain>/admin/<BULL_AUTH_KEY>/queues

Use it to monitor queue throughput, inspect active / completed / failed jobs, and replay errors. Find your BULL_AUTH_KEY on the api service env tab.

Versioning

Firecrawl upstream does not publish semver tags — only :latest, which is rebuilt every time main advances. To give you reproducible deploys, this template pins each image by digest (@sha256:...). The current pin corresponds to the build of 2026-05-04. Updates ship as new template revisions.

Reference