Fireplexity

Guide

Self-hosting a Perplexity alternative,
end to end.

Running an answer engine on your laptop takes four minutes. Running one that a team relies on is a different job, with a monthly bill and a list of things that break. Here are three ways to do it, what each costs at three usage levels, and where each one gives out first.

The short version

Cheapest that works

Fireplexity on Firecrawl's free plan and Groq's free tier: about 60–100 queries a month for $0. Past that, a team pays a little more than the same number of Perplexity Pro seats would cost, and a product pays about 1.5¢ a query.

Nothing leaves your machines

Vane with Ollama, or Fireplexity on self-hosted Firecrawl and a local model. No per-query cost, but you need about 16 GB of memory for the model, and renting that GPU around the clock costs $250–540 a month.

Self-hosting is a control decision, not a savings plan. The numbers below are there so you can make it with your eyes open.

Three ways to self-host

“Self-hosted” covers a wide range. At one end you run the app and rent everything it calls; at the other, no query, page or token leaves hardware you control. Most deployments sit in one of three places.

Path A

Fireplexity on managed APIs

The app runs on your server. Search and scraping go to Firecrawl's cloud, inference to Groq. The way the repo ships.

~1.1–1.7¢/query + hosting
Path B

Fireplexity, fully local

Self-hosted Firecrawl with your own SearXNG for search, and a model served by Ollama. Two code changes.

$0/query + servers + GPU
Path C

Vane with a local model

One Docker container with SearXNG bundled, pointed at Ollama. The quickest fully local option.

$0/query + GPU

Paths A and B use the same Fireplexity app, so you can start on A and move pieces to B later. Path C is a different app: it searches through a metasearch engine rather than scraping full pages for each query. Our comparison of Fireplexity, Vane and the other alternatives covers that difference. This page is about running them.

What one query actually does

Costs follow from the code, so start there. Fireplexity's whole backend is one file, app/api/fireplexity/search/route.ts. For each question it:

  • Makes one call to Firecrawl's /v2/search, asking for web, news and image results with limit: 6 and each result scraped to markdown. With several sources, Firecrawl applies the limit per source type, so that is up to 18 results. Search costs 2 credits per 10 results, so 4 credits, and each scraped page 1 more: 6 for the web results, and up to 6 more if news results are scraped too. 10–16 credits in all.
  • Trims every scraped page to about 2,000 characters around the question's keywords, then streams an answer from the model. That is roughly 4,000 tokens in and 800 out.
  • Makes a second model call to write five follow-up questions, from the first 500 characters of the answer and the source titles. Roughly 500 tokens in and 100 out. It is easy to miss, and it means every question is two model requests.

There is no database, queue or cache. Two environment variables and a Node process are the whole deployment, which is the reason it is easy to self-host and also the reason it has none of the protection a public service needs (more on that below).

One query, Path A, at list prices

  • Firecrawl search — up to 18 results, 2 credits per 104 credits
  • Firecrawl scrapes — 6 web results, plus news if scraped6–12 credits
  • 10–16 credits on Standard ($99 / 100k, billed monthly)$0.0099–0.0158
  • Groq answer — 4k in @ $0.15/M, 800 out @ $0.60/M$0.0011
  • Groq follow-ups — ~500 in, ~100 out$0.0001
  • Total per query≈ $0.011–0.017

On the Hobby plan ($19 for 5,000 credits) the same credits cost $0.038–0.061, so a query is about 4–6¢. The plan you are on matters several times more than the model you choose.

Path A: Fireplexity on managed APIs

This is the repo as shipped, with one fix. Groq retired the Kimi K2 model the code names; openai/gpt-oss-120b is the replacement it recommends, and the code calls the model twice, so the change has to catch both.

On a server with Node 18+
$ git clone https://github.com/firecrawl/fireplexity.git && cd fireplexity
$ npm ci
$ cp .env.example .env.local      # set FIRECRAWL_API_KEY and GROQ_API_KEY
$ perl -pi -e 's#moonshotai/kimi-k2-instruct#openai/gpt-oss-120b#g' \
    app/api/fireplexity/search/route.ts
$ npm run build && npm start     # port 3000; put Caddy or nginx in front for TLS

Or push the repo to Vercel, which needs no build configuration. Both work; the choice turns on who is going to use it.

Where to host the app

Hosting the Fireplexity app — October 2026
WhereMonthlyFitsWatch for
Your own machine$0Trying it outNobody else can reach it, which is the point
Vercel Hobby$0Personal use onlyTerms forbid commercial use. Over a usage cap you wait out the 30 days; you can't buy more
Vercel Pro$20 per developer seat, $20 of usage includedTeams and productsUsage above the included credit is billed on top
Small VPS (Hetzner CX33, 4 vCPU / 8 GB)€8.49 + VATAnything, if you will patch itListed as unavailable when we checked; the CPX equivalents cost several times more

Function time is not a problem on either Vercel plan. With Fluid compute, which is on by default, functions run for up to 300 seconds on Hobby and up to 800 on Pro. A streamed answer finishes in a fraction of that. The constraint on Hobby is the licence, not the runtime: an internal tool at your company counts as commercial use.

The free tier, and where it stops

Firecrawl's free plan gives 1,000 credits a month, refreshed monthly, with no card. At 10–16 credits a query that is 60 to 100 questions. Groq's free tier allows 30 requests a minute but only 8,000 tokens a minute and 200,000 a day for gpt-oss-120b. One Fireplexity question uses about 5,400 tokens across its two calls, so the free tier manages one question a minute and about 37 a day. A second person asking in the same minute gets a rate-limit error.

For one person, that is a perfectly usable search engine at no cost. For two or more, plan to pay.

Paid, at three usage levels

Credits are sold in plans, not by the query, so real monthly cost moves in steps:

Path A, monthly, at list prices — October 2026
UsageFirecrawlGroqHostingTotal
Solo — 60 queriesFree (600–960 of 1,000 credits)Free tier$0$0
Team — 3,000 queriesStandard, $99 (30–48k of 100k credits)~$4, paid tier~$10–20~$115–125
Product — 30,000 queriesGrowth, $399 (300–480k of 500k credits)~$36from $20~$455

Two things stand out. The team row costs a little more than five $20 Perplexity Pro seats, for a much smaller product, so a team should self-host for control or privacy, not to save money. The product row works out at about 1.5¢ a query. Billing Firecrawl annually ($83 for Standard, $333 for Growth) takes about a sixth off the Firecrawl line. Perplexity's Sonar API, the metered option you would otherwise build on, comes to roughly 1.9–2.7¢ a query for Sonar Pro, so self-hosting is cheaper there, but not by a wide margin.

Levers that change the bill

Scrape fewer pages. The limit: 6 applies to web, news and images separately, and scraping is most of the bill. Dropping news from the sources saves up to six credits a query. Lowering the limit to 4 saves two to four more. Answers rarely cite all six web pages anyway.

Use the smaller model. gpt-oss-20b is half the price of 120b on Groq ($0.075 / $0.30 per million). The whole model bill is about a tenth of the total, though, so this saves less than one fewer scrape.

Drop the follow-up questions. Half the model requests and about a tenth of the tokens go away. That matters most on Groq's free tier, where the daily token cap is what you hit first.

Path B: Fireplexity, fully local

Moving to Path B means taking Firecrawl and Groq out of the loop. Both are possible, neither is a setting, and each costs you something.

1. Run Firecrawl yourself

Firecrawl's engine is open source under AGPL-3.0. Running it unmodified for yourself is straightforward. If you change it and offer it to others over a network, the AGPL requires you to publish your changes. Fireplexity itself is MIT, so its licence does not change.

The repo's Docker Compose file starts seven services: the API and its workers, a Playwright browser service, Redis, RabbitMQ, Postgres and FoundationDB (with a separate init container). Only port 3002 is exposed. The file caps the API at 4 CPUs and 8 GB and Playwright at 2 CPUs and 4 GB. Firecrawl says those caps are not tested minimums and publishes no recommended host size. Our estimate is a 16 GB machine such as Hetzner's CX43 at €15.99 a month, and you should check it against your own load.

What you lose is the part that makes scraping work on hard sites. The self-hosted docs say Fire-engine is not included, so pages load through the bundled Playwright with a basic fetch fallback. Screenshots, page actions and advanced anti-bot handling depend on Fire-engine. Expect sites behind aggressive bot protection to come back empty.

Search works differently too. The self-hosted /search tries Fire-engine if it is configured, then SearXNG if you set SEARXNG_ENDPOINT, and otherwise falls back to DuckDuckGo. If a search errors it returns an empty list, not an error. Fireplexity then answers from no sources, and nothing tells you why. Run your own SearXNG and set the endpoint (the SearXNG section below covers the catch). Fireplexity also asks for news and image results. Check what your self-hosted search returns for those before you rely on them.

Then point Fireplexity at it. The Firecrawl address is hardcoded, and the route refuses to run without a key, so give it a placeholder:

app/api/fireplexity/search/route.ts
- const searchResponse = await fetch('https://api.firecrawl.dev/v2/search', {
+ const searchResponse = await fetch(`${process.env.FIRECRAWL_API_URL}/v2/search`, {

# .env.local
FIRECRAWL_API_URL=http://localhost:3002
FIRECRAWL_API_KEY=local      # any value; the route only checks that one is set

2. Serve the model yourself

Ollama exposes an OpenAI-compatible API on port 11434, so the change is a provider swap. Keeping the variable name means the two model calls further down work unchanged:

route.ts — after npm i @ai-sdk/openai-compatible
- import { createGroq } from '@ai-sdk/groq'
+ import { createOpenAICompatible } from '@ai-sdk/openai-compatible'

- const groq = createGroq({ apiKey: groqApiKey })
+ const groq = createOpenAICompatible({ name: 'ollama', baseURL: 'http://localhost:11434/v1' })

  // and in both calls: groq('gpt-oss:20b')
  // GROQ_API_KEY still has to be set to something: the route checks for it

The model is the real cost. Ollama's gpt-oss:20b is a 14 GB download that runs in about 16 GB of memory. gpt-oss:120b is 65 GB and wants a single 80 GB GPU. On 16–24 GB consumer cards such as an RTX 4080 or 4090, community benchmarks put the 20b model at roughly 60–140 tokens a second, depending on context length and drivers. That is slower than Groq, but fine for an 800-token answer. The 20b also writes weaker answers than the 120b Path A uses, and that shows on questions that need several sources combined.

What fully local costs if you rent the GPU

  • RTX 4090 (24 GB), RunPod community, 730 h~$248/mo
  • RTX 4090 (24 GB), RunPod secure cloud, 730 h~$540/mo
  • Firecrawl + SearXNG host, Hetzner CX43€15.99/mo
  • Break-even against Path A at ~1.1–1.7¢~16k–51k queries/mo

If you already own the GPU, the marginal cost is electricity. If you have to rent one around the clock, Path B only matches Path A's price at product volume. Choose it because queries can't leave your network, not to save money.

Path C: Vane with a local model

If what you need is a private Perplexity alternative to use, rather than code to build on, Vane is less work than Path B. It is one container with SearXNG bundled, MIT licensed, and at about 37,000 GitHub stars it is the most widely used project of its kind. The repository moved from Perplexica to ItzCrazyKns/Vane, and the old address redirects.

From the Vane README
$ docker run -d -p 3000:3000 -v vane-data:/home/vane/data \
    --name vane itzcrazykns1337/vane:latest
# Linux, with Ollama on the host: add --add-host=host.docker.internal:host-gateway
# and use http://host.docker.internal:11434 as the Ollama URL in Vane's settings

Its README names Ollama, OpenAI, Anthropic, Gemini and Groq as providers. That gives Path C a cheap middle option: Vane on an €8.49 VPS with a Groq key. Groq is the only metered part, since SearXNG is free, so the bill is the server plus fractions of a cent per query. No other option here is cheaper for something more than one person can use. For an existing SearXNG instance there is a slim image that takes SEARXNG_API_URL. That instance needs the JSON output format and the Wolfram Alpha engine enabled.

The catch with SearXNG

Paths B and C both depend on SearXNG, and SearXNG does not have its own index. It forwards your query to Google, Bing, DuckDuckGo and others from your server's IP address. A single data-centre IP sending steady automated traffic gets blocked, and SearXNG's own defaults assume it. When an engine returns a CAPTCHA, SearXNG suspends that engine for a day. For a Google reCAPTCHA it is seven days, for a Cloudflare challenge fifteen, and for a 429 an hour.

In practice, a self-hosted answer engine can get worse over a week without telling you. Google drops out, then another engine, and answers start coming from whichever engines are still responding. Three things help:

  • Turn on the limiter (server.limiter: true, backed by Valkey) so bots hitting your instance don't get your IP banned upstream. It reads client IPs from X-Forwarded-For, so set up your reverse proxy properly or every user looks like the same person.
  • Enable JSON output. It is off in the default settings.yml, and both Firecrawl and Vane need it.
  • Watch the engine stats. SearXNG reports suspended engines. Check them weekly rather than waiting for users to notice answers getting thinner.

Path A doesn't leave this to you. Keeping search working is Firecrawl's job there, and that is a large part of what the credits pay for.

Before you put it on the internet

Every path above gives you an app that does no access control. Fireplexity in particular was published as a demo. If you deploy it publicly as it is, strangers' queries run on your API keys. The minimum before a public URL:

  • Put it behind a login. For an internal tool, your SSO proxy or Cloudflare Access in front of the whole app is less work than adding auth to the code.
  • Rate-limit the search route per user or IP. Upstash's free Redis tier (500,000 commands a month) is enough for @upstash/ratelimit at small scale.
  • Remove the key pass-through. The route uses a firecrawlApiKey from the request body if one is sent. It is harmless, but it is an input you don't need, and a public deployment shouldn't accept it.
  • Cap spend upstream. On Firecrawl the free plan stops at its limit and pay-as-you-go lets you set a monthly ceiling. Set one. A crawler looping on your search box costs real money in a few hours.
  • Treat scraped pages as untrusted input. Page text goes straight into the prompt. A page can contain instructions aimed at your model, and the bigger risk is that it quietly changes the answer. Remember this before you give the model any tools.

And keep checking answers. Self-hosting changes who runs the system, not whether the citations are right. The failure modes in our guide to checking AI search citations apply to every path here.

So which path

Path A if…

  • You want answers from full, freshly scraped pages, including JavaScript-heavy sites.
  • You are building a product and need predictable per-query cost.
  • You don't want to run search infrastructure. Firecrawl does that.
  • It is just you, and 60–100 free queries a month is enough.

Path C if…

  • You want a private Perplexity alternative to use, not code to modify.
  • You already have a GPU with 16 GB or more, or don't mind a Groq key.
  • Snippet-based answers are good enough for your questions.
  • You will check SearXNG's engine health now and then.

Path B is for the narrower case where the code has to be Fireplexity's and nothing can leave the network: a regulated environment, a corpus you can't send to a third party, or a product whose customers require it. It is the most work and, unless you own the hardware, the most expensive. When it is a hard requirement, though, none of the other options here meet it.

Common questions

Can I self-host a Perplexity alternative for free?

Yes, with limits. Fireplexity on Firecrawl's free plan (1,000 credits a month) and Groq's free tier answers about 60 to 100 queries a month at no cost. Vane with a local model through Ollama has no per-query cost at all, but needs roughly 16 GB of memory for a usable model such as gpt-oss-20b. Its bundled SearXNG can also be throttled by the engines it queries.

How much does it cost to self-host Fireplexity?

About 1.1–1.7¢ a query on Firecrawl's Standard plan billed monthly: 10–16 credits for one search across web, news and images plus its scraped pages, and about a tenth of a cent of Groq inference. You buy credits by the plan, though, so a 3,000-query team spends roughly $115–125 a month including hosting, and a 30,000-query product about $455.

Can Fireplexity run completely offline?

Not as shipped. The Firecrawl address and the Groq provider are hardcoded in one file. Self-hosted Firecrawl with your own SearXNG, plus an OpenAI-compatible provider pointed at Ollama, makes it fully local with two small edits. The self-hosted Firecrawl lacks the cloud's Fire-engine anti-bot layer, so some sites won't scrape.

Can I host Fireplexity on Vercel's free plan?

For personal use, yes. Hobby is restricted to personal, non-commercial use, and its 300-second function limit is more than enough for a streamed answer. A business, even for an internal tool, needs Vercel Pro at $20 per developer seat a month, or a small VPS.

Is self-hosting cheaper than Perplexity Pro?

Rarely, for a team. Five people running 3,000 queries a month through Fireplexity spend a little more than five Pro seats cost, and get a much smaller product. Self-hosting comes out cheaper at product scale, where per-seat pricing doesn't apply. The case for it at any scale is control.

Sources

  1. firecrawl/fireplexity on GitHub — route.ts and content-selection.ts, read at commit 7b7bb97 for the call pattern, trimming and hardcoded endpoints
  2. Firecrawl pricing — plans, monthly free credits, 2 credits per 10 search results, 1 per scraped page (read 8 October 2026); search docs — limit applies per source type
  3. Self-hosting Firecrawl and the firecrawl repository — AGPL-3.0, Fire-engine exclusion, Compose services and limits, search backend order
  4. Groq rate limits and model pages for gpt-oss-120b and gpt-oss-20b — free-tier limits and token prices
  5. Groq model deprecations — Kimi K2 retirement
  6. Vercel Hobby plan, function duration and pricing — non-commercial terms, 300 s / 800 s limits, Pro seat price
  7. Hetzner price adjustment and cost-optimized cloud — CX33 and CX43 prices after 15 June 2026
  8. RunPod pricing — RTX 4090 hourly rates
  9. Ollama: gpt-oss — download sizes and memory requirements; throughput range from community benchmarks
  10. Vane on GitHub — install command, bundled SearXNG, providers, slim image requirements
  11. SearXNG search settings and limiter — JSON format, suspension times, Valkey requirement
  12. Upstash Redis pricing — free tier

Prices are list prices checked on 8 October 2026 and change often; Firecrawl's Standard and Growth prices are billed monthly; annual billing is about a sixth cheaper. Per-query figures are estimates built on the stated assumptions, not measurements. Token counts vary with the question and the pages scraped. This site is an independent community page and is not affiliated with Firecrawl, Perplexity AI or the Vane project.

Start on the free tier.
Move pieces later.

Path A runs locally in about four minutes with two API keys, and every part of it can be swapped out afterwards.