Skip to content

Latest commit

 

History

5 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Perplexity Scraper

Scrapeless Perplexity Scraper - collect Perplexity answers with one API call

Try Scrapeless Blog License

Send a prompt to Perplexity and get the answer back as structured JSON — the Markdown response, the ranked web sources behind it, and Perplexity's own follow-up questions — through the Scrapeless LLM Chat Scraper API.

Perplexity is a search engine wearing a chatbot's clothes: every answer is grounded in ranked web results, and those results are returned to you as data. That makes this the most directly SEO-shaped of the LLM actors — web_results is effectively an AI-mediated SERP for your query.

Quick start

curl 'https://api.scrapeless.com/api/v2/scraper/execute' \
  --header 'Content-Type: application/json' \
  --header 'x-api-token: YOUR_API_TOKEN' \
  --data '{
    "actor": "scraper.perplexity",
    "input": {
      "prompt": "Best supplements for better sleep",
      "country": "US"
    }
  }'

To receive the result asynchronously, add a webhook object:

"webhook": { "url": "https://www.your-webhook.com" }

Request parameters

Parameter (input.*) Type Required Description
prompt string Yes The prompt to send to Perplexity.
country string Yes Country / region code (e.g. US, JP, DE) the prompt is sent from.

There is no web_search flag on this actor, and adding one changes nothing. We ran the same prompt with web_search: true and with the key omitted entirely: both returned 10 web_results and the same task_result shape. Perplexity always searches — that is what it is. The key is accepted and silently dropped, so a 200 response is not evidence that a parameter did anything. web_search is real on scraper.chatgpt; it is not real here.

Response

{
  "status": "success",
  "task_id": "57a822f6-ee56-4d5f-b952-34fc3ac74b9f",
  "task_result": {
    "prompt": "Best supplements for better sleep",
    "result_text": "...Markdown answer with ([domain][1]) markers and a link footer...",
    "web_results": [
      {
        "name": "Supplementing your sleep",
        "url": "https://www.health.harvard.edu/healthy-aging-and-longevity/supplementing-your-sleep",
        "snippet": "..."
      }
    ],
    "related_prompt": [
      "Which sleep aids have the strongest evidence in adults"
    ],
    "media_items": []
  }
}

A complete, unedited response from a real run is committed at results/perplexity-sample.json — see results/TRIMMED.md.

task_result fields

Field Type Description
prompt string The prompt you submitted, echoed back.
result_text string Markdown answer. Carries inline ([domain][n]) markers and a trailing reference block mapping [n] to URL and title.
web_results array Ranked sources behind the answer. This is the field most people actually want.
web_results[].name string Page title. Note it is name, not title — the other LLM actors use title.
web_results[].url string Page URL. Clean, no text fragments.
web_results[].snippet string Extracted snippet.
related_prompt array Perplexity's suggested follow-up questions.
media_items array Images/videos surfaced with the answer. Frequently empty.

Field notes from real runs

  • web_results[].name, not .title. Every other LLM Chat Scraper actor calls this key title. If you are writing one collector across engines, normalise here or you will silently write nulls. This is the single most common porting bug on this actor.
  • result_text embeds its own bibliography. Beyond the inline ([domain][n]) markers, the answer ends with a [1]: https://… "Title" block. If you render the Markdown as-is, that block shows up as visible text. Strip it, or parse it — it is a second, independent copy of the source list.
  • 10 web_results is the shape we consistently saw, with or without extra input keys. Do not treat it as a guaranteed count.
  • media_items is usually []. Present but empty on our runs. Guard it anyway.
  • A cold call takes roughly 15 seconds. The examples here use a 180s client timeout.

Code examples

Language File Run
Python example.py pip install requests && python example.py
Node.js example.js node example.js (Node 18+)
Go example.go go run example.go
Java Example.java java Example.java (Java 11+)
PHP example.php php example.php
export SCRAPELESS_API_TOKEN="your_api_token"

Copy .env.example to .env if you prefer a file. Never commit the real token.

Use cases

AI-mediated SERP tracking. web_results is a ranked source list for a natural-language query. Run your keyword set on a schedule and you have rank tracking for AI search.

Citation share of voice. Count how often your domain appears in web_results across a prompt set, against competitors.

Content gap analysis. The pages Perplexity cites for a question are the pages it considers authoritative. If yours is not among them, that is the brief.

Follow-up mining. related_prompt is Perplexity's own model of what a user asks next — a clean source of long-tail topics.

Cross-engine benchmarking. Same envelope and auth as chatgpt-scraper and gemini-scraper, so one collector covers all three.

FAQ

What is the Perplexity Scraper? A Scrapeless LLM Chat Scraper actor (scraper.perplexity) that submits a prompt to Perplexity and returns the answer plus its ranked sources as JSON.

Do I need a browser or proxy pool? No. One HTTP POST on your side.

Can I force web search on or off? No — Perplexity always searches, and there is no flag that changes that. See the note under Request parameters.

Can I get results asynchronously? Yes — add a webhook object with your callback URL.

Is this legal? You are responsible for your own use. Check applicable law, platform terms, and your organisation's data policy, and do not collect private or personal data.

Verification

Everything documented above was captured live on 2026-09-01 against POST https://api.scrapeless.com/api/v2/scraper/execute.

Check Result
scraper.perplexity live call HTTP 200 in 15.4s, status: success
task_result keys observed media_items, prompt, related_prompt, result_text, web_results
web_results returned 10 entries, keyed name / url / snippet
web_search control test omitted vs true → identical shape, 10 results both times; flag is a no-op
examples/example.py run live, returned a real answer + sources
examples/example.js run live, returned a real answer + sources
examples/example.php syntax-checked (php -l)
examples/Example.java compiled (javac)
examples/example.go not run — no Go toolchain on the verification box

Learn more

Contact

  • Discord
  • Telegram
  • For repo issues or improvements, open an issue or pull request.

About

Collect Perplexity answers, ranked web sources, and follow-up questions through the Scrapeless LLM Chat Scraper API for AI search monitoring and source analysis.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors