Files
openhuman/gitbooks/features/native-tools/web-scraper.md
T
2026-05-08 23:37:29 -07:00

1.0 KiB

description, icon
description icon
A purpose-built "GET-and-read" tool that returns clean text, not raw HTML. globe

Web Scraper

A purpose-built fetch tool, separate from generic http_request / curl. It exists because the agent doesn't want raw HTML — it wants the article.

What it does

  • Fetches a URL.
  • Strips boilerplate (nav, ads, footer, scripts).
  • Returns clean text the agent can reason over.

Guardrails

  • Caps response at 1 MB — large pages get truncated, not silently dropped.
  • 20-second timeout — slow servers don't stall the conversation.
  • Subject to the same proxy and URL-guard rules as other network tools.

What it's good for

  • Reading articles, blog posts, docs pages, GitHub READMEs without the noise.
  • Following up on a Web Search result.
  • Summarising a single page on demand.

See also