mirror of
https://github.com/tinyhumansai/openhuman.git
synced 2026-07-27 21:08:00 +00:00
1019 B
1019 B
description, icon
| description | icon |
|---|---|
| A purpose-built "GET-and-read" tool that returns clean text, not raw HTML. | globe |
Web Scraper
A purpose-built fetch tool, separate from generic http_request / curl. It exists because the agent doesn't want raw HTML - it wants the article.
What it does
- Fetches a URL.
- Strips boilerplate (nav, ads, footer, scripts).
- Returns clean text the agent can reason over.
Guardrails
- Caps response at 1 MB - large pages get truncated, not silently dropped.
- 20-second timeout - slow servers don't stall the conversation.
- Subject to the same proxy and URL-guard rules as other network tools.
What it's good for
- Reading articles, blog posts, docs pages, GitHub READMEs without the noise.
- Following up on a Web Search result.
- Summarising a single page on demand.
See also
- Web Search - find URLs to feed into the scraper.
- Smart Token Compression - what trims long pages before they hit the model.