GHSA-X2RJ-828P-HX9M

Vulnerability from github – Published: 2026-08-21 20:56 – Updated: 2026-08-21 20:56
VLAI
Summary
Xinference vulnerable to remote code execution via unsafe `eval()` in Llama3 tool-call parsing
Details

Summary

Xinference used Python's unsafe eval() function when parsing Llama3 tool-call output generated by a large language model. Because the model output can be influenced by attacker-controlled prompts sent to the chat completion API, a remote attacker can craft prompts that cause the model to return a Python expression. Xinference then evaluates that expression on the server while post-processing the tool-call result. In the tested default deployment, authentication was not enabled, so the vulnerability was exploitable by an unauthenticated remote attacker through the /v1/chat/completions endpoint.

Details

Users can interact with deployed models through Xinference's OpenAI-compatible /v1/chat/completions API. The request entry point is implemented in xinference/api/restful_api.py; non-streaming requests call the model instance's chat() method and return the inference result.

When the Transformers backend is used, inference results flow through the batching logic in xinference/model/llm/transformers/core.py. Non-streaming chat results are handled by handle_chat_result_non_streaming(). If the request contains a tools field, Xinference calls _post_process_completion() to parse tool-call output from the model response.

The Llama3 tool-call parser is implemented in xinference/model/llm/tool_parsers/llama3_tool_parser.py. In affected versions, extract_tool_calls() parsed model output with eval():

def extract_tool_calls(
    self, model_output: str
) -> List[Tuple[Optional[str], Optional[str], Optional[Dict[str, Any]]]]:
    try:
        data = eval(model_output, {}, {})
        return [(None, data["name"], data["parameters"])]
    except Exception:
        return [(model_output, None, None)]

The intended behavior was to convert a Python dictionary-like string generated by the model into a dictionary object. However, eval() executes the input as a Python expression, and eval(model_output, {}, {}) is not a security sandbox. If an attacker can influence the model output through prompt injection or direct chat input, the attacker can cause the model to return an expression such as:

__import__('os').system('touch /tmp/hacked')

When the expression reaches eval(), it is executed in the Xinference server process context. The harmless touch /tmp/hacked command can be replaced with other payloads, such as a reverse shell, malware download, sensitive file read, or lateral-movement payload.

Score

Severity: Critical

CVSS v3.1: 10.0

Vector: CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:C/C:H/I:H/A:H

Rationale:

  • AV:N: the vulnerable API is remotely reachable over the network;
  • AC:L: exploitation only requires a crafted chat-completion request and tool-call parameter;
  • PR:N: the tested default configuration did not require authentication;
  • UI:N: no user interaction is required;
  • S:C: command execution can affect resources beyond the Xinference application boundary;
  • C:H/I:H/A:H: remote code execution can fully compromise confidentiality, integrity, and availability.

Credit

This vulnerability was discovered by:

  • XlabAI Team of Tencent Xuanwu Lab (xlabai@tencent.com)
  • Atuin Automated Vulnerability Discovery Engine
  • Guannan Wang (wgnbuaa@gmail.com), Zhanpeng Liu (pkugenuine@gmail.com), Guancheng Li (lgcpku@gmail.com)
Show details on source website

{
  "affected": [
    {
      "database_specific": {
        "last_known_affected_version_range": "\u003c= 2.5.0"
      },
      "package": {
        "ecosystem": "PyPI",
        "name": "xinference"
      },
      "ranges": [
        {
          "events": [
            {
              "introduced": "0"
            },
            {
              "fixed": "2.7.0"
            }
          ],
          "type": "ECOSYSTEM"
        }
      ]
    }
  ],
  "aliases": [
    "CVE-2026-61539"
  ],
  "database_specific": {
    "cwe_ids": [
      "CWE-95"
    ],
    "github_reviewed": true,
    "github_reviewed_at": "2026-08-21T20:56:37Z",
    "nvd_published_at": null,
    "severity": "CRITICAL"
  },
  "details": "### Summary\n\nXinference used Python\u0027s unsafe `eval()` function when parsing Llama3 tool-call output generated by a large language model. Because the model output can be influenced by attacker-controlled prompts sent to the chat completion API, a remote attacker can craft prompts that cause the model to return a Python expression. Xinference then evaluates that expression on the server while post-processing the tool-call result. In the tested default deployment, authentication was not enabled, so the vulnerability was exploitable by an unauthenticated remote attacker through the `/v1/chat/completions` endpoint.\n\n### Details\n\nUsers can interact with deployed models through Xinference\u0027s OpenAI-compatible `/v1/chat/completions` API. The request entry point is implemented in `xinference/api/restful_api.py`; non-streaming requests call the model instance\u0027s `chat()` method and return the inference result.\n\nWhen the Transformers backend is used, inference results flow through the batching logic in `xinference/model/llm/transformers/core.py`. Non-streaming chat results are handled by `handle_chat_result_non_streaming()`. If the request contains a `tools` field, Xinference calls `_post_process_completion()` to parse tool-call output from the model response.\n\nThe Llama3 tool-call parser is implemented in `xinference/model/llm/tool_parsers/llama3_tool_parser.py`. In affected versions, `extract_tool_calls()` parsed model output with `eval()`:\n\n```python\ndef extract_tool_calls(\n    self, model_output: str\n) -\u003e List[Tuple[Optional[str], Optional[str], Optional[Dict[str, Any]]]]:\n    try:\n        data = eval(model_output, {}, {})\n        return [(None, data[\"name\"], data[\"parameters\"])]\n    except Exception:\n        return [(model_output, None, None)]\n```\n\nThe intended behavior was to convert a Python dictionary-like string generated by the model into a dictionary object. However, `eval()` executes the input as a Python expression, and `eval(model_output, {}, {})` is not a security sandbox. If an attacker can influence the model output through prompt injection or direct chat input, the attacker can cause the model to return an expression such as:\n\n```python\n__import__(\u0027os\u0027).system(\u0027touch /tmp/hacked\u0027)\n```\n\nWhen the expression reaches `eval()`, it is executed in the Xinference server process context. The harmless `touch /tmp/hacked` command can be replaced with other payloads, such as a reverse shell, malware download, sensitive file read, or lateral-movement payload.\n\n### Score\n\nSeverity: Critical\n\nCVSS v3.1: 10.0\n\nVector: `CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:C/C:H/I:H/A:H`\n\nRationale:\n\n- AV:N: the vulnerable API is remotely reachable over the network;\n- AC:L: exploitation only requires a crafted chat-completion request and tool-call parameter;\n- PR:N: the tested default configuration did not require authentication;\n- UI:N: no user interaction is required;\n- S:C: command execution can affect resources beyond the Xinference application boundary;\n- C:H/I:H/A:H: remote code execution can fully compromise confidentiality, integrity, and availability.\n\n### Credit\n\nThis vulnerability was discovered by:\n\n- XlabAI Team of Tencent Xuanwu Lab (xlabai@tencent.com)\n- Atuin Automated Vulnerability Discovery Engine\n- Guannan Wang (wgnbuaa@gmail.com), Zhanpeng Liu (pkugenuine@gmail.com), Guancheng Li (lgcpku@gmail.com)",
  "id": "GHSA-x2rj-828p-hx9m",
  "modified": "2026-08-21T20:56:37Z",
  "published": "2026-08-21T20:56:37Z",
  "references": [
    {
      "type": "WEB",
      "url": "https://github.com/xorbitsai/inference/security/advisories/GHSA-x2rj-828p-hx9m"
    },
    {
      "type": "WEB",
      "url": "https://github.com/xorbitsai/inference/pull/4786"
    },
    {
      "type": "WEB",
      "url": "https://github.com/xorbitsai/inference/commit/1b3d220f342ce68d34cec4586d9409d457dadc42"
    },
    {
      "type": "PACKAGE",
      "url": "https://github.com/xorbitsai/inference"
    },
    {
      "type": "WEB",
      "url": "https://github.com/xorbitsai/inference/releases/tag/v2.7.0"
    }
  ],
  "schema_version": "1.4.0",
  "severity": [
    {
      "score": "CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:C/C:H/I:H/A:H",
      "type": "CVSS_V3"
    }
  ],
  "summary": "Xinference vulnerable to remote code execution via unsafe `eval()` in Llama3 tool-call parsing"
}



Log in or create an account to share your comment.




Tags
Taxonomy of the tags.


Loading…

Loading…

Loading…

Forecast uses a logistic model when the trend is rising, or an exponential decay model when the trend is falling. Fitted via linearized least squares.

Sightings

Author Source Type Date Other

Nomenclature

  • Seen: The vulnerability was mentioned, discussed, or observed by the user.
  • Confirmed: The vulnerability has been validated from an analyst's perspective.
  • Published Proof of Concept: A public proof of concept is available for this vulnerability.
  • Exploited: The vulnerability was observed as exploited by the user who reported the sighting.
  • Patched: The vulnerability was observed as successfully patched by the user who reported the sighting.
  • Not exploited: The vulnerability was not observed as exploited by the user who reported the sighting.
  • Not confirmed: The user expressed doubt about the validity of the vulnerability.
  • Not patched: The vulnerability was not observed as successfully patched by the user who reported the sighting.

Loading…

Detection rules are retrieved from Rulezet.

Loading…

Loading…

Loading…