<?xml version='1.0' encoding='UTF-8'?>
<?xml-stylesheet href="/static/style.xsl" type="text/xsl"?>
<rss xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/" version="2.0">
  <channel>
    <title>Most recent entries from all</title>
    <link>https://db.gcve.eu</link>
    <description>Contains only the most 10 recent entries.</description>
    <docs>http://www.rssboard.org/rss-specification</docs>
    <generator>python-feedgen</generator>
    <language>en</language>
    <lastBuildDate>Fri, 09 Oct 2026 13:01:42 +0000</lastBuildDate>
    <item>
      <title>fkie_cve-2026-105760</title>
      <link>https://db.gcve.eu/vuln/fkie_cve-2026-105760</link>
      <description>&lt;p&gt;vLLM is an inference and serving engine for large language models. Prior to 0.30.0, a caller can use the request-level media_io_kwargs field to select the GLMGA video backend and supply large values for the fps and max_frames options without a strict work ceiling. GLMGA constructs and deduplicates an attacker-sized pre-decode frame-index list, allowing a compact request and tiny valid video to consume disproportionate CPU time and memory in the shared media-loading executor. This issue is fixed in version 0.30.0.&lt;/p&gt;</description>
      <content:encoded>&lt;p&gt;vLLM is an inference and serving engine for large language models. Prior to 0.30.0, a caller can use the request-level media_io_kwargs field to select the GLMGA video backend and supply large values for the fps and max_frames options without a strict work ceiling. GLMGA constructs and deduplicates an attacker-sized pre-decode frame-index list, allowing a compact request and tiny valid video to consume disproportionate CPU time and memory in the shared media-loading executor. This issue is fixed in version 0.30.0.&lt;/p&gt;</content:encoded>
      <guid isPermaLink="false">https://db.gcve.eu/vuln/fkie_cve-2026-105760</guid>
    </item>
    <item>
      <title>GHSA-58v5-2m8f-94pr — vLLM: GLMGA video sampling permits request-driven CPU and memory exhaustion</title>
      <link>https://db.gcve.eu/vuln/ghsa-58v5-2m8f-94pr</link>
      <description>&lt;p&gt;&lt;strong&gt;Affected:&lt;/strong&gt; PyPI: vllm&lt;/p&gt;
&lt;p&gt;## Summary&lt;/p&gt;
&lt;p&gt;The OpenAI-compatible chat endpoint accepts request-level video loader options through `media_io_kwargs`. A caller can select the GLMGA sampler and provide large `fps` and `max_frames` values. GLMGA first constructs an attacker-sized Python index list and then deduplicates it, even when the supplied video contains only a few frames. This work occurs in the shared media-loading executor before frame reads, allowing a tiny valid video and compact JSON options to consume disproportionate CPU time and memory and delay unrelated media requests.&lt;/p&gt;
&lt;p&gt;The demonstrated impact is partial denial of service. No memory corruption, data disclosure, or code execution is claimed.&lt;/p&gt;
&lt;p&gt;## Affected configuration&lt;/p&gt;
&lt;p&gt;The server must expose chat completions for a video-capable model and accept request-level `media_io_kwargs`. The model does not need to use GLMGA by default because the request value overrides the model&amp;#39;s loader mapping. The proof of concept uses an inline two-frame video and explicitly selects the OpenCV decoder, so no external media host or GPU decoder is required.&lt;/p&gt;
&lt;p&gt;The request is reachable without vLLM credentials when no API key is configured. With API-key authentication enabled, any caller holding a valid key can reach the same path.&lt;/p&gt;
&lt;p&gt;## Attack surface&lt;/p&gt;
&lt;p&gt;A remote caller submits a valid chat-completion request with a small video and the following request-level options:&lt;/p&gt;
&lt;p&gt;```json
{
  &amp;#34;media_io_kwargs&amp;#34;: {
    &amp;#34;video&amp;#34;: {
      &amp;#34;video_backend&amp;#34;: &amp;#34;glmga&amp;#34;,
      &amp;#34;backend&amp;#34;: &amp;#34;opencv&amp;#34;,…&lt;/p&gt;</description>
      <content:encoded>&lt;p&gt;&lt;strong&gt;Affected:&lt;/strong&gt; PyPI: vllm&lt;/p&gt;
&lt;p&gt;## Summary&lt;/p&gt;
&lt;p&gt;The OpenAI-compatible chat endpoint accepts request-level video loader options through `media_io_kwargs`. A caller can select the GLMGA sampler and provide large `fps` and `max_frames` values. GLMGA first constructs an attacker-sized Python index list and then deduplicates it, even when the supplied video contains only a few frames. This work occurs in the shared media-loading executor before frame reads, allowing a tiny valid video and compact JSON options to consume disproportionate CPU time and memory and delay unrelated media requests.&lt;/p&gt;
&lt;p&gt;The demonstrated impact is partial denial of service. No memory corruption, data disclosure, or code execution is claimed.&lt;/p&gt;
&lt;p&gt;## Affected configuration&lt;/p&gt;
&lt;p&gt;The server must expose chat completions for a video-capable model and accept request-level `media_io_kwargs`. The model does not need to use GLMGA by default because the request value overrides the model&amp;#39;s loader mapping. The proof of concept uses an inline two-frame video and explicitly selects the OpenCV decoder, so no external media host or GPU decoder is required.&lt;/p&gt;
&lt;p&gt;The request is reachable without vLLM credentials when no API key is configured. With API-key authentication enabled, any caller holding a valid key can reach the same path.&lt;/p&gt;
&lt;p&gt;## Attack surface&lt;/p&gt;
&lt;p&gt;A remote caller submits a valid chat-completion request with a small video and the following request-level options:&lt;/p&gt;
&lt;p&gt;```json
{
  &amp;#34;media_io_kwargs&amp;#34;: {
    &amp;#34;video&amp;#34;: {
      &amp;#34;video_backend&amp;#34;: &amp;#34;glmga&amp;#34;,
      &amp;#34;backend&amp;#34;: &amp;#34;opencv&amp;#34;,…&lt;/p&gt;</content:encoded>
      <guid isPermaLink="false">https://db.gcve.eu/vuln/ghsa-58v5-2m8f-94pr</guid>
    </item>
  </channel>
</rss>
