Decision Models vs. CVE Data

Ollama’s new decision endpoint, Cloudflare’s Clef models on a MacBook, and two questions asked of all 27,489 CVEs published in the last 60 days: does the description say why the bug matters, and is it clear?

I read the announcement and had a weekend project before I finished it, because the question I care about most in the CVE Program has always been a classification problem: can a defender read this CVE and know why it should care?

The announcement was Ollama 0.35.1 adding Cloudflare’s Clef and Clef Flash to a new kind of model it started supporting in 0.35. They don’t write text. You hand them some state and a set of typed questions, and they hand back choices, yes-or-no probabilities and scores, with nothing to parse. That is a classifier that runs on my laptop, which is exactly the shape of the problem.

I already track how CNAs fill in their records at cnascorecard.org, but that only covers the structured fields a parser can read. The part I have never been able to check at scale is the prose, the one field every scanner shows an operator. So I asked Clef two things about every CVE published between August 4 and October 3, 2026. Does the description establish a security impact? And how clear is it for a defender?

Full disclosure: I serve on the CVE Program’s Quality, Automation and Consumer working groups. Nothing here is a CVE Program position.

TL;DR

Clef Flash, the 9B model, answered both questions for every new CVE at a median of under a second each on a laptop, with no API key and no bill. Outside the Linux kernel, 989 of 23,736 descriptions (4.2%, about one in 24) never say what an attacker could do. The kernel is its own case: its published policy leaves impact to the user, so 3,661 of its 3,753 records don’t state one either, which takes the all-CVE figure to 4,650 of 27,489 (16.9%). Flash on its own is too strict, so Clef 27B re-reads everything it flags as “no impact” outside the kernel, and 27B overturned 1,198 of 2,187 (54.8%). Against 129 published CVEs I labeled by hand, that two-step answer caught 15 of the 16 I said state no impact and never called one of my yeses a no.

Key Statistics at a Glance

Metric Value
CVEs published, August 4 to October 3, 2026 27,489, from 283 CNAs
No security impact stated, outside the Linux kernel 989 of 23,736 (4.2%)
No security impact stated, Linux kernel CNA (by policy) 3,661 of 3,753 (97.5%)
No security impact stated, all CVEs 4,650 of 27,489 (16.9%)
Large CNAs with zero such descriptions Microsoft, Oracle, Chrome, Wordfence, WPScan (5,666 records)
Descriptions rated Good or better for clarity (Flash, which runs generous) 69.7%
Clef Flash vs Clef 27B, same impact answer (published records, development sample) 402 of 428 (93.9%)
Impact answer vs my hand labels (Flash, then 27B) 128 of 129 (99.2%)
Clarity vs my hand labels (Flash), correlation 0.71
Clef Flash median time per description 0.84 seconds
Hardware A MacBook Pro I already own (M5 Pro, 64 GB)
API bill None

A Model That Only Answers Multiple Choice

Clef is a 27B model fine-tuned from Qwen 3.8, and Clef Flash is a 9B model fine-tuned from Qwen 3.5. Both are Apache 2.0, with weights on Hugging Face (Clef, Clef Flash), and both run locally from the Ollama library (clef, clef-flash; ollama pull clef-flash is about 10 GB). Cloudflare also hosts them on Workers AI, but everything here ran on my laptop. They are called through /v1/systemone, the decision endpoint Ollama added in 0.35, with a block of state and up to 64 named questions. Each question is one of three types:

Type What comes back
choice One of 2 to 26 options, with a probability for each
noul A single probability that the answer is true
score A position on an ordered scale of 2 to 26 levels, with a probability for each level

Here is one real request from the run, trimmed. The state is the CVE’s title and description, and the questions are written in plain English with a description for each answer:

{
  "model": "clef-flash",
  "state": {
    "title": "Arbitrary File Read in Google Agent Development Kit (ADK)",
    "description": "A Path Traversal vulnerability in the builder endpoint in Google Cloud Agent Development Kit (ADK) versions 1.9.0 through 1.21.0 on Python allows an unauthenticated remote attacker to read arbitrary files using a crafted file_path query parameter."
  },
  "questions": {
    "security_impact_stated": {
      "type": "noul",
      "instructions": "Does the description establish a security impact, that is, explain or clearly imply how an attacker could gain access, capabilities, or information they are not authorized for, or deny service to others? ...",
      "criteria": {
        "false": "No security impact is established: it describes a bug, a code change, or a fix ...",
        "true": "A security impact is stated or clearly implied ..."
      }
    },
    "impact_basis": {
      "type": "choice",
      "instructions": "How does the description support (or fail to support) a security impact?",
      "criteria": {
        "stated_explicitly": "...",
        "implied_by_weakness": "...",
        "bug_fix_only": "...",
        "not_security_relevant": "...",
        "insufficient_information": "..."
      }
    },
    "desc_clarity": {
      "type": "score",
      "instructions": "Rate how clear, specific, and useful the description is for a defender ...",
      "criteria": [
        "Unusable ...",
        "Poor ...",
        "Adequate ...",
        "Good ...",
        "Excellent ..."
      ]
    }
  }
}

And what Flash returned for CVE-2026-79707 (from my development run, where these three questions rode along in a bigger request):

{
  "security_impact_stated": { "type": "noul", "noul": 0.882 },
  "impact_basis": {
    "type": "choice",
    "choice": "stated_explicitly",
    "confidence": 0.63,
    "probabilities": {
      "stated_explicitly": 0.845,
      "implied_by_weakness": 0.098,
      "bug_fix_only": 0.029,
      "insufficient_information": 0.015,
      "not_security_relevant": 0.013
    }
  },
  "desc_clarity": {
    "type": "score",
    "score": 3.35,
    "probabilities": {
      "0": 0.010, "1": 0.019, "2": 0.042,
      "3": 0.465, "4": 0.464
    }
  }
}

That is the whole appeal. No regex over a chat reply, no “please respond in JSON”, no answer that wanders outside the options. In every request I sent, the output token count came back zero, because the model scores the options rather than writing anything.

Worth knowing before you build on it: a choice question tops out at 26 options, and cost scales with the number of questions. On this laptop the 27B model takes about 0.6 seconds per question once a request carries eight or more, and 2.0 seconds even for a request with only two. Sending requests in parallel bought me nothing, because on my setup Ollama answers decision requests one at a time. Ask the questions that earn their time.

Two Questions and Nothing Else

Does the description establish a security impact? Not “is this a vulnerability”, which the text often can’t settle. Just: does the description say, or clearly imply, what an attacker could do? A follow-up choice asks how: it says outright, it names a vulnerability class like SQL injection whose impact is understood, it describes only the bug or its fix, it has no security angle, or there isn’t enough to tell.

How clear is the description for a defender? A score from 0 (Unusable) to 4 (Excellent), with the levels written out: Adequate means it names the product and the problem but leaves out versions, the attacker or the impact; Excellent has all of them.

Clef only ever sees the title and description. It never sees the CVSS vector, the CWE, the affected-product data or the references. That is deliberate.

Outside the Kernel, One in 24 Never Says Why It Matters

Outside the Linux kernel, 989 of 23,736 descriptions (4.2%, about one in 24) don’t establish a security impact, and they come from 87 of the other 282 CNAs. Count the kernel and it is 4,650 of 27,489 (16.9%), but the kernel’s 3,661 are a policy, not an oversight, and they get their own paragraph below.

Where the 4,650 descriptions with no stated impact come from: Mozilla 144, MITRE 115, VulDB 107, GitHub 83, Apache 74, Cisco 53, VulnCheck 35, 378 from 80 other CNAs, and 3,661 (78.7%) from the Linux kernel CNA by policy

A selection, including three small CNAs (30 to 43 records) whose percentages move a lot on a few records:

CNA Descriptions No security impact stated
Linux 3,753 3,661 (97.5%)
Tanium 30 28 (93.3%)
Qualcomm 31 22 (71.0%)
Mozilla 250 144 (57.6%)
VMware 99 29 (29.3%)
Drupal 43 12 (27.9%)
Cisco 254 53 (20.9%)
Apache 454 74 (16.3%)
MITRE (CNA of Last Resort) 1,011 115 (11.4%)
Apple 295 30 (10.2%)
VulDB 1,587 107 (6.7%)
GitHub 2,836 83 (2.9%)
Microsoft, Oracle, Chrome, Wordfence, WPScan 5,666 0

The Linux kernel CNA tops that table on purpose. Its published policy is to assign a CVE to any bug fix it identifies and to leave applicability to the user, so its records are the fix commit message under “In the Linux kernel, the following vulnerability has been resolved”. CVE-2026-90222 explains in careful detail how a request buffer can be freed while the NFC driver is still sending it, and never says what an attacker could do with that. Clef 27B gives it 0.042 for a stated impact and files it under “bug or fix only” at 0.921. That policy is also why the kernel accounts for 78.7% of every description in the window that states no impact.

Qualcomm’s flagged records read like “Memory corruption while processing rear sensor IOCTL calls.” Mozilla’s flagged ones are longer but built the same way: “Use-after-free in the DOM: Streams component. This vulnerability was fixed in Firefox 156…” A bug class and a component, with no attacker and no consequence. Mozilla’s severity lives in its linked advisories, which Clef never sees.

Tanium’s row is the one I don’t trust. Its records read “Tanium addressed a SQL injection vulnerability in Asset.” A named SQL injection counts as an impact everywhere else: of the 1,968 descriptions in the window that name SQL injection or cross-site scripting, five are scored “no impact”, and four of those five are Tanium’s. Even “Tanium addressed an information disclosure vulnerability in Discover” names its consequence. My guess is that “Tanium addressed…” reads to the model as a patch note. None of my hand labels is a Tanium record, so treat that 93.3% as the model’s quirk, not Tanium’s.

All 12 of Drupal’s are “Unsupported” advisories, which the Drupal Security Team publishes when a module has a known security issue its maintainer hasn’t fixed, and the advice is to uninstall it. CVE-2026-76782 reads, in full: “Vulnerability in Drupal Screenshot. This issue affects Screenshot versions: *.*.” Even “uninstall it” would say more than that.

Clarity Follows the Template

The corpus averages 2.7 on the 0 to 4 scale, and 69.7% of descriptions rate Good or better. That is Flash’s line for Good, and Flash is generous: on my 129 hand labels it called 55% Good or better and I called 34%, so read 69.7% as high.

Average Clef Flash clarity for the 19 CNAs with 250 or more CVEs, from Oracle and Wordfence at 3.26 down to the Linux kernel at 1.88

The CNAs at the top write to a template that names the product, the versions, the attacker and the outcome, and it shows. Clarity also catches things no rule check would. 53 of Cisco’s 254 CVEs open with “As part of Cisco’s ongoing commitment to proactive security and product quality”, then describe a hardening release that fixes “multiple” internally found issues. Those 53 are exactly the Cisco records with no stated impact; the other 201 all have one.

Microsoft is the other case worth a look. Every one of its 1,456 descriptions states an impact, yet it averages 2.45 for clarity. Its descriptions follow one short formula (weakness, component, attacker, outcome), none of the 1,456 carries a version number, and 52 of them are the same sentence. The versions are in the structured affected data, which the rules allow; Clef just can’t see them.

The Small Model Needs a Second Look

My plan was to run Flash on everything and use 27B as a spot check. Apple killed that plan. Flash on its own said 126 of Apple’s 295 descriptions (42.7%) stated no security impact. Apple’s records end with a line like “An app may be able to bypass sandbox restrictions.” That is a stated impact. Flash gave CVE-2026-84577 a 0.328 and filed it under “bug or fix only”. It was reading Apple’s terse house style as an absence.

So the impact answer is a cascade. Flash reads everything. Clef 27B, about three and a half times slower on the same two questions (2.0 seconds against 0.59), re-reads every description outside the Linux kernel that Flash flagged as stating no impact. For the kernel, it re-reads a sample of 333 as a check. Where 27B has re-read a record, its answer wins.

Flash said “no impact stated” Re-read by 27B 27B disagreed
Outside the Linux kernel 2,187 1,198 (54.8%)
Linux kernel 333 19 (5.7%)

After the re-read, Apple drops from 42.7% to 10.2%, Mozilla from 70.8% to 57.6%, and the corpus from 5,867 “no impact” calls to 4,650. 27B gives CVE-2026-84577 a 0.873.

It doesn’t fix everything. All 30 Apple records still counted as “no impact” name a consequence, and five of them are “An app may be able to cause unexpected system termination”, which 27B scores 0.448 while filing it under “stated explicitly”. The two answers contradict each other. That’s the edge of the instrument, and those 30 are in the 4,650.

The cascade also has a blind spot: 27B only ever re-reads Flash’s “no” calls. Outside the kernel, where it happened to read Flash’s “yes” calls, it flipped 1 of 341. Scaled to all 21,549, that is about 60 records wrongly counted as stating an impact, so 4.2% may be a quarter point low. In the kernel sample it flipped 19 of 333 “no” calls, which would be about 190 kernel records wrongly counted as stating none, so the all-CVE 16.9% is, if anything, half a point high.

Dot plot of each CNA's no-impact share, Clef Flash alone versus after the 27B re-read; Apple falls from 42.7% to 10.2%

Clarity stays on Flash alone, so every CNA is scored by the same model. On the 529 development records both models scored, their clarity scores correlate at 0.87.

Checked Against My Own Labels

Agreement between two models only tells you they agree. So I built a small local web app and labeled 150 CVEs by hand: up to three per CNA from the last 30 days, plus 21 rejected records (everything below uses the 129 published ones). I answered the same two questions without seeing either model’s answer until I had committed to mine, and 149 of the 150 labels came after the question wording was frozen.

On the 129 published records Agrees with me Said “no impact” when I said yes Said “impact” when I said no
Flash alone 121 (93.8%) 7 1
Clef 27B 128 (99.2%) 0 1
Flash, then 27B on what it flags 128 (99.2%) 0 1

Every one of Flash’s seven misses was the too-strict kind, and the 27B re-read fixed all seven, which is the case for the cascade in one row. Worth being precise about what 128 of 129 means, though: I said “no impact” on only 16 of the 129, so a model that always said “impact” would score 113 (87.6%). The real result is that the cascade caught 15 of my 16 and never called one of my yeses a no. I also tried a stricter version of the impact answer, which put the headline at 15.4%. It agreed with my labels on 124 instead of 128, so the plain answer stayed.

The one miss is CVE-2026-77249, where a caller-controlled Jira URL “can redirect that unhooked request to an internal address, bypassing the redirect checks…” I read that as not saying what an attacker gets. Both models lean the other way, and the record’s own title says “redirect-based SSRF”. They have a point, which makes that miss more likely mine than theirs.

On clarity, Flash tracks me with a correlation of 0.71 and lands on exactly my level, after rounding, 72.1% of the time; always guessing 2 would get 62.8%. I never used the top or bottom of the scale, so “within one level” would mean nothing here. Flash runs about a quarter of a level generous: it averaged 2.37 on the ones I called 2 and 3.05 on my 3s. 27B was about as close as Flash.

As for the sample: capping it at three records per CNA means it holds far fewer Linux kernel fix logs than the corpus does, so it tests the hard cases outside the kernel more than the easy ones inside it. The labels also come from the same 30-day slice I used to develop the questions, so this is an in-sample check. And 150 records labeled by one person on a Sunday morning is a check, not a benchmark.

The lesson I would hand to anyone building on Clef: a small model is a fine screen and a poor judge. Running 27B on only the 2,520 records it needed to see was about a tenth of the work of running it on everything.

What It Cost

Flash takes a median of 0.84 seconds per description (14,870 timed requests, 90th percentile 0.98 seconds), which is about 6.4 hours of work for all 27,489 on an M5 Pro MacBook Pro with 64 GB. 2,434 of the 2,520 Clef 27B re-reads happened during the mess described next, at a median of 4.4 seconds. Measured on its own, 27B answers the same two questions in a median of 2.0 seconds (661 earlier requests), which is about an hour and a half for all 2,520. There was no API bill, no key to manage and no data leaving the laptop.

The wall clock was longer than that, and it was my fault. For about nine hours I had two Claude Code sessions both babysitting the run, each chaining the same chunks against the same queue in the same order. So 12,003 Flash answers and 2,438 27B answers were each asked twice, a few seconds apart, and each copy ran at half speed. The answers were identical, so nothing in the results changed. ps would have told me in a second.

Every CNA, Check by Check

The two answers sit in a per-CNA table alongside the parts a parser handles without any model: 23 checks, 16 of them mapped to CNA Operational Rules 4.1.0 and the rest data-format checks, run on every record in three seconds. Among them, 5,097 records (18.5%) leave the CNA’s problem-type field empty, including every record from the Linux kernel CNA and from Mozilla. Rule 5.1.7 makes identifying the type a MUST, but a CNA can meet it in the description, and Mozilla’s usually do. And 1,947 records (7.1%) share their description word for word with another CVE, which rule 5.1.1 discourages but the rest of the record can make up for. Adobe has 72 Experience Manager CVEs sharing one sentence. Microsoft has 52 sharing “Heap-based buffer overflow in Windows Biometric Service allows an authorized attacker to elevate privileges locally.” There is no combined score. Each rule check and each Clef answer is its own column, so you can sort by the one you care about. The per-CNA page is here.

What the Instrument Can’t See

Clef reads the title and description and nothing else, so a terse record that links a thorough advisory scores as terse. That is the right test for “is this record usable by itself” and the wrong test for “did the CNA do its homework”.

The check on those numbers is 150 hand labels from one labeler, and both questions are judgment calls. Reasonable people will put the line for “states an impact” in slightly different places, as CVE-2026-77249 shows.

If You Run a CNA

Put the attacker and the consequence in the description: who can do it, from where, and what they get. “An app may be able to bypass sandbox restrictions” is enough. “Memory corruption while processing rear sensor IOCTL calls” is not. If 52 of your CVEs share one sentence, the sentence isn’t doing its job.

If You Consume CVE Data

Outside the kernel, about one description in 24 won’t tell you why the bug matters. For kernel CVEs, by the kernel’s own policy, almost none will. When it doesn’t, you’ll need the CVSS vector, the references or the vendor advisory to find out. A model can tell you which records those are in under a second each, on a laptop.

What Didn’t Make the Cut

This started as a proof of concept, and I asked Clef plenty of other things on a sample of 300 CVEs first: whether the CVSS score was right, whether the CWE fit, and whether the record described a real vulnerability at all. CVSS came out as a reasonable second reviewer, about level with the other local models I tried. CWE caught real mistakes, like an out-of-bounds write filed as an out-of-bounds read, but I had nothing independent to confirm its other calls. And “is it a vulnerability” can’t be answered from the text, since CVEs later rejected as invalid read just as convincingly as the rest. Asking whether the description states an impact is the version of that question the text can answer. The experiments are in the repo if you want to pick them up.

Methodology and Reproducibility

  • Data: CVE List V5 from a local clone of CVEProject/cvelistV5 at commit 70f91a3c4fb4 (October 3, 2026, 13:55 UTC). All 27,489 PUBLISHED records with a datePublished in the 60 days before that commit, selected by publish date rather than ID year.
  • Models: Clef Flash (9B) and Clef 27B through Ollama’s /v1/systemone decision endpoint (Clef needs 0.35.1 or later), on an M5 Pro MacBook Pro with 64 GB of memory. Inputs are the record’s title and description only. Questions are versioned YAML in the repo, and answers are cached by question wording, so changing a question re-asks only that question.
  • Cascade: Flash answers both questions for every record in random order. Clef 27B re-answers the impact question for every record Flash marked “no impact” outside the Linux kernel CNA, plus a sample of 333 Linux kernel records as a check, and its answer is used where it exists. It never re-reads Flash’s “impact” calls. Every one of the 27,489 descriptions has a final answer.
  • Rule checks: 23 checks, 16 mapped to CNA Operational Rules 4.1.0 and 7 format checks on CWE IDs (against the CWE v4.20 catalog) and CVSS vectors, each reported as the share of a CNA’s records that pass.
  • Hand labels: 150 CVEs (129 published, 21 rejected), up to three per CNA from the last 30 days, labeled blind to both models’ answers; 149 of them after the question wording was frozen on October 3. They are in the repo.
  • Announcements and docs: Cloudflare’s Introducing Clef, Ollama’s decision models post and decision API docs, and the release notes for Ollama 0.35.0 and 0.35.1.
  • Code and per-CNA page: github.com/jgamblin/CLEFCVE and the per-CNA page.
  • Data collected October 3, 2026; Clef runs completed October 4, 2026.

I set out to test CVE descriptions and spent most of the weekend testing my own questions. What I came away with is a new kind of tool. A decision model can’t ramble and can’t answer outside the options, and Clef Flash made 27,489 judgment calls on a laptop in about six and a half hours of compute. Nimble, Tev1 and Clef all arrived in Ollama within a week of each other, so this is the start of the category, not the end of it. I will for sure keep pointing these models at CVE data, and the questions are in the repo if you want to beat me to it.