Why session classification needs a model
A company can know exactly what it spent on AI and still have little idea what work that spending supported. Classifying the conversations helps close that gap. Our tests with Clef show that this additional analysis can run on a developer’s laptop, without sending the conversations to another AI provider.
A conversation full of SQL might be about fixing a bug, comparing database designs, or planning a migration. Grouping all three as database work hides what the engineer was trying to accomplish. If those labels feed a report on AI usage, the difference between coding and research changes the story the report tells.
A language model lets us describe the categories we care about and ask where a conversation belongs. We can refine what “research” or “planning” means as the product develops, then test those definitions against real conversations without retraining the model for every revision.
Using a hosted model for each label means sending it the conversation and paying for another inference call. A local classifier can work beside the captured session data, with no per-call API charge or dependence on another provider being available. We tested Cloudflare’s Clef, a dedicated decision model, to see how much quality and speed we would give up for that control.
Why use an LLM to choose a label?
There is a long history of doing classification well with much smaller models. BERT's original paper showed how a pretrained language model could be fine-tuned for strong results across a range of language-understanding tasks. In a conventional classifier built this way, the output layer learns a particular set of labels from examples.
Suppose your product starts with a single “coding” category, then needs to distinguish new features from maintenance. For a classifier trained on the old categories, adding two names to the application does not teach it the new distinction. You need examples that establish the boundary, an updated model, and a way to check whether the change helped. A team still working out which distinctions its customers need can end up revisiting that training process repeatedly.
Researchers have spent years reducing that dependence on fixed labels. Entailment-based classification reframes the job as asking whether a passage supports a description, allowing labels the model did not see in task-specific training. More recent approaches such as GLiClass pursue flexible classification efficiently. These ideas predate the current enthusiasm for decision models.
General-purpose LLMs made the same flexibility accessible through an ordinary prompt. You can explain what “maintenance” means for your product, include an ambiguous example, and test the revised definition without training new weights. Each classification still involves running the model over the conversation and generating a response.
TypeSafe's Jev turns that request into a dedicated decision API. Give it the context and descriptions of the allowed answers, and it returns a choice and probabilities that code can use directly. Its training targets decisions and calibrated probabilities.
Clef brought the decision model onto our Mac
Cloudflare acknowledged the interest around Jev when it introduced Clef, then released the weights under Apache 2.0. We could download a model with this decision interface and test our own category definitions on the Mac without training a classifier from scratch.
Clef uses a 27-billion-parameter language model to read the input, then a trained decision head scores the permitted answers from the model’s internal representations. It returns those scores without generating an intermediate response token by token. Distinguishing a bug fix from an architectural question can still require substantial language understanding, even when the output is one word.
We tested it on 336 real AI-session excerpts, asking for the activity currently requested: coding, writing, research, operations, planning, or other. All candidates received the same excerpts, task instructions, and category definitions.
We used saved Astra xhigh classifications as the reference. Astra answered independently; it did not grade the other models, and none saw its labels. Throughout this article, a match means agreement with Astra, not a human verdict that the classification is correct.
Local Clef matched 84.8% of that reference at a 1.45-second median. Our strongest local Gemma 4B configuration, with thinking enabled, reached 61.3% and took 8.11 seconds.

Of the 60 excerpts Astra called coding, Gemma with thinking called 28 research; Clef did that once. The same group of requests would look much more research-heavy under Gemma.
The smaller local Clef-flash reached 59.8% agreement at a median of 0.42 seconds.
How much do we give up by keeping it local?
Local Clef finished about six percentage points ahead of Jev and four behind Sol, which had the highest agreement with our Astra reference.

Jev kept a substantial speed advantage. Its median decision took 211 milliseconds, compared with roughly a second and a half for Clef. If the user is waiting for a routing decision before anything else can happen, that difference deserves weight. A label added to a captured request in the background has a different latency requirement.
At about four cents per thousand classifications, Jev is already inexpensive. Avoiding that bill alone would be a weak reason to reserve a large part of a developer’s memory for a local model. The stronger reason is control over the additional analysis: it runs beside the captured session data without another inference provider receiving the excerpt. The original AI assistant may still be cloud-hosted.
A local model still has to earn its memory
Deploying Clef means reserving memory that would otherwise be available to the developer’s own tools.
We ran Clef on a 48 GB Mac using the MLX community 4-bit conversion. The pinned download is 16.3 GB, and peak allocation in the MLX runtime reached 16.1 GiB, before accounting for other applications and system memory. Four-bit quantization reduces the storage needed for the backbone's weights, but 27 billion parameters still occupy a substantial part of the machine.
That leaves a concrete deployment question: how much memory can the classifier use without slowing down the work it is there to understand? Our Mac had enough capacity for this experiment, but we did not measure contention with a representative developer workload. Clef’s p95 was 3.49 seconds, and a workstation with spare memory for background jobs is a different environment from a laptop already under pressure during a build.
For session analytics, we would put classification in the background, where a delayed label does not hold up the engineer’s next action. Test fresh sessions with the product’s own category definitions and inspect disagreements that would change the resulting reports. Then measure completion time and memory pressure while the machine does its usual work. Larger queues also need a throughput test; our single-request timings cannot establish capacity.
A model the application can own
What excites us about Clef is the combination of flexible categories and control over where inference runs. We downloaded open weights, supplied our descriptions, and ran the classification on a Mac. It agreed with our frontier reference more often than Jev did on the same excerpts, with a median decision time of about a second and a half. We did not have to train a classifier specifically for those labels.
For a developer building session analytics, that changes the architecture worth considering. A cloud assistant can generate the code while a local model classifies the captured request. The classification feature can evolve through category descriptions and local evaluation, without requiring a separate inference service for every label.
The broader opportunity is to build more of these decisions into the application itself. Clef gives developers an open model they can run and categories they can define, with results in this study that held up against hosted alternatives. For background session classification, local inference deserves a place in the design from the start.
Study notes
We used one bounded excerpt from each of 336 eligible sessions in one account's available collection: 222 Codex, 101 Claude, and 13 Cursor. The earlier 42-session pilot was excluded. These labels describe the currently requested activity, not whole-session outcomes or productivity.
Frontier configurations requested xhigh reasoning. Gemma used ordinary generation with thinking off or on; this panel did not use the OpenJev harness. The Clef extension reused frozen excerpts and saved Astra labels without prompt tuning. All scheduled first attempts remain in the denominator, including failures. The Astra-reference analysis and the category-level reading of its results are exploratory.
These are comparisons of complete configurations with different sizes, training, and inference paths. They do not isolate the effect of a decision head or establish equivalence between Clef and Sol. The model-reference labels also cannot establish calibration against human-verified outcomes.
About Faros
Faros connects AI spend to engineering outcomes. Its Token Engineering platform brings together AI sessions and software delivery data so teams can see where tokens go, evaluate model choices against their own work, and govern AI usage. Book a demo.




.webp)
