The case for bringing AI classification on-device

Understanding AI spend means knowing what the work was. Faros tested whether an open model can label sessions by activity locally, keeping conversations in-house.

An illustration showcasing a laptop running on on-device AI model.

The case for bringing AI classification on-device

Understanding AI spend means knowing what the work was. Faros tested whether an open model can label sessions by activity locally, keeping conversations in-house.

An illustration showcasing a laptop running on on-device AI model.
Chapters
No items found.

Why session classification needs a model

A company can know exactly what it spent on AI and still have little idea what work that spending supported. Classifying the conversations helps close that gap. Our tests with Clef show that this additional analysis can run on a developer’s laptop, without sending the conversations to another AI provider.

A conversation full of SQL might be about fixing a bug, comparing database designs, or planning a migration. Grouping all three as database work hides what the engineer was trying to accomplish. If those labels feed a report on AI usage, the difference between coding and research changes the story the report tells.

A language model lets us describe the categories we care about and ask where a conversation belongs. We can refine what “research” or “planning” means as the product develops, then test those definitions against real conversations without retraining the model for every revision.

Using a hosted model for each label means sending it the conversation and paying for another inference call. A local classifier can work beside the captured session data, with no per-call API charge or dependence on another provider being available. We tested Cloudflare’s Clef, a dedicated decision model, to see how much quality and speed we would give up for that control.

Why use an LLM to choose a label?

There is a long history of doing classification well with much smaller models. BERT's original paper showed how a pretrained language model could be fine-tuned for strong results across a range of language-understanding tasks. In a conventional classifier built this way, the output layer learns a particular set of labels from examples.

Suppose your product starts with a single “coding” category, then needs to distinguish new features from maintenance. For a classifier trained on the old categories, adding two names to the application does not teach it the new distinction. You need examples that establish the boundary, an updated model, and a way to check whether the change helped. A team still working out which distinctions its customers need can end up revisiting that training process repeatedly.

Researchers have spent years reducing that dependence on fixed labels. Entailment-based classification reframes the job as asking whether a passage supports a description, allowing labels the model did not see in task-specific training. More recent approaches such as GLiClass pursue flexible classification efficiently. These ideas predate the current enthusiasm for decision models.

General-purpose LLMs made the same flexibility accessible through an ordinary prompt. You can explain what “maintenance” means for your product, include an ambiguous example, and test the revised definition without training new weights. Each classification still involves running the model over the conversation and generating a response.

TypeSafe's Jev turns that request into a dedicated decision API. Give it the context and descriptions of the allowed answers, and it returns a choice and probabilities that code can use directly. Its training targets decisions and calibrated probabilities.

Clef brought the decision model onto our Mac

Cloudflare acknowledged the interest around Jev when it introduced Clef, then released the weights under Apache 2.0. We could download a model with this decision interface and test our own category definitions on the Mac without training a classifier from scratch.

Clef uses a 27-billion-parameter language model to read the input, then a trained decision head scores the permitted answers from the model’s internal representations. It returns those scores without generating an intermediate response token by token. Distinguishing a bug fix from an architectural question can still require substantial language understanding, even when the output is one word.

We tested it on 336 real AI-session excerpts, asking for the activity currently requested: coding, writing, research, operations, planning, or other. All candidates received the same excerpts, task instructions, and category definitions.

We used saved Astra xhigh classifications as the reference. Astra answered independently; it did not grade the other models, and none saw its labels. Throughout this article, a match means agreement with Astra, not a human verdict that the classification is correct.

Local Clef matched 84.8% of that reference at a 1.45-second median. Our strongest local Gemma 4B configuration, with thinking enabled, reached 61.3% and took 8.11 seconds.

Local Clef reached 84.8% agreement at 1.45 seconds; Gemma with thinking reached 61.3% at 8.11 seconds; local Clef-flash reached 59.8% at 0.42 seconds; Gemma without thinking reached 49.7% at 0.65 seconds.
Figure 1. Clef improved on both tested Gemma configurations in reference agreement, with much less delay than Gemma's thinking mode. All 336 first attempts count toward agreement; latency covers successful calls after model loading. Both axes start at zero.

Of the 60 excerpts Astra called coding, Gemma with thinking called 28 research; Clef did that once. The same group of requests would look much more research-heavy under Gemma.

The smaller local Clef-flash reached 59.8% agreement at a median of 0.42 seconds.

How much do we give up by keeping it local?

Local Clef finished about six percentage points ahead of Jev and four behind Sol, which had the highest agreement with our Astra reference.

Local Clef matched 285 of 336 Astra reference labels, between hosted Sol's 299 matches and hosted Jev's 264.
Figure 2. Local Clef sat between Jev and Sol on the same excerpts. Scores retain failed attempts in the denominator. Clef's lead over Jev has uncertainty: the paired session 95% interval was +1.2 to +11.3 percentage points, while grouping related sessions by project widened it to 0.0–12.4. Full results and statistical details are below.

Jev kept a substantial speed advantage. Its median decision took 211 milliseconds, compared with roughly a second and a half for Clef. If the user is waiting for a routing decision before anything else can happen, that difference deserves weight. A label added to a captured request in the background has a different latency requirement.

Configuration Median decision time API cost per 1,000
Clef · local 4-bit 1.45 s $0
Jev 1.13 · hosted 0.21 s $0.042
Sol 5.6 · xhigh, hosted 1.86 s $8.303
‍Successful-call medians, excluding model loading. Jev and Sol costs use available provider/gateway charge fields at the tested input mix. Local hardware and electricity were not priced. Hosted and local timings use different execution paths and were recorded on different dates.

At about four cents per thousand classifications, Jev is already inexpensive. Avoiding that bill alone would be a weak reason to reserve a large part of a developer’s memory for a local model. The stronger reason is control over the additional analysis: it runs beside the captured session data without another inference provider receiving the excerpt. The original AI assistant may still be cloud-hosted.

A local model still has to earn its memory

Deploying Clef means reserving memory that would otherwise be available to the developer’s own tools.

We ran Clef on a 48 GB Mac using the MLX community 4-bit conversion. The pinned download is 16.3 GB, and peak allocation in the MLX runtime reached 16.1 GiB, before accounting for other applications and system memory. Four-bit quantization reduces the storage needed for the backbone's weights, but 27 billion parameters still occupy a substantial part of the machine.

That leaves a concrete deployment question: how much memory can the classifier use without slowing down the work it is there to understand? Our Mac had enough capacity for this experiment, but we did not measure contention with a representative developer workload. Clef’s p95 was 3.49 seconds, and a workstation with spare memory for background jobs is a different environment from a laptop already under pressure during a build.

For session analytics, we would put classification in the background, where a delayed label does not hold up the engineer’s next action. Test fresh sessions with the product’s own category definitions and inspect disagreements that would change the resulting reports. Then measure completion time and memory pressure while the machine does its usual work. Larger queues also need a throughput test; our single-request timings cannot establish capacity.

A model the application can own

What excites us about Clef is the combination of flexible categories and control over where inference runs. We downloaded open weights, supplied our descriptions, and ran the classification on a Mac. It agreed with our frontier reference more often than Jev did on the same excerpts, with a median decision time of about a second and a half. We did not have to train a classifier specifically for those labels.

For a developer building session analytics, that changes the architecture worth considering. A cloud assistant can generate the code while a local model classifies the captured request. The classification feature can evolve through category descriptions and local evaluation, without requiring a separate inference service for every label.

The broader opportunity is to build more of these decisions into the application itself. Clef gives developers an open model they can run and categories they can define, with results in this study that held up against hosted alternatives. For background session classification, local inference deserves a place in the design from the start.

Study notes

We used one bounded excerpt from each of 336 eligible sessions in one account's available collection: 222 Codex, 101 Claude, and 13 Cursor. The earlier 42-session pilot was excluded. These labels describe the currently requested activity, not whole-session outcomes or productivity.

Frontier configurations requested xhigh reasoning. Gemma used ordinary generation with thinking off or on; this panel did not use the OpenJev harness. The Clef extension reused frozen excerpts and saved Astra labels without prompt tuning. All scheduled first attempts remain in the denominator, including failures. The Astra-reference analysis and the category-level reading of its results are exploratory.

These are comparisons of complete configurations with different sizes, training, and inference paths. They do not isolate the effect of a decision head or establish equivalence between Clef and Sol. The model-reference labels also cannot establish calibration against human-verified outcomes.

About Faros

Faros connects AI spend to engineering outcomes. Its Token Engineering platform brings together AI sessions and software delivery data so teams can see where tokens go, evaluate model choices against their own work, and govern AI usage. Book a demo.

‍

Chase Norton

Chase Norton

Chase is the Head of AI at Faros.

Graduation cap with a tassel over a dark gradient background.
AI ENGINEERING REPORT 2026
The Acceleration 
Whiplash
The definitive data on AI's engineering impact. What's working, what's breaking, and what leaders need to do next.
  • Engineering throughput is up
  • Bugs, incidents, and rework are rising faster
  • Two years of data from 22,000 developers across 4,000 teams
Research
10
MIN READ

The Speed Trap: 8 takeaways from our latest AI engineering research

AI made software development faster, but review gaps, QA bottlenecks, and rising incident volume reveal a new risk: the Speed Trap.

Research
10
MIN READ

Why AI coding agents actually fail (it's not the model)

Why do coding agents fail? We analyzed 4,000 errors across 6 models and discovered the real culprits.

Research
12
MIN READ

Routing Claude Code Opus 4.8 requests to GLM 5.2: a five-day live pilot

GLM 5.2 cut direct Claude Code request cost from $0.146 to $0.032 over five days. Why an image compatibility boundary still paused a broader rollout.