Jev vs Local AI: CLM and Laya Put to the Test
RUNTIME.
0:00 / 0:00
Jev vs Local AI: CLM and Laya Put to the Test
8 934 просмотра · 3 дня назад
RUNTIME.
344 подписчика
8 934 просмотра · 3 дня назад
How do local decision models compare with Jev when their choices have consequences? We test CLM and Laya against hosted Jev on short decisions, sorting tasks and a small driving simulator—and show the successes, failures and limits.
RESEARCH, RECORDED RESULTS + INTERACTIVE REPLAY
https://github.com/Runtime-weekly/run...
Includes frozen requests/results, scoring, a data visualizer, and both versions of the code-only driving controller. The companion replays recorded model decisions; fresh inference requires the upstream model setup below.
WATCH NEXT
Kev vs Laya — our earlier local decision-model comparison:
• Open-Source Jev? Kev vs Laya — We Tested Both
Install Laya locally + useful demos (English checkpoint; this episode tests the separate typed-decisions variant):
• Open-Source Jev? Install Laya Locally + 3 ...
Tev1 vs Qwen — another local decision-model test:
• Open-Source Jev? Tev1 vs Qwen — Tested Loc...
CHAPTERS
0:00 Jev vs local decision models
0:22 What we are running
1:00 How CLM is built
1:37 Published results and our test plan
2:46 Short decision tests
4:07 Checking the CLM implementation
5:01 Sorting: rules, choices and results
7:51 Driving: rules, choices and results
10:02 Response times and deployment costs
10:45 What these results actually support
11:30 Next tests, GitHub and your results
MODELS + TEST CONDITIONS
• Jev: hosted jev-1.13.0 through TypeSafe's API.
• CLM: Contrastive-LM/CLM-v0.1-8B, with trained contrastive heads and a frozen Qwen3-8B backbone, running locally on NVIDIA DGX Spark.
• Laya: version 0.3.4, typed-decisions variant, running locally on CPU.
Our diagnostic set contains 24 authored short-decision cases, paired reversed-option runs, six reset sorting cases, and three text-based toy driving routes. This is a small task-specific comparison, not a general model leaderboard or a real-world driving test. The animations replay saved decisions; they are not live camera input or live inference.
CLM pilot/sorting runs used a Transformers adapter. We checked choice agreement on 60 saved requests against the reference vLLM implementation; the driving runs used vLLM directly. Choice agreement does not establish full numerical or general equivalence.
The code-only driving controller was revised after feedback to minimize lane changes, then rerun. Model results and the original controller remain unchanged in the repository. The code controller has privileged access to the simulator's exact transition rules; it is a reference for this environment, not a fair model leaderboard entry.
Response times reflect different local GPU, local CPU and hosted-network setups. Loading, first calls and caching differ, so these measurements are not a hardware-controlled speed ranking. Model scores are not calibrated accuracy estimates.
PRIMARY SOURCES + SETUP
CLM code: https://github.com/Contrastive-LM/CLM
CLM weights: https://huggingface.co/Contrastive-LM...
Qwen backbone: https://huggingface.co/Qwen/Qwen3-8B
Laya: https://huggingface.co/convaiinnovati...
Jev API: https://docs.typesafe.ai/api
Exact tested revisions and source notes:
https://github.com/Runtime-weekly/run...
Subscribe for the next model-architecture deep dive and more practical local-AI experiments. If this was useful, a like helps us know which tests to build next.
Built something interesting, found a better test, or got different results? Share your work in the comments, including prompts, model versions and settings so we can investigate and improve together.
#Jev #LocalAI #CLM