TypeSafe’s own benchmarks put it at 193.6x faster and 444.6x cheaper than LLMs. Independent tests usually find smaller gaps, though still large ones.
. Agent infrastructure and harness engineering
This is the largest category by adoption. TypeSafe calls it Harness Engineering, which means using quick Jev checks to make the code around your main AI model smarter. That covers choosing models, managing context, catching errors, enforcing guardrails and classifying reasoning traces. Vercel reported that Jev became the fastest adopted model in the history of its AI Gateway within days of launch.
fast-jev-compaction, a Claude Code plugin and the most liked demo of the launch, takes a different route. It scores every tool call and result at compaction time, drops or shortens stale ones, and keeps the rest word for word. It passed 4,000 GitHub stars within days. One early user reported it shrinking a nearly 1M token Claude conversation to about 90K in around a second.
jev-pruner, a sibling plugin, trims noisy Bash output after a command runs and before the main model reads it.
Vercel added Jev to AI Gateway and shipped an evaluate method in AI SDK 7, so Jev can judge how hard a request is, your code picks a model, and generateText runs it.
Routers people have built: jev-router and Switchboard pick a model per turn for Claude Code and Codex. Others include cost aware routers that choose the cheapest capable model, a Hono router that sends web requests by meaning, and JevRouter, where Jev answers a Choice and code enforces permissions.
Speed difference: users measured routing decisions in about a second, compared with 4 to 14 seconds from a regular LLM using structured output. Another logged decisions in 145 to 271 milliseconds.
Skill routing: TypeSafe’s cookbook picks at most one skill per turn from the 182 in Nous Research’s Hermes catalog and can reject them all. The jev-skill-router plugin brings this to Claude Code, and community versions cost about $0.001 per routed turn.
Guardrails and verification
Because each check is so cheap, you can run one on every LLM input, output and tool call. Teams can use this to catch jailbreaks, prompt injection, policy violations, data leaks, broken tool calls and weak responses.
Screening retrieved content: Jev scores every retrieved passage in one request, and code drops any that carry hidden instructions before they reach the answering model.
Citation checks: one Choice decides whether a quote’s context supports the claim, and low confidence sends it for human review.
Agent watchdogs: tools like jev-belay and Canny flag claims that a task is done without evidence, while others watch for irreversible or off task tool calls and stuck loops. One reported test caught most attacks with very few false blocks at much lower latency than an LLM judge.
Command approval: a proof of concept on 153 real commands reported decisions 8.7x faster and 4.4x fewer approval prompts.
All of these share one pattern. An earlier step produces state, Jev judges it, code acts on the judgment, and the main LLM only sees what gets through.
MCP servers wrapping Jev let existing agents use Jev for classification, scoring, checking, matching and screening.
Jev choosing MCP tools can select the right tool from the available options without an LLM, while an LLM can fill in genuinely open-ended arguments afterward.
Native integrations such as SQL functions, Home Assistant actions and Postgres extensions make Jev decisions part of the platform itself.
Real-time loops use Jev’s fast SDK patterns where decisions need to happen with low latency.
No comments:
Post a Comment