🔍 Read the full analysis: Where Does Jev Fit In AI Decision-Making? 24 Use Cases on ThorstenMeyerAI.com
Turn your quiet moments into listening time
- Thousands of audiobooks, podcasts and originals
- Listen on your phone, tablet or Echo — also offline
- Cancel anytime
TL;DR
Thorsten Meyer published a map of 24 potential uses for Jev, a tool that returns typed answers to narrow questions so software can route decisions. He says three uses are live in his publishing operation, 12 meet his fit test, seven need measurement, and two are poor fits. The examples supplied include results from publishing checks, but the wider list’s details are not available in the source material.
Thorsten Meyer published a map of 24 possible Jev use cases, reporting that three are already running in his publishing operation and that about 90,000 decisions have been processed there. The report describes Jev as a tool for answering narrow, typed questions at scale, with software acting on confident answers and routing uncertain cases elsewhere.
Meyer divides the 24 cases into three live uses, 12 strong fits, seven that need measurement and two poor fits. He says strong fits meet his four conditions: high volume, a narrow question, low-cost errors or a route for uncertain cases, and evidence that the existing heuristic fails. The supplied material describes the test and selected publishing examples; it does not include the full set of 24 cases.
In his account, Jev receives a state, such as text or JSON, plus typed questions. It returns a yes-or-no probability, a choice with probabilities and confidence, or a score on ordered levels. Meyer says a call takes about 0.3 to 0.9 seconds and costs about $0.04 per million input tokens. These are figures reported by the author, not independently verified in the supplied material.
The three live examples cover relevance, language and topic classification. Meyer says a relevance check assessed about 10,000 story-and-site pairings in three days, with 22% judged clearly on topic. A language check scanned 78,889 articles for $2.01, found 1,576 non-English articles and fixed 1,553. A classifier used Jev as a fallback when a primary large language model made errors; Meyer reports 89% agreement with a frontier model overall and 97% to 99% agreement when Jev’s confidence was at least 0.8.
24 use cases for Jev at a glance
Every use case, coloured by how well it fits
Proven in production
1Relevance gate: story and site2Language check3Classifier fallbackPublishing and content
4Thin-source detector5Same-event dedupe6Product fits the roundup7Disclosure present8Headline quality9Comment moderationCommerce and support
10Support-ticket routing11Return-reason coding12Review to feature complaints13Catalogue taxonomy14Order-fraud pre-triageSoftware and AI systems
15LLM guardrail16RAG passage filter17Citation check18Tool and intent routing19Log-line triage20PR risk triageBusiness ops and home
21Inbox triage22Expense categorisation23Lead qualification24Smart-home intent15 of 24 are ready to build or already running
How Confidence Shapes the Workflow
Meyer describes Jev as a tool for high-volume, narrow decisions, rather than for writing or summarizing material. His proposed workflow lets software act on clear cases and sends uncertain ones through another route. He reports using this approach to run checks across large collections, with uncertain classifications handled separately.
For publishers, the examples concern language, relevance and fallback classification. Meyer labels disclosure checks and comment moderation as strong fits, and headline quality as a case that needs measurement first. He links the proposed uses to operational problems, including evidence that an existing rule is failing.
Meyer reports article counts, costs and correction totals for a language scan. The supplied material does not include an independent audit, a detailed accuracy review for every check or information about the consequences of the changes. The figures describe his operation.
The Four Conditions Behind the Map
Meyer says each use should meet four conditions before deployment: high volume, a narrow question that does not require multi-step reasoning, errors that are inexpensive or uncertain cases that can be escalated, and a visibly failing heuristic. He says teams should keep a keyword rule when it works.
Before wiring in Jev, Meyer proposes replaying 300 to 500 past decisions, comparing results overall and by confidence band, and reviewing 20 disagreements to judge which system was right. He recommends deploying only where the high-confidence band reaches 95%, then using a separate feature flag, starting with a 5% to 10% canary and expanding afterward. These are the author’s proposed steps, not evidence that every listed use has completed them.
The supplied publishing examples show how the labels differ. A detector for thin sources is marked measure first because an existing 300-character rule may misread short items. Same-event deduplication is marked poor fit after Meyer says a canary found zero duplicates. Disclosure checks and comment moderation are labeled strong fits, with uncertain outcomes sent for human review or queued.
“Jev is the right tool wherever a system needs thousands of small judgements and can hand the unclear ones to something smarter.”
— Thorsten Meyer
What the Published Evidence Covers
The source material does not provide the names, questions, rules or results for all 24 use cases. It ends during its introduction to commerce and customer operations, so those cases and the rest of the map cannot be assessed from the material provided. The reported figures also come from Meyer’s own operation; independent validation is not included.
For several publishing ideas, the author identifies a need to measure the current error rate before deployment. The supplied account does not show whether those tests have since been completed, how performance varies across the live checks, or how often human reviewers overturn Jev’s answers. It also does not give the baseline behind every performance comparison.
Measure Before Wider Deployment
Meyer’s proposed next step for a prospective use is to test it against past decisions and review disagreements before enabling it in production. A team following that method would then deploy behind a flag, begin with a small canary and expand only after checking the results. The report does not announce a deployment date or say that the measure-first cases have passed those checks.
Readers seeking the remaining use cases would need the complete article or further detail from Meyer. Based on the supplied material, the immediate question is whether other workflows can demonstrate both a failing existing rule and reliable results on the cases Jev handles confidently.
Key Questions
What is Jev, according to the report?
Meyer describes Jev as a tool that takes text or JSON plus typed questions and returns answers, probabilities or scores that software can use to make routing decisions. He says it does not write, summarize or extract content.
How many of the 24 uses does Meyer say are ready or live?
He reports three live uses and 12 strong fits, for 15 in those categories. He labels seven measure first and two poor fit. The supplied source does not include details for every case.
What were the reported results from the language check?
Meyer says the check scanned 78,889 articles for $2.01, found 1,576 non-English articles and fixed 1,553. Those figures are reported for his own publishing operation.
How does Meyer recommend testing a new use?
He proposes replaying 300 to 500 past decisions, comparing outcomes by confidence band, and reviewing 20 disagreements. His suggested deployment starts with a feature flag and a 5% to 10% canary.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
