Jev: The Complete Guide to TypeSafe's System One Model

57 min read

Jev makes typed, calibrated decisions in about 100ms for a fraction of a cent. How it works, real output, and an editor skill I built and tested.

Part of the AI Agents topic hub.

Hero image for Jev: The Complete Guide to TypeSafe's System One Model
Table of Contents

I ran this entire post through an AI model 24 times in 1.1 seconds. The bill was $0.0023.

The model is Jev, from TypeSafe AI. You give Jev some context, like a support ticket or a draft, plus a few specific questions. It gives back typed answers with probabilities attached. Here are three real answers from the quickstart further down:

  • Which team should handle this ticket: technical, with 0.99 confidence
  • Whether the message asks for a refund: 0.93
  • How severe the bug is, from 0 to 2: 2.0

Each call took about 100 milliseconds and cost two thousandths of a cent.

I care about this because of my day job. At Refound we build production AI agents, and nearly all of them have decisions buried inside. One client’s agent reads every inbound email and decides what kind of email it is. Another scans social media and decides whether a brand fits a VC firm’s investment criteria.

Today each of those decisions is a full LLM call. It’s slow, it adds up at volume, and when the model says it’s “confident,” that number isn’t calibrated to anything.

Jev might just be the answer to that.

To find out, I built something I wanted: an editor for my blog posts. Jev judges a draft on AI tells, voice, editing, and SEO, and an LLM does the rewriting. We’ll build a small version in code here, one that edits a single paragraph, and you can download the full version as a skill.

By the end of this guide you’ll know what Jev is, how to make your first call, where it breaks, and whether it belongs in your stack.

What Is Jev?

Jev is a model from TypeSafe AI that answers typed questions about a piece of text or data, and returns a calibrated probability with every answer. It doesn’t generate text.

TypeSafe calls it a “System One model,” and if you’ve read Daniel Kahneman’s Thinking, Fast and Slow you already know where this is going. System One is your fast, intuitive brain. It reads a face and knows it’s angry before you’ve thought about it. System Two is slow and deliberate. You use it for long division or planning a trip. ChatGPT, Claude, and Gemini are all System Two models. They reason, plan, and generate one token at a time, which is why they’re slow and expensive.

TypeSafe’s bet is that a big share of what we ask System Two models to do doesn’t need System Two. Routing a ticket, flagging a comment, and scoring a lead are judgment calls with a small number of possible answers. Jev is built to make that kind of call quickly.

I buy the framing, because it matches what I see in client work. Most of the LLM calls inside our agents exist to decide something, and only a few of them write or require deep reasoning.

How is Jev different from an LLM?

An LLM answers by writing. It produces one token, feeds that token back in, and produces the next one, until the answer is finished. If what you wanted was a decision, your code then has to dig it out of the prose, and nothing in that prose tells you how much to trust it.

Jev never writes. Instead, you send Jev a piece of state (any text or structured data) and one or more typed questions about that state. Every question is evaluated independently, in parallel, in the same request. Five questions or fifty, it’s roughly the same call. You get back an answer for each one.

Here’s the same support ticket going through both:

The Jev numbers are the real output from the quickstart further down. The LLM answer is illustrative.

Watch the token counter on the left. The LLM’s answer grows one token at a time, so a longer answer takes longer, and when it’s done your code still has to find “technical” and “blocking” inside a paragraph. On the right, all three answers arrive together, and a fourth question wouldn’t slow anything down. The values go straight into an if-statement.

How Does Jev Work?

What’s fascinating here is that I haven’t defined what “technical” or “blocking” means yet Jev understands and makes the right call. Now, TypeSafe hasn’t published an architecture paper or a parameter count, so everything in this section is their explanation.

Jev is trained with a method they call RLCD, Reinforcement Learning for Calibrated Decisions. RLHF, the method behind ChatGPT, optimizes for what a human rater prefers. RLCD optimizes for how well the model’s stated probabilities match reality. If Jev says “70% confident” a thousand times, about 700 of those calls should be right. You can’t say that about a chat model’s stated confidence, and it matters when you’re building. You can put a threshold on a calibrated number and know roughly how often you’ll be wrong.

The other piece is what they call a parallel sampler. A normal LLM generates one token at a time, each conditioned on everything before it. Jev evaluates all of your questions against the same state in one shot. TypeSafe quotes 70 to 500 milliseconds per call, whether you ask one question or a dozen. My own calls came in between 85 and 412 milliseconds.

The three question types

Every question you send Jev is one of three shapes.

Choice picks one option out of a closed set:

Choice(
    instructions="Which team should handle this ticket?",
    criteria={
        "billing": "Charges, invoices, and refunds",
        "technical": "Bugs, outages, integration failures",
        "account": "Login, permissions, profile changes",
    },
)

You get back the selected option, plus a probability for every option.

Score places a judgment on a rubric:

Score(
    instructions="How severe is this bug?",
    criteria=[
        "Cosmetic or informational",
        "Degraded, but a workaround exists",
        "Blocking, no workaround",
    ],
)

Levels are numbered from 0, so this three-level rubric runs from 0 to 2, and the returned score can land between two levels, like 1.6. To put any Score on a 0 to 1 scale, divide by len(criteria) - 1.

Noul is TypeSafe’s name for a yes/no question, returned as a probability rather than a boolean:

Noul(instructions="Does this message request a refund?")

0.99 means yes. 0.5 means the model doesn’t know, not “medium.”

Confidence

Every Choice and Score answer comes back with a confidence score, a single number from 0 to 1 describing how concentrated the probability distribution is. Nouls don’t carry one. With only two outcomes, the probability itself already tells you everything. If Choice puts 91% of its mass on one option, confidence is high. If it’s spread evenly across four options, confidence is low, even if one of them technically wins.

TypeSafe’s own guidance is roughly: above 0.9, act automatically; in the middle, confirm or flag for review; below 0.5, don’t act at all, route to a human or fall back. I’d go further: the threshold should move with the cost of being wrong. A misrouted “check my balance” request is a minor annoyance, and a misrouted “wire this transfer” request can cost someone money.

Here’s that idea with three different messages going through the same question:

The question and thresholds come from the example in TypeSafe’s confidence docs. The three messages and their numbers are illustrative.

Those three question types and the confidence score are the whole API, and the rest of this guide combines them.

Quickstart

Install the SDK and set your key:

pip install typesafe-sdk
export TYPESAFE_API_KEY=your-key-here

Then ask all three question types about one support ticket, in one call:

from typesafe_sdk import Choice, Noul, Score, TypeSafeClient

client = TypeSafeClient()  # reads TYPESAFE_API_KEY

ticket = (
    "Hi, I've been trying to connect my Stripe account for 3 days and it keeps "
    "failing with a 500 error. We launch Monday and I can't take payments. "
    "If this isn't fixed today I need my money back."
)

response = client.system_one(
    state=ticket,
    questions={
        "team": Choice(
            instructions="Which team should handle this ticket?",
            criteria={
                "billing": "Charges, invoices, and refunds",
                "technical": "Bugs, outages, integration failures",
                "account": "Login, permissions, profile changes",
            },
        ),
        "severity": Score(
            instructions="How severe is this bug?",
            criteria=[
                "Cosmetic or informational",
                "Degraded, but a workaround exists",
                "Blocking, no workaround",
            ],
        ),
        "wants_refund": Noul(instructions="Does this message request a refund?"),
    },
)

team = response.answers["team"]
print(team.choice, team.probabilities, team.confidence)
print(response.answers["severity"].score)
print(response.answers["wants_refund"].noul)

Here’s what came back when I ran it:

technical {'billing': 0.0, 'account': 0.0, 'technical': 1.0} 0.99
2.0
0.93

Technical, blocking, and a likely refund request. The ticket mentions money, and Jev still gave billing a probability of zero. The refund signal showed up in the Noul instead. Asking several narrow questions instead of one broad one keeps each answer clean.

I timed five calls in a row from my laptop: 245ms for the first, then 85, 119, 91, and 96. The request was 450 input tokens. At $0.042 per million tokens, I’d have to run this ticket about 50,000 times to spend a dollar. That’s wild.

Building a Self-Editing Article Loop

So far I’ve shown Jev in isolation: ask a question, get a typed answer back. In practice Jev is one part of a larger system. It’s the fast, cheap judgment layer next to something slower and more capable.

This is a fuzzy problem. You can’t decide “is this draft ready to publish” with a word count or a keyword filter. Someone has to judge whether the opening earns attention, whether the piece makes an argument, and whether it reads like a person wrote it or like the fortieth “in today’s fast-paced digital landscape” paragraph. Jev is built for calls like this. They’re closed-ended and semantic, and the usual alternative is a full LLM call per question.

I wanted my Jev editor (Jevitor?) to judge my content in 4 areas:

  1. AI tells. Does it read like an AI model wrote it?
  2. Voice. Does it sound like me?
  3. Editorial. Does the opening earn attention, does the piece make an argument, is there padding?
  4. SEO. Is it set up to be found?

A whole post is a lot to judge, so I’ll build this at two sizes.

First, a small version in code that edits a single paragraph. It only handles the editorial part, and it’s short enough to show every moving part: the questions, the scoring, the feedback, and the rewrite loop.

Then the full version, which handles whole posts and all four areas. A post is too long to judge the way you judge a paragraph, so that one works differently and asks different questions. It’s a skill you can download, and it gets its own section after this build.

Jev Editing A Paragraph

Each pass, Jev gets two things as state: the draft, which is a single paragraph here, and my notes, which are the facts the writer is allowed to use. Then it answers eight editorial questions in one call:

from typesafe_sdk import Choice, Noul, Score, TypeSafeClient

jev = TypeSafeClient()  # reads TYPESAFE_API_KEY

QUESTIONS = {
    "has_specific_hook": Noul(
        instructions="The opening 1-2 sentences of `draft` make a specific, concrete claim or observation, not a generic statement anyone could write about this topic."
    ),
    "makes_real_argument": Noul(
        instructions="`draft` stakes out an actual opinion or takeaway, rather than just describing or surveying the topic neutrally."
    ),
    "has_filler": Noul(
        instructions="`draft` contains noticeable padding, repetition, or throat-clearing that doesn't add new information."
    ),
    "matches_audience": Noul(
        instructions="`draft` speaks to a technically literate reader's actual level and concerns, rather than a generic reader."
    ),
    # The grounded check: Jev compares the draft to the notes it was given.
    "unsupported_claims": Noul(
        instructions="`draft` contains a specific fact, number, quote, person, company, or anecdote that is not supported by `notes`."
    ),
    "insight_density": Score(
        instructions="How much non-obvious insight does `draft` contain, beyond describing the topic?",
        criteria=[
            "Purely descriptive -- no insight beyond restating the topic",
            "One genuinely useful insight or example",
            "Several useful insights or examples",
            "Dense with non-obvious insight throughout",
        ],
    ),
    "structural_clarity": Score(
        instructions="How well organized and scannable is `draft`?",
        criteria=[
            "Disorganized -- hard to follow the throughline",
            "Loosely organized",
            "Clearly organized with a visible throughline",
            "Tight, well-sequenced, nothing out of place",
        ],
    ),
    "primary_blocker": Choice(
        instructions="What is the single biggest thing holding `draft` back from being published as-is?",
        criteria={
            "ready": "Nothing significant; it can be published as-is",
            "needs_stronger_hook": "The opening is generic and doesn't earn attention",
            "needs_more_specificity": "Too vague or generic; lacks concrete detail",
            "needs_cutting": "Padding or repetition should be cut",
        },
    ),
}

Seven of these judge the writing. The eighth, unsupported_claims, compares the draft against the notes. Jev can’t check a claim against the world. It can check a draft against source material you hand it. The loop needs that because an LLM asked for specifics will happily invent them, and Jev will reward the invention. There’s an example under Where Jev Breaks.

None of these questions know about each other, and that’s deliberate. Jev reads the state fresh for each one, so the answer to has_filler can’t leak into insight_density. It’s also why this is one API call instead of eight.

Try It Yourself

Run the editorial gate on your own paragraph

Paste a paragraph below and run the seven writing questions from the build above against the live Jev API. (The eighth question checks a draft against your notes, and there are no notes here, so it sits this one out.)

Bring your own TypeSafe API key. You can get one at typesafe.ai. Your key is sent straight to a small server-side proxy that forwards it to TypeSafe and is never stored or logged on our end. It's kept in your browser's session storage only, so you don't have to re-enter it if you come back to this tab.

Questions to Ask

Turning eight answers into one decision

Jev hands back eight independent answers. Deciding what they mean together is work your code does, on purpose:

WEIGHTS = {
    "has_specific_hook": 0.20,
    "makes_real_argument": 0.20,
    "has_filler": -0.15,
    "matches_audience": 0.10,
    "insight_density": 0.20,  # normalized to /3
    "structural_clarity": 0.10,  # normalized to /3
}

PASS_THRESHOLD = 0.75
MAX_UNSUPPORTED = 0.30  # a hard condition, kept out of the weighted score on purpose
MAX_ITERATIONS = 3



def composite_score(answers: dict) -> float:
    score = 0.0
    for name, weight in WEIGHTS.items():
        a = answers[name]
        score += weight * (a.score / 3 if a.type == "score" else a.noul)
    return max(0.0, min(1.0, (score + 0.15) / 0.95))  # rescale the -0.15..0.8 range onto 0..1


def passes(answers: dict) -> bool:
    return (
        composite_score(answers) >= PASS_THRESHOLD
        and answers["primary_blocker"].choice == "ready"
        and answers["unsupported_claims"].noul <= MAX_UNSUPPORTED
    )

The weights are my editorial opinion, since Jev doesn’t have one. If you’d rather forgive a weak hook on a piece with strong insight, turn has_specific_hook down and insight_density up. The notes check sits outside the weighted score as a hard condition, because no amount of good writing should let an invented fact through.

The part Jev can’t do: turning a diagnosis into instructions

Jev never writes anything. It hands back has_specific_hook: 0.12 and nothing else. If you want an LLM to act on that number, your code has to translate it:

BLOCKER_INSTRUCTIONS = {
    "needs_stronger_hook": "Rewrite the first two sentences so they open on the most specific fact in the notes.",
    "needs_more_specificity": "Replace general statements with concrete details from the notes.",
    "needs_cutting": "Cut any sentence that restates an earlier one.",
}


def build_feedback(answers: dict) -> list[str]:
    """Turn Jev's numbers into instructions. An empty list means there is
    nothing concrete to ask for, and the loop should stop rather than guess."""
    issues = []
    if answers["unsupported_claims"].noul > MAX_UNSUPPORTED:
        issues.append("Remove every fact, number, person, company, or anecdote that is not in the notes. Do not replace it with another invented one.")
    if answers["has_specific_hook"].noul < 0.5:
        issues.append("Open with a specific, concrete claim taken from the notes, not a generic statement.")
    if answers["makes_real_argument"].noul < 0.5:
        issues.append("Stake out an actual opinion or takeaway instead of surveying the topic neutrally.")
    if answers["has_filler"].noul > 0.5:
        issues.append("Cut padding and repeated phrasing. Every sentence should add something new.")
    if answers["insight_density"].score < 1.5:
        issues.append("Build the paragraph around the most non-obvious point in the notes.")
    blocker = answers["primary_blocker"].choice
    if blocker in BLOCKER_INSTRUCTIONS and BLOCKER_INSTRUCTIONS[blocker] not in issues:
        issues.append(BLOCKER_INSTRUCTIONS[blocker])
    return issues

Every instruction maps to a signal Jev raised. When nothing concrete is wrong, the function returns an empty list, and the loop stops there.

The rewrite, and the loop

import anthropic

claude = anthropic.Anthropic()  # reads ANTHROPIC_API_KEY


def evaluate(draft: str, notes: str) -> dict:
    return jev.system_one(state={"notes": notes, "draft": draft}, questions=QUESTIONS).answers


def rewrite(draft: str, issues: list[str], notes: str) -> str:
    feedback = "\n".join(f"- {i}" for i in issues)
    message = claude.messages.create(
        model="claude-sonnet-5",
        max_tokens=1024,
        messages=[
            {
                "role": "user",
                "content": (
                    "Revise the draft below to address every issue listed. "
                    "Keep roughly the same length.\n\n"
                    "Rules:\n"
                    "- Use only facts that appear in the author's notes. Never invent an example, "
                    "number, person, company, or anecdote.\n"
                    "- If an issue can't be fixed with what's in the notes, leave it unfixed.\n"
                    "- Return only the revised draft. No commentary, no notes to the editor.\n\n"
                    f"Author's notes:\n{notes}\n\n"
                    f"Issues to fix:\n{feedback}\n\n"
                    f"Draft:\n{draft}"
                ),
            }
        ],
    )
    return message.content[0].text


def self_edit(draft: str, notes: str) -> tuple[str, str, int]:
    best = (-1.0, draft)  # (score, draft) among drafts with no unsupported claims

    for i in range(1, MAX_ITERATIONS + 1):
        answers = evaluate(draft, notes)
        score = composite_score(answers)
        unsupported = answers["unsupported_claims"].noul
        print(f"[iteration {i}] score={score:.2f}  blocker={answers['primary_blocker'].choice}  unsupported={unsupported:.2f}")

        if unsupported <= MAX_UNSUPPORTED and score > best[0]:
            best = (score, draft)
        if passes(answers):
            return draft, "ready_to_publish", i

        issues = build_feedback(answers)
        if not issues:
            # Jev has no concrete complaint left. Another rewrite would be a guess.
            return best[1], "needs_human_review", i
        print("\n".join(f"    - {x}" for x in issues))
        if i < MAX_ITERATIONS:
            draft = rewrite(draft, issues, notes)

    return best[1], "needs_human_review", MAX_ITERATIONS

The rewrite prompt limits Claude to the facts in my notes. The loop remembers the best draft that passed the notes check, and returns that one, which isn’t always the last one. And since we can’t guarantee this converges, we set a MAX_ITERATIONS flag.

This is also where cost and latency start to matter. Judging a draft with a full LLM call three times, on top of the rewrites, adds seconds and cents to every piece. The two Jev calls in my run added a third of a second, and a fraction of a cent.

Running it on a bad paragraph

Let’s feed it something deliberately bad, along with seven bullet points of things I know to be true:

slop_draft = """In today's fast-paced digital landscape, AI agents are revolutionizing
the way businesses operate. These powerful tools leverage cutting-edge technology to
streamline workflows and boost productivity. It's important to note that AI agents
offer a wide range of benefits, from automating repetitive tasks to providing valuable
insights. By harnessing the power of AI agents, companies can unlock new levels of
efficiency and stay ahead of the competition."""

final_draft, status, iterations = self_edit(slop_draft, NOTES)
NOTES = """- I build production AI agents for clients.
- Most of the LLM calls inside those agents make a decision. Few of them write anything.
- Example: one agent reads every inbound customer email for an e-commerce brand and decides what kind of email it is before anything else happens.
- Example: another agent scans social media and decides whether each brand it finds meets a VC firm's investment criteria.
- Today each of those decisions is a full LLM call. That is slow, and it adds up at volume.
- When an LLM says it is "confident" in a classification, that number is not calibrated.
- These decisions are closed-ended: a fixed set of categories, or a list of yes/no criteria."""

Here’s the real output:

[iteration 1] score=0.12  blocker=needs_more_specificity  unsupported=0.14
[iteration 2] score=0.76  blocker=ready                   unsupported=0.09

Final status: ready_to_publish after 2 iteration(s)

Every score and blocker in this animation is from the real run.

It went from 0.12 to 0.76 in one rewrite, and Jev called it ready on the second pass. Here’s the draft:

Most of the LLM calls inside the production agents I build aren’t writing anything — they’re making a decision. One agent reads every inbound customer email for an e-commerce brand and decides what kind of email it is before anything else happens. Another scans social media to decide whether each brand it finds meets a VC firm’s investment criteria. Both are closed-ended: a fixed set of categories, or a list of yes/no criteria. Yet today each of those decisions runs as a full LLM call — slow, and it adds up at volume. Worse, when the model says it’s “confident” in a classification, that number isn’t calibrated. Treating a closed-ended decision like open-ended generation is the real waste in these systems.

Pretty good, right? The draft above still has two em-dashes and an “aren’t writing anything, they’re making a decision” line but I’ll get to that in the next section.

The Jev Editor Skill

The loop above edits one paragraph. It can ask eight broad questions about “the draft” because a paragraph is small enough for Jev to take in at once.

A whole post is too long for that. So the full Jevitor works differently: it cuts the post into pieces and asks each piece a few narrow questions. That means its questions are different from the eight you just saw. It also covers all four areas, where the loop only covered editing.

I’ve packaged it as an agent skill for Claude Code and other agents. Your coding agent plays the writer, Jev plays the judge, and it stops after three passes, but you can change that.

It has three layers:

  • Plain code handles anything countable: em-dashes, title length, link counts, leftover TODOs. Jev is unreliable at counting, and code is exact and free.
  • Jev handles narrow judgments on small pieces of the post: the opening, each section, the title, the closing.
  • Your agent, which can reason, gets the jobs Jev failed in testing: checking facts, checking the intro’s promises against the body, and asking whether a newcomer could follow along.

The scores are this post’s real results: its first draft, and the version you’re reading.

The questions it asks

Here are the pieces, and what Jev is asked about each one:

Piece of the postWhat Jev is asked about itArea
The opening, everything before the first headingIs the hook specific? Does it say what you’ll get?Editorial
Is it a cliche?AI tells
Does it state the takeaway in the first 100 words?SEO
Each section, one at a time9 questions, one for each AI-writing patternAI tells
How closely does it match your own writing?Voice
The closing sectionDoes it give a next action? Is it only a recap? Does it land?Editorial
The title and descriptionIs the title honest? Is it specific? Does the description give a reason to click?SEO
The whole postDoes it point to parts of itself that don’t exist?Editorial
Is there a one-sentence definition of the main concept?SEO

So every question belongs to one of four areas: AI tells, voice, editorial, or SEO. Here’s what each area checks, including the parts that are plain code and never go to Jev.

AI tells. These are the annoying writing tics that AI uses. (I’ve explained more about this in how I write with AI without creating slop.) You can put these rules into a skill and Claude will still miss some, so Jev is a great choice here:

TellWhat it looks like
Negative parallelism”It’s not X, it’s Y” and its cousins
Empty intensifiers”genuinely,” “truly,” “actually” used for emphasis
Question, then answer”The result? Devastating.”
SignpostingAnnouncing what you’re about to say
Superficial analysisCommentary on what a fact “shows” that adds no new fact
Rule of threeTriplets used for rhythm
Significance inflation”Plays a crucial role,” with no fact behind it
Mic-drop closersOne-line paragraph endings that restate the point
Trailing “-ing” clauses”…, highlighting the importance of”
Cliche openerAsked about the opening only

A tell fails the post when it shows up in more than a third of your sections, so one mic-drop closer is fine and one in every section gets flagged. Code counts the rest: em-dashes, intensifier and AI-vocabulary density, stock phrases, stacked hedges, and whether your sentence and paragraph lengths vary.

Voice. Jev has no memory and has never read my blog, so it can’t know what I sound like unless I show it. The skill sends each section to Jev alongside five excerpts from my published posts, and asks how closely the section matches the author of those excerpts. It judges tone, rhythm, humor, and word choice, and ignores topic.

I tested it on writing it hadn’t seen:

TextVoice match, 0 to 1
Four sections from posts of mine that weren’t in the samples0.63 to 0.82
Three sections from the first draft of this post0.31 to 0.49
TypeSafe’s launch post0.12
The slop paragraph0.00

The first draft of this post didn’t sound like me. Makes sense. A section passes at 0.60, and a post passes when 60% of its sections do. You point the skill at your own published posts, so it learns your voice.

Editorial. The questions about the opening and the closing from the table above, plus the check for references to parts of the post that don’t exist, which only runs on posts under 6,000 words. Code catches leftover TODOs, paragraphs over 120 words, and code fences with no language. It also checks your calls to action: that the post has one, that there’s one near the end, and that no long stretch goes without one. Everything that needs a reader goes to your agent: checking facts, checking the intro’s promises against the body, and asking whether a newcomer could follow along.

SEO. Most of this is counting, so most of it is code. Jev only gets the four checks that need a reader.

CheckChecked by
Title is 30 to 60 charactersCode
Meta description is 110 to 160 charactersCode
At least 3 internal links and 3 outbound links to sourcesCode
Link text says where the link goes, with no “click here”Code
Every internal link points at a post that existsCode
Target keyword is in the title, description, first 100 words, and a headingCode
Images have alt text, there’s a hero image, and headings don’t skip levelsCode
At least two headings are phrased as questionsCode
The title matches what the post delivers, and says what you’ll getJev
The description gives a concrete reason to clickJev
The first 100 words name the topic and the takeawayJev
One sentence defines the central concept and stands on its ownJev

The last two are there for AI search, because answer engines tend to quote the top of a page and sentences that make sense out of context. My guide to generative engine optimization goes deeper on that.

Each area scores the share of its checks that pass. A post passes at 70 or higher on tells, editorial, and SEO, and 60 on voice.

Why so many questions?

That’s about two dozen questions, asked of small pieces, where you might expect four or five asked of the whole post. There are three reasons.

Jev goes blind on long inputs. I took my 9,000-word Claude Code guide, which has a troubleshooting section and plenty of caveats, and asked Jev whether the post names at least one limitation. It said 0.98. Then I deleted every paragraph that mentioned a limitation, an error, or a caveat, about 1,300 words, and asked again. It said 0.99. The answer didn’t depend on what was in the post.

It’s sharp on small ones. When I swapped a real title for a clickbait one, “the title matches the post” dropped from 0.77 to 0.14. So the skill asks Jev about the smallest piece that can answer the question, and only two questions ever see the whole post.

A narrow answer tells you what to fix. “This section uses ‘it’s not X, it’s Y’ twice” is an edit your agent can make in a minute. “This post reads a bit like AI” gives it nothing to work with. Every question in the skill names one pattern, and every flag comes with the section it was found in.

They’re cheap enough that there’s no reason to hold back. Feel free to add more questions if you need more specificity. I tested 54 candidates against posts I’m proud of and raw AI drafts, and kept the ones that could tell them apart.

What it said about this post

The first draft scored 64 on AI tells, 30 on voice, 73 on editorial, and 56 on SEO. It also had two hard failures: leftover TODO comments, and a reference to a section that didn’t exist. The version you’re reading scores 86, 60, 93, and 100. Voice only just clears its bar: four of ten sections are still under it. The skill told me which ones and I could have done a human pass on them to bring the score up but I left it as is to see if you can find it. A full run is about 24 requests and costs a quarter of a cent.

It’s free and open source. Drop your email below and I’ll send you the GitHub link with install and usage instructions. It works with Claude Code and any agent that reads skills.

More Patterns

The self-editing loop is one version of a general pattern: send a batch of typed questions, combine the answers in code, and route on confidence. Here are other places it fits. The animations in this section use illustrative numbers, not real runs.

Decisions inside production agents

This is the use case I care about most. The agents we build at Refound spend most of their money on judgment calls. One classifies every inbound customer email for an e-commerce brand before anything else happens. Another watches social media for a VC firm and decides whether each brand it finds meets the firm’s investment criteria. Both run a full LLM call per decision today.

Both decisions are closed-ended. The first is a fixed set of categories, which is a Choice. The second is a list of yes/no criteria, which is a handful of Nouls in one request. I haven’t moved either one to Jev yet. If I do, I’ll start in shadow mode: run Jev next to the existing LLM calls, compare answers on a few thousand real decisions, and switch the ones where they agree. If you’re new to how these agents are put together, start with my guide to AI agents.

Inside coding agents

Claude Code, Cursor, and Codex all have some version of “is this tool call safe to run,” and today most of that logic is either hardcoded or a full LLM call. SuperQode, an agent harness, describes wiring Jev in as a dedicated gate. A few questions run against each proposed tool call (is it within the granted permissions, is it destructive, could it leak data, is it on task), and the answers map to allow, deny, or ask a human. I wrote more about how harnesses work in my agent harness guide, and The Anatomy of Claude Code traces a request through a real one.

There’s also an open-source skill pack, awesome-jev-agent-skills, that plugs Jev into Claude Code and Codex for five narrower jobs: triaging which diagnostic to run on a failure, spotting missing test scenarios, reviewing a diff against stated invariants, checking a structured extraction against its source text, and checking a “done” claim against evidence. The maintainers are clear about the limits: “Jev evaluations are advisory.”

Screening what goes in and out of an LLM

TypeSafe’s guardrails cookbook runs every message through a set of independent Noul checks (jailbreak attempt, harmful request, medical dosage question, self-harm signal) plus one Score for how much harm complying would do. Your code routes to pass, review, block, or a support path based on where the numbers land. This handles cases a keyword block gets wrong. A mild “what’s the max dose of melatonin” question goes to human review instead of an auto-block, and a self-harm signal goes to a support path. It takes several independent Nouls, combined in code, to make that distinction.

Structured data pipelines

TypeSafe’s structured data extraction (SDE) cascade uses a cheap LLM for a first-pass extraction. Jev then checks each extracted field with a narrow Noul (“is this value plausible,” “does this match the source,” “is this a hallucination”). Only records that fail go to an expensive reasoning model. The cascade escalates when any single field trips its threshold, because averaging would let one confident red flag disappear into a healthy-looking mean.

Beyond Engineering

The same three question types fit plenty of business decisions:

Lead qualification. Reading a form submission or an intro email and judging seriousness, budget signal, and urgency is a semantic read on unstructured text, the same kind of call as the hook and argument questions in the build.

Trust and safety triage. Deciding whether a comment, review, or listing is spam, harassment, or a policy violation is the guardrails pattern pointed at user-generated content instead of LLM output.

Candidate screening. TypeSafe’s own composite-scoring example does this: score a resume independently on Python depth, leadership signal, system-design experience, and generalist range, then blend those into different weighted totals depending on whether you’re hiring an IC or a manager. Same weighted-formula idea as the composite score in the build, just pointed at a different problem.

Freelancer-to-project matching. Judging fit between a project brief and a freelancer’s portfolio, skill match, communication style, likely reliability, from unstructured text on both sides.

Legal review. Judging if a website change or marketing campaign passes legal compliance, and flagging exactly which requirement it fails so an LLM or a person can say what to change. This works like the notes check in the build: hand Jev the policy and the campaign, and ask whether the campaign breaks it.

Expense review. Most expense requests have free form text to explain the expense that requires humans to read. Jev can analyze it against the expense policy and auto-pass or fail.

Where Jev Breaks

TypeSafe publishes its own list of known failure modes for Jev 1.13. It’s specific, and my testing matched several items on it:

  • Literal reading. It answers what the instructions say, not what you meant. Implied conditions get missed.
  • Math and numbers. It can’t reliably count, or use a Score’s numeric gap as an actual magnitude. Do the arithmetic in code.
  • Dates and times. Treated as text, not ordered values. Relative dates and mixed formats trip it up.
  • Indirection. Multi-hop reasoning and double negatives are weak spots.
  • Large, noisy state. Accuracy drops as irrelevant context piles up. Filter before you send it.
  • Adversarial content. It assumes the input is honest by default, so prompt injection is a real risk.
  • Contradictory instructions. Confused output when your instructions and criteria don’t agree.
  • No structural guarantees. P(this is true) and 1 minus P(this is false) aren’t guaranteed to match. Don’t build logic that assumes they do.
  • No generation. It won’t write, and forcing it to try is slow and bad.

I’d read this list before I started, and I still walked into several of them.

Literal reading. An early version of my blocker question had an option called needs_fact_check: “Contains a claim that needs verification.” It fired on three of my four published posts. Read literally, every technical article contains a claim that needs verification, so that option could never lose. The notes check in the build replaced it, because that’s a question Jev can answer from the text in front of it.

Large, noisy state. That’s the 9,000-word experiment in the skill section. Whole-post questions stopped responding to what was in the post.

One that isn’t on their list: Jev can’t tell whether something happened. When I ran the loop without notes, the feedback asked for “a specific example,” and Claude wrote: “A logistics company I spoke with cut invoice-reconciliation time from three days to eleven minutes.” There is no logistics company. Jev scored that draft 0.92, because it judges what’s on the page, and an invented anecdote is very specific. With my notes in the state, unsupported_claims scores that same paragraph 0.98.

I saw it again with a question for “this section includes the author’s first-hand experience.” Raw AI drafts scored 0.81 on it, and my published posts scored 0.60. AI drafts write “I tested this” without blinking.

It isn’t deterministic. Send Jev the same state and the same question twice, and you can get two different numbers. I sent one identical request twice and unsupported_claims came back 0.15, then 0.18. Across four runs on that paragraph it ranged from 0.14 to 0.18. Confident answers hold still: the invented logistics paragraph scored 0.98 both times.

The wobble only matters near a threshold. One section of this post sits right on the 0.60 voice bar, so the post’s voice score flips between 60 and 70 from one run to the next. Don’t put a hard cutoff where your scores cluster. Leave a middle band that goes to a human, or ask a few times and average. TypeSafe’s self-consistency cookbook measures this by repeating each call 15 times.

Some caveats come from outside the docs. “Hallucination-free” gets repeated a lot, and it’s true in a narrow sense: Jev can’t return a value outside your schema. It can still pick the wrong value inside that schema, with high confidence.

There’s also a question about how TypeSafe measures itself. Their workflow evaluations use the average of two other LLMs’ outputs (GPT-6 and Fable 5.1) as the reference answer, with no independently verified answer key. TypeSafe argues this works against them, since it biases the reference toward OpenAI’s and Anthropic’s models. Either way, agreeing with two other LLMs is a different thing from being correct.

Composing several calibrated answers into one decision, which is what my loop does, doesn’t keep the combined decision calibrated. And Jev never explains itself. You get a number to audit later, with no rationale.

None of this has put me off. I’m keeping Jev away from math, dates, and untrusted input for now, and I’d run it in shadow mode against your own data before you trust a threshold. Treat TypeSafe’s benchmark numbers as a ceiling reached under ideal conditions.

Jev Pricing and Access

Jev costs $0.042 per million input tokens, and output is free since nothing is generated. A full editorial pass on this post, 24 requests and about 55,000 tokens, cost $0.0023. The context window is 32,000 tokens.

It’s open to everyone now, through TypeSafe’s own API and SDKs, Vercel’s AI Gateway, Cloudflare Workers AI, and OpenRouter. My recommendation: use whichever platform you’re already deployed on.

Should You Use Jev?

I’d use Jev today for routing, moderation, guardrails, high-volume tagging and scoring, and gating agent tool calls. These are closed-ended judgments where an occasional wrong answer is cheap.

I’d wait, or run it in shadow mode first, on anything involving arithmetic or date logic, anything facing adversarial input without extra hardening, and anything where a wrong decision has to be explained afterward, like a regulated approval.

From my own testing I’d add two rules. Keep the state small, because Jev was sharp on a title or a single section and close to blind on a 9,000-word post. It’s the same lesson as context engineering for agents: pick the right information, don’t send everything. And test every question against a known-good and a known-bad input before you trust it, because more than half of mine failed.

What sold me is the price. A full editorial pass on this post took 24 requests and 1.1 seconds, and cost $0.0023. At that price I can check every section of every draft on every pass, and I can put a judgment call inside an agent loop without thinking about the bill.

Here’s what to do next: take the quickstart, swap in a real ticket or email from your own product, and ask Jev three questions about it.

After that, get the Jev Editor skill, point it at a draft and three of your published posts, and see what it says about your writing.

Related Posts

Read Claude Managed Agents: Cloud Agent Tutorial
Hero image for Claude Managed Agents: Cloud Agent Tutorial
guide ai-agents claude

Claude Managed Agents: Cloud Agent Tutorial

Learn how Claude Managed Agents provides hosted containers, tools, sessions, and multi-agent orchestration for building autonomous agents.

13 min