Special review · Agent systems · Early experiment analysis

Fast Decisions Inside Agents: Jev’s Structured Outputs and Practical Limits

Examining System One interfaces, calibration and early evaluations for research harnesses

Responsible editor: Hyun-Jung Kim · AI-assisted · Evidence cutoff 2026-09-21

Conceptual illustration of unstructured material passing through a sorting apparatus into distinct organized forms
Figure 1. AI-generated metaphor for turning material into bounded decisions. It depicts no real apparatus, neural architecture or measured data. Created with OpenAI imagegen.

Hyun-Jung Kim · AI-assisted · primary sources checked · evidence cutoff 2026-09-21 · Methods & agent disclosure

Agents repeatedly decide whether evidence is relevant, a failed task deserves another attempt, or a stronger model is needed. Generating a response for every small decision can make waiting for the model a substantial part of the workflow. Jev aims to make these judgments available through a short software call.

Assessment: worth examining as a way to change the cost and frequency of decisions in agent harnesses.

This review compares documentation and an early public experiment and checks ratios in published aggregates. It does not run the Jev API or reproduce the benchmark. The evidence does not establish general reasoning or scientific-computing superiority.

Expanded September 21, 2026: the shared LinkedIn post’s flight-search, driving and trading examples were checked against primary material and public code, with measurement boundaries added.

Decisions that software can consume

TypeSafe AI announced Jev on September 15, 2026. Its label System One emphasizes fast, focused judgments; it should not be read as a demonstrated model of human cognition or an established scientific taxonomy. Jev takes its name from William Stanley Jevons. [1]

Requests separate the material being assessed, state, from questions. Unlike a conventional large language model (LLM), Jev uses predefined answer spaces. It currently accepts text, including text in JSON objects and arrays, but no images, audio or video. Game demonstrations therefore do not establish visual perception or real-world robotic control. [2]

PrimitiveOutputReviewer-designed example
ChoiceAn option, option probabilities and confidenceClassify a run as normal, a retry candidate, needing investigation, or none of these
ScoreA distribution over ordered levels, its mean score and confidenceEvidence: unrelated, indirect or direct
NoulA 0–1 estimate that a statement is trueDoes the document state the requested experimental condition?

Choice allows up to 255 options. Designers can include other or not_stated. Noul has no separate confidence field. Thresholds need to respect these different meanings. [3] [4]

Parallel questions, composition in code

The documentation describes isolated, parallel evaluations against shared state. A broad question about whether a document can be processed can be decomposed into relevance, required information and ambiguity. Code combines the answers. Additional questions consume tokens but reportedly add little latency. Separate evaluation does not establish statistical independence or logical consistency across answers. [5]

Prepared evidence enters focused questions; code checks results before choosing automated handling, additional reasoning or human review
Figure 2. A proposed harness design based on the documented interface. Execution permissions and mandatory checks remain in code. This is not Jev's internal neural architecture.

Conceptually, latency includes preparing evidence, evaluating a model, generating sequential output and postprocessing. Removing open-ended generation can reduce one component. Input processing, network round trips, queueing and subsequent model calls remain. A model-call speedup is therefore different from a complete-workflow speedup.

If a replaceable decision stage accounts for fraction f of runtime and becomes s times faster, an elementary model gives total speedup 1 / [(1 − f) + f / s], assuming other costs remain fixed. In an illustrative scenario, making a stage that consumes 20% of runtime 100 times faster produces only about 1.25 times overall acceleration. Measure where time is spent first.

Probability, calibration and confidence

TypeSafe calls its training method reinforcement learning for calibrated decisions (RLCD). It distinguishes this goal from reinforcement learning from human feedback (RLHF) and reinforcement learning with verifiable rewards (RLVR). Merely enforcing a format does not capture RLCD’s stated aim: the returned probabilities should correspond to observed outcomes. The consulted material does not specify enough about rewards, training procedures, model size or architecture to reimplement Jev. A particular encoder or loss function cannot be inferred from the interface. [6]

Calibration is a property evaluated across predictions. Among cases assigned roughly 0.8 probability of relevance, approximately 80% should be relevant. This does not guarantee any individual answer. Calibration predates Jev: Guo and colleagues studied how neural-network accuracy and calibration can diverge and evaluated temperature scaling. [7]

The API's confidence field has another meaning. For Choice and Score, it summarizes the concentration of the returned distribution. The explanation page does not give its exact formula. A confidence value of 0.9 must not be translated into 90% accuracy. Calibrate probabilities and abstention policies on the intended workload. [8]

For Score, the default result is Σ k·pₖ, where k is a level index and pk its probability. Ordered labels do not automatically have equally spaced meanings. This score is not a validated estimate of an energy or another physical quantity. [9]

Promising observations with narrow comparison conditions

TypeSafe's headline 193.6-times speed and 444.6-times cost improvements come from its own workflow evaluation. It uses strong external models' mean predictions as references and requests probability-bearing structured outputs from competing LLMs. This differs from a broad benchmark with independently established human answers. [1]

LangChain's September 20 experiment instead evaluated five frozen weather-agent responses 100 times per judge. Binary agreement with one human reviewer and variation in continuous scores were measured separately. LLM sampling settings were left at provider defaults; the Jev service version was not retained in the experiment metadata. [10]

JudgeHuman-label agreementMean latency / callMean cost / call
Jev100.0%0.44 s$0.00035
GPT-5.6 Luna96.4%2.50 s$0.00039
GPT-5.6 Terra99.8%2.83 s$0.00289
Claude Sonnet 4.680.0%2.16 s$0.02811

These are reported repository values. There are only five distinct cases, not 500 independent problems. A judge can also be consistently wrong. [11]

Zero-baseline bars comparing mean per-call latency in the early LangChain experiment
Figure 3. A reconstruction of the table's latency values, not new measurements. The source table supplies no uncertainty intervals. Jev's speed advantage in this setting is approximately 4.9–6.4 times.

Dividing the rounded mean costs gives advantages of approximately 1.11 times over Luna, 8.26 over Terra and 80.31 over Sonnet. Published JSON aggregates also give a mean within-case quality variance of approximately 1.494×10−5 for Jev; competitors' variances are about 92–913 times larger. [12]

Repeatability deserves investigation, but low score variance proves neither calibration nor general accuracy superiority. Tuned baselines, many more independent cases, multiple human reviewers and longitudinal evaluation are needed to determine whether the gains transfer.

What the seven-second flight search actually combines

A LinkedIn post by Peter Cha (차예솔) [17] connects repeated classification to flight search, driving and trading demos. Its linked Jev Ultrafast project [18] makes the distinction between a decision model and an agent concrete: Jev returns decisions; surrounding software supplies observation, memory, execution and completion checks.

The browser agent converts observed document object model (DOM) controls into an indexed list. It asks for an operation and operation-specific targets together, then executes only the target matching the selected operation. This is speculative fan-out: evaluate possible branches in one request and discard answers that are irrelevant to the chosen branch. It saves a subsequent network round trip. [19]

Observed page controls feed parallel operation and target decisions, with a text-model branch only when typing is needed, followed by checked execution and renewed observation
Figure 4. A reconstruction of the Jev Ultrafast action loop. The executor turns model choices into actions. This is neither a neural-architecture diagram nor a performance measurement.

For TYPE_TEXT, a separate language model generates the field value. The recorded run used inception/mercury-2.5 to produce Zurich and London. Jev did not generate those strings. The executor checks whether the selected target remains valid and re-observes stale pages. A model-selected DONE is not independent evidence of task success. Agent loop [20]

Recorded itemValue and boundary
TaskShow one-way Zürich → London flight results for September 20, 2026. No flight selection, booking or payment.
Recorded elapsed time7.073 s, from first prediction after initial observation to final DONE. Browser setup, initial navigation and independent final verification are excluded.
Included work17 Jev requests, two text-model calls, ten browser interactions and one explicit WAIT, including discarded stale decisions and results loading.
Usage90,558 Jev input tokens; $0.00006272 reported across the two text-model calls.

These values come from the recorded JSON [21] and measurement description [22]. Applying the published input price gives 90,558 × $0.042 / 1,000,000 = $0.003803436 for Jev. Adding the text helper gives about $0.003866, or $0.0039 rounded. The Jev dollar component is an estimate because its responses contain tokens without a billed dollar amount. Browser, hosting and development costs are excluded. This review made no such calls or payments.

A runtime comparison retained both models and alternated three runs per implementation. Median time fell from 9.450 s to 7.092 s, about 25%, while median browser protocol calls fell from 1,092 to 101. Both arms passed 3/3 checks. [23] This small experiment concerns surrounding code on one task. It neither compares Jev with another decision model nor estimates reliability across websites.

The transferable design is to construct eligible actions from current evidence, select among them using semantic judgment, and check execution conditions and outcomes in code. Faster inference cannot repair incorrect candidate construction or stale state.

What the driving and trading demos establish

The driving link leads to JevPilot [24], a Three.js simulator. Software samples steering-and-speed candidates and passes compact road, traffic and destination tables to Jev. Its README describes up to four requests per second near traffic or turns and roughly 1.5 on clear roads. These are author-described settings, not throughput measured by this review.

The request builder [25] resolves single-option decisions locally. The server [26] validates state and selections and rejects moving candidates with predicted imminent collisions. Detailed geometry, control and collision computations remain in software. This demonstrates bounded decisions inside a simulated control loop; it does not evaluate camera perception or real-road autonomous-driving safety.

The public Jev Trader page [27] describes choosing a bid or ask on Kuru's MON/USDC book every block. At inspection it displayed dry run and was waiting for a connection. No evidence sufficient to verify live fills, profit or fees was obtained. Fast direction selection and profitable trading require different evaluations.

The post's 20–200-times speed and 40–400-times cost ranges are not adopted as universal guarantees. Vendor workflows, the LangChain repeated-judge experiment and the flight runtime comparison use different baselines and units. Free output is also a current pricing policy: input processing and decisions still require computation, and any accompanying generation model adds its own charges.

Compare against classifiers as well as LLMs

Classification without text generation is well established. Encoder models such as ModernBERT provide relevant efficient baselines. Returning a category alone does not establish a novel learning principle. Jev should be assessed as a combination of a question-defined interface, probability outputs, parallel evaluation and a managed service. [13]

ApproachSuitable conditionsCosts to compare
Rules and parsersExplicit formats, arithmetic, dates and permissionsMaintaining exceptions
Encoder classifiersStable categories and labeled examplesLabeling, training and infrastructure
Structured-output LLMsJudgment combined with reasoning, explanation or generationInput/output length, reasoning and retries
JevFrequent focused semantic decisions with bounded answersService calls, question design, calibration and escalation

This table offers selection criteria, not measured rankings. LangChain's TypeSafeClassifier integration illustrates uses such as model routing. Retaining the generation model while replacing selected decisions is a natural experiment. [14]

Schema compliance leaves semantic errors unresolved

A constrained answer space prevents inventing an option outside that space. It does not prevent choosing the wrong permitted answer. This review therefore separates schema compliance from correctness when interpreting the company's “no hallucinations” language.

The documented Jev 1.13 weaknesses include arithmetic, counting, date comparison, indirect references, irrelevant long context and adversarial input. Independently asking a proposition and its negation does not guarantee probabilities summing to one. When the intended events are true complements, compute 1 − p in code instead of expecting separate model responses to enforce probability identities. [15]

Permissions and objective success criteria should remain explicit. File existence, exit codes, test results and numeric tolerances belong in code. The proposed role for Jev is interpreting unstructured descriptions. This architecture is a reviewer recommendation, not evidence that Jev is a validated security control.

First experiments for research and coding harnesses

One candidate is screening public papers: does an abstract report measured device lifetime, or does the supplied passage directly support a particular claim? Keep the model that explains the paper while comparing the cost of repeated selection decisions.

Another is categorizing execution logs. A parser checks the exit code; Jev classifies unstructured error descriptions; code determines whether to retry, stop or call another model. In density functional theory (DFT) workflows, convergence tolerances, energy differences and forces remain numerical tasks. There is no evidence here that Jev replaces electronic-structure calculations or a validated property predictor.

The documented fixed model is jev-1.13.0, priced at $0.042 per million input tokens with no output charge. The whole request permits 64k tokens, while state plus its longest question permits 32k. English is the primary training language; Korean needs separate evaluation. The consulted documentation did not establish deployable on-premises weights. [16]

This illustrative request screens a public abstract. It has not been sent to the API. The instructions and criteria carry the intended meaning.

{
  "model": "jev-1.13.0",
  "state": {"abstract": "Public abstract to evaluate"},
  "questions": {
    "reports_device_lifetime": {
      "type": "noul",
      "instructions": "Does the abstract explicitly report measured OLED device lifetime?",
      "criteria": {
        "true": "A measured device lifetime result is explicitly stated.",
        "false": "No such result is stated, or only a calculated proxy is given."
      }
    }
  }
}

Establish human labels for representative cases and compare rules, a small classifier, the current LLM and Jev on the same inputs. Record false positives, missed cases, observed accuracy by probability bin, abstention, median and 95th-percentile latency, and costs including retries and escalation. Separate repeated judgments of identical cases from evaluation on new independent cases.

The operational metric is cost per completed task meeting the required quality. Cheap classification can lose its advantage when errors create expensive rework. The proposed starting point is observation without changing execution, followed by limited deployment after error types and consequences are understood.

Related discussion and reading

Sergii Shcherbak's public LinkedIn introduction includes a Doom demonstration and links to the vendor announcement. Social content is treated as a discovery source; technical judgments above rely on primary documentation and the public experiment.

Authorship, AI assistance & verification

How this review was made

Responsible editor
Hyun-Jung Kim
AI system
OpenAI Codex Work Mode; exact model identifier not retained
Verifiable agent roles
Codex — primary-source review, bilingual writing, scientific copyediting, original figures, aggregate arithmetic and publication checks
Editorial harness
AI Tech Review Editorial Harness v2026.08 · public method
Verification scope
TypeSafe documentation and release, LangChain experiment and public repository pinned at adfea74905f721ea2594e22804c8c8edf1693163; reported means, public JSON aggregate variance ratios and five-case/100-repeat design; no live API or benchmark reproduction; shared LinkedIn post; Jev Ultrafast 1231850 and JevPilot e1beeb1 documentation/code; flight cost arithmetic; trader public page only, no demo execution; bilingual HTML, generated hero, deterministic SVGs, metadata and local assets; one agent with no independent verification agent
Human review record
topic and conditional publication requested; no separate line-by-line human review
Evidence cutoff
2026-09-21

A technical review of official documentation and an early experiment, without new API measurements. Code computes reported-value ratios and figures; harness applications are reviewer proposals. scientific-stop-slop-ko publication copyediting applied. Exact authoring-model identifier not retained.

This public HTML includes the article, figures, and public external references. Private working notes and message metadata are not published.