Company News

Company News

Company News

Evaluating Jev for answering and reviewing clinical research

Evaluating Jev for answering and reviewing clinical research

Evaluating Jev for answering and reviewing clinical research

We evaluated Jev across evidence answering, abstract screening, and full-text screening to see where its probability estimates can help clinical and regulatory research workflows.

Of the three tasks we tested, answering evidence questions is where Jev's probability estimates helped most: they decided which questions were worth sending to a costlier model.

At Log10’s Applied AI group, we are constantly studying new models and assessing their relevance for clinical and regulatory workloads. Despite the recent excitement about Jev, our initial tests suggest a limited role for it in these workflows. Below are our findings from tests on a subset of ClinReg.

We tested whether Jev, a low-cost model that returns probabilities, could answer clinical research questions, check other models’ answers, and select answers for review by a stronger model. We chose Claude Opus for that role. Our tests covered evidence answering, abstract screening, and full-text screening.

The clearest benefit came when Jev answered evidence questions first and Opus supplied new answers to the questions Jev was least confident about (Configuration 1 below). Jev was unreliable as the sole judge of other models’ answers on all three tasks. Using its judgments to select answers for Opus led to more corrections than random selection on abstract and full-text screening, but not on evidence answering overall. The full-text finding is exploratory; the abstract-screening test did not establish an advantage over the original model’s own confidence.

What is Jev?

Jev is TypeSafe AI’s model for structured decisions. Given context and a question, it returns probabilities in one of three formats:

  • Choice: probabilities for a list of answer options. We used the highest-probability option as the answer to evidence questions.

  • Noul: the estimated probability that a yes/no answer is “yes.” We used it to ask whether a study should be included or another model’s answer was correct.

  • Score: a numerical rating against ordered levels you define, with probabilities and confidence. We did not evaluate this format.

These probabilities are Jev’s estimates. We tested whether ranking answers by them helps select a fixed share for review more effectively than random selection.

Three tasks and their datasets


Task

Dataset

Input

Required answer

Evidence answering

TrialPanorama

Clinical question, supporting study evidence, and answer options

Choose the best-supported option

Abstract screening

TrialPanorama

Review eligibility criteria, study title, and abstract

Include or exclude

Full-text screening

CLEF eHealth TAR, topic CD009593

Review eligibility criteria and a complete paper

Include or exclude


Screening criteria specify which studies belong in a review, including the patient group, treatment or test, and study design. The full-text dataset contains 47 papers for one tuberculosis diagnostic-test review: 20 should be included and 27 excluded. The TrialPanorama tasks cover multiple clinical questions and reviews and are part of our ClinReg evaluations.

We checked outputs against reference answers never shown to the models. This includes Opus: it can introduce errors as well as fix them. The real responses below illustrate particular outcomes; they are not a random sample. Follow-up tests used published training splits, so “previously unused” means new to our experiments, not necessarily unseen during model training.

Three model configurations

We tested each configuration on all three tasks:

Configuration

How it works

Practical benefit we wanted to test

1. Jev answers; Opus handles uncertain cases

Jev answers every question. Opus supplies replacement answers for the lowest-confidence 20%.

Lower overall answering cost. Use inexpensive Jev answers for most questions and reserve Opus for uncertain cases, rather than paying for Opus on every question.

2. Jev judges another model’s answer

Jev labels existing answers as right or wrong. We measure whether those judgments are correct.

Evaluation on a smaller budget. Test whether Jev can check existing answers at lower cost than having experts or a stronger model evaluate every answer.

3. Jev selects another model’s answers for Opus to review

Jev estimates whether each existing answer is correct. Opus reviews or replaces the 20% Jev considers most likely wrong.

More effective use of a limited review budget. When Opus can review only a fraction of answers, test whether Jev can select answers that lead to more corrections than random review.


In configuration 1, Jev supplies the original answer. In configurations 2 and 3, it evaluates an answer another model has already produced. The sections below distinguish follow-up tests on previously unused data from exploratory analyses of existing results.

Configuration 1: Jev answers and Opus handles uncertain cases

Jev answers every question. We keep its most confident 80% of answers and replace the remaining 20% with Opus’s answers. We compare this with Jev alone, Opus alone, and sending the same number of randomly selected questions to Opus.

For evidence questions, confidence is the probability Jev assigns to its chosen answer. For yes/no screening decisions, uncertainty is greatest near 50%.

Evidence answering

We fixed the selection rule after an initial analysis, then tested it on 1,000 evidence questions not used in that analysis. Both models answered every question so we could compare all four approaches on the same data. Jev used Choice with four answer options.

Confident and correct: On a thyroid-cancer test question, Jev chose A with probability 1.00 and assigned zero to the other options. The reference answer was A, so the workflow kept a correct answer.

Uncertain and corrected: On an in vitro fertilization question, Jev returned:

{

"choice": "C",

"probabilities": {"A": 0.33, "B": 0.29, "C": 0.38, "D": 0},

"confidence": 0.17

}

The reference answer was B. Jev’s probability of 0.38 on its chosen answer put this question among the lowest-confidence 200. Opus chose B, correcting the mistake. Selection used the chosen option’s probability, not the separate API confidence field. We selected the lowest-confidence fifth; 0.38 was not a fixed cutoff.

Accuracy rose from 80.0% with Jev alone to 87.9% with confidence-based selection, compared with 83.0% for random selection. The advantage over random was 4.9 percentage points (95% confidence interval: 3.6-6.2). Opus corrected 87 errors and introduced 8 new ones, leaving 79 fewer errors across the 1,000 questions.

Opus alone reached 94.9%. Using it for the selected 20% captured about half its accuracy improvement over Jev, at about one-fifth of the cost: $5.46 per 1,000 questions, versus $25.33 for Opus alone. These are estimates for the calls each workflow would need, not the cost of evaluating every alternative. We did not measure live response times or human time savings.

Abstract screening

Jev received each study’s title and abstract alongside the review’s eligibility criteria. We used Noul to ask: “Should this study be included?” A probability of at least 50% meant include; below 50% meant exclude.

We replaced the 20% of decisions closest to 50% with Opus’s decisions. These were Jev’s most uncertain answers: a 48% inclusion probability is almost undecided, whereas 5% strongly favors exclusion.

In an analysis of existing responses, F1 increased from 70.7% for Jev alone to 74.9% with uncertainty-based selection, compared with 72.8% when Opus replaced a random 20%. The advantage over random selection was 2.1 points (95% confidence interval: 0.3-3.7).

F1 combines the share of included studies that are eligible with the share of all eligible studies retained. It can improve by removing ineligible studies, even without recovering more missed eligible studies. We did not establish that uncertainty-based selection reduced missed eligible studies more than random selection did. This result has not been confirmed on previously unused reviews.

One real response illustrates which decisions our selection rule would prioritize. Jev screened seven studies together, answering a separate inclusion question for each. All seven should have been included:

Study

Probability of inclusion

Jev’s decision

Reference

study_1

0.97

Include

Include

study_2

0.05

Exclude

Include

study_3

0.09

Exclude

Include

study_4

0.11

Exclude

Include

study_5

0.75

Include

Include

study_6

0.48

Exclude

Include

study_7

0.93

Include

Include


Jev excluded four eligible studies. Study_6, with an inclusion probability of 0.48, would receive higher review priority than studies 2-4, whose probabilities strongly favored exclusion. Selection depends on Jev’s uncertainty; the actual mistakes are identified by comparison with the reference labels.

Full-text screening

We applied the same approach to 47 full papers from CLEF eHealth TAR, using Claude Opus 5.5 as the second model. Noul determined inclusion at the same 50% cutoff. We sent the nine papers closest to 50% to Opus, rounding the 20% budget to a whole number.

Accuracy rose from 39 of 47 (83.0%) to 43 of 47 (91.5%), versus 85.4% on average with random selection and 95.7% with Opus alone. All four approaches retained all 20 eligible papers. Opus corrected four false inclusions and introduced no new errors in the selected group.

This was an exploratory test on one 47-paper review; the observed gain does not establish how well the approach would transfer to other reviews. The estimated workflow cost was $0.64 for the 47 papers, versus $3.39 for Opus alone.

Each call also asked a Choice question: select INCLUDE or one of 13 exclusion reasons. These were two questions in the same call; Choice was not a second-stage check. Noul remained the primary decision.

Paper (short description)

Noul probability of inclusion and decision

Choice option and probability

Ref.

Rapid molecular detection of tuberculosis and rifampin resistance

0.96: Include

INCLUDE: 0.99

Include

Urine-based Xpert MTB/RIF in hospitalized patients with HIV

0.51: Include

Ex_SPECIMEN: 0.79

Excl.

Xpert MTB/RIF in a tuberculosis prevalence survey

0.93: Include

INCLUDE: 0.99

Excl.


The first paper shows agreement on a correct inclusion. For the second, Choice identified an unsuitable specimen: the study used urine, while the review required respiratory specimens. Choice assigned only 0.17 to INCLUDE, yet Noul’s 0.51 led to inclusion. We compared formats after the run; this does not establish that Choice is generally more accurate.

For the third paper, both formats were confident and wrong. It remained outside the nine papers selected for Opus. A second model cannot correct an error it never reviews.

Configuration 2: Jev judges another model’s answer

The model that originally produced an answer is its source model. For abstract screening and evidence answering, Jev evaluated saved answers from 26 models in our TrialPanorama benchmark runs, including Claude Opus 5, GPT-6 Astra, Gemma 4 31B, and Kimi K2.6. The full-text analysis used 12 source models.

Jev received the original task inputs and one proposed answer, with the source model’s identity and confidence hidden. We asked Noul: “Is this answer correct?” A probability below 50% meant “wrong”; at least 50% meant “correct.”


We checked these judgments against the reference answers. Judgment accuracy is the share of Jev’s right/wrong labels that are correct. For example, if Jev assigns 10% probability to a correct answer being correct, its “wrong” judgment is an error. There is no Opus correction stage here.

The abstract screening test covered 90 reviews; the evidence test covered 24,698 proposed answers to 950 questions. These reviews and questions were not used to develop the judging prompt. The exploratory full-text analysis reused the 47-paper review: its 12 source models supplied 563 valid decisions, of which 74 (13.1%) were wrong.


On all three tasks, Jev’s judgment accuracy was below that of labeling every answer correct. That simple rule scores well because most source answers are right, although it cannot identify any errors.

Most answers Jev flagged were also correct: only 36.7% of its abstract-screening flags, 34.4% of full-text flags, and 11.5% of evidence-answering flags were real errors. For evidence answering, that was close to the overall error rate of 11.0%.

Full-text flags nevertheless contained a higher share of errors than the full set: 34.4% versus 13.1%. This distinction matters. Jev can concentrate errors in a subset without reliably judging every answer. Configuration 3 tests whether that subset is useful for Opus.

Configuration 3: Jev selects answers for Opus to review

Another model answers first. Jev receives that answer and the original task inputs, then estimates whether the answer is correct. We select the 20% with the lowest estimated probability of correctness. An answer scored at 10% takes priority over one scored at 50%: we are selecting suspected mistakes.


We keep unselected answers unchanged and measure net corrections: errors fixed minus new errors introduced. The correction step differed across tests:

  • Abstract screening: Opus judged each selected decision. If it called the decision wrong, we switched include to exclude or vice versa.

  • Full-text screening and evidence answering: Opus’s saved answer replaced each selected answer. These exploratory analyses reused previously analyzed data.

Abstract screening

The follow-up test covered 1,000 studies across 100 previously unused reviews. Half were eligible; real screening pools may contain far fewer. GPT-5.6 Luna, Gemma 4 31B, and Gemini 3.8 Flash each supplied initial decisions. For each model, Opus reviewed 200 selected randomly, by the source model’s lowest confidence, or by Jev’s lowest correctness estimates.


Jev’s selections led to 27.7 net corrections per model, compared with 10.0 for random selection and 18.7 for the source model’s confidence. These are equally weighted averages across the three models, measured over each model’s full set of 1,000 decisions.

The advantage over random was 17.7 corrections (95% confidence interval: 7.5-27.4). Compared with source-model confidence, the gain was 9.0, but the interval ranged from 1 fewer to 19 more corrections. We therefore did not establish an advantage over using the source model’s confidence.

That matters for cost: source-model confidence was already available, whereas Jev evaluated all 1,000 decisions before selecting 200. Its observed gain over source-model confidence mainly came from excluding ineligible studies. Both methods restored similar numbers of wrongly excluded eligible studies.

Full-text screening

For each of the 12 source models in the 47-paper review, we selected the nine decisions Jev considered most likely wrong (the rounded 20% budget) and replaced them with Opus’s saved decisions. We compared this with replacing nine randomly selected decisions.

Selection method

Average net corrections per source model

Random nine decisions

0.80

Jev’s nine most-suspect decisions

2.42


Jev’s selections produced 1.62 more net corrections per model than random selection (95% confidence interval: 0.46-2.91, resampling papers). This supports a benefit within this review, but the absolute gain was modest and the analysis covered only one review. It did not compare Jev with the source models’ own confidence.

Evidence answering

This analysis reused the 950-question judging set. We selected 20% of each source model’s answers using Jev’s correctness estimates and replaced them with Opus’s saved answers. The comparison covered 25 source models; we excluded Opus 5 because replacing its answers with the same saved answers would change nothing.

Selection method

Average net corrections per source model

Random 20% of answers

13.09

Jev’s 20% most-suspect answers

13.04


Jev did not outperform random selection overall. The difference was -0.05 net corrections per model (95% confidence interval: -4.51 to 4.25).

Results varied by source model. For Gemma 4 31B, Jev’s selections led to 70 net corrections, versus 35 for random selection. For GPT-6 Astra, they introduced seven more errors than they fixed, versus 2.2 net corrections for random selection.

One possible explanation is that Jev favors answers matching its own choice, so a correct answer from a stronger model can be flagged because Jev disagrees. The selection result is consistent with that interpretation; it does not establish the cause. This analysis also did not compare selection against source-model confidence.

Where Jev may fit

Jev’s usefulness depended on both the task and its role. Its clearest benefit came from answering evidence questions first and using its confidence to select questions for a stronger model. Using Jev to judge other models’ evidence answers and select replacements did not outperform random selection overall.

For screening, Jev’s judgments selected decisions that led to more corrections than random selection on both abstracts and full texts. The abstract result came from a follow-up test; the full-text result was exploratory and limited to one review. An advantage over the source model’s own confidence remains unproven for abstracts and was not tested for full texts.

Jev was unreliable as the sole quality check on all three tasks. A study can mention the right disease but involve the wrong patient group or study design. Checking that decision, or whether an answer is supported by clinical evidence, can require much of the same interpretation as answering the original question. Our tests do not establish that task difficulty caused the judging failures.

Jev may fit narrower decisions with explicit criteria, such as routing a support ticket to billing or technical support, identifying an explicit refund request, or checking whether a document contains a required field. These are plausible applications, not tasks we evaluated here.

These tests support selective use of Jev’s probabilities, with performance checked for the specific task and role.

Ready to see Everest in action?

Discover how Everest delivers accuracy, automation, and compliance other tools can’t match.

Schedule a demo

Ready to see Everest in action?

Discover how Everest delivers accuracy, automation, and compliance other tools can’t match.

Schedule a demo