Company News

Company News

Company News

Open-weight models pulled ahead — where the work is hardest

Open-weight models pulled ahead — where the work is hardest

Open-weight models pulled ahead — where the work is hardest

We published the ClinReg Leaderboard with six new models — and found that an open-weight model scored highest on one of the most difficult tasks: TLF programming.

The three ClinReg tasks, best open-weight model vs best proprietary. Open-weight leads the hardest one — TLF programming — while proprietary still tops IND drafting, and literature screening is a tie.

A living scoreboard for clinical and regulatory AI

Today we're giving the ClinReg Leaderboard its own home at clinreg.log10.io.  Six weeks ago we introduced the benchmark behind it, with a finding that cuts against the prevailing assumption: on clinical and regulatory tasks — drafting an IND, programming the tables, listings, and figures behind a CSR, screening literature for a review — open-weight models had already caught up to the closed frontier.

The obvious question was whether that would hold as new models shipped. So rather than answer it once, we're publishing the leaderboard as a living resource — and we'll keep adding models (as well as new tasks) and updating what we find as the field moves. This is the first such update: we added six models — GPT-6 Astra, GLM-5.3, GLM-5.3 Flash, Gemini 3.8 Flash, Gemini 3.7 Flash, and Muse Spark 1.3.

Here's what changed.

Overall, it's neck-and-neck

GPT-6 Astra takes the top overall score at 89.9. But the next model on the board is open-weight: GLM-5.3 at 89.3 — six tenths of a point behind. Open-weight models now hold four of the top ten, and the whole top ten sits within five points of each other. At the very top, the open-vs-closed line no longer predicts who leads.


The complete field — all 25 models, colored by vendor. The top ten sit within five points of one another; scores fall off only further down the board.

On TLF programming, open-weight is now #1

The overall number hides the more interesting result. TLF generation is the hardest, most agentic task on the benchmark: from raw CRF data and specs, a model has to write and run the full SDTM → ADaM → TLF pipeline in a single session, debugging its own code until the output validates against the published reference.

On that task, GLM-5.3 (open-weight) is #1 outright at 88.2 — ahead of every closed model, including GPT-6 Astra at 86.2. Open-weight models take three of the top four TLF scores. On the task that most resembles the real, long-horizon work of a study programmer, open-weight didn't catch up this round. It pulled ahead.

The picture isn't uniform, and that's the point of reading the columns rather than the average. On IND drafting, a closed model still leads (GPT-5.6 Terra at 91.2), and the best open-weight model, Kimi K3, is a few points back at 86.7. Literature screening is effectively saturated: open and closed models alike cluster at the top. 

The cost gap is the real headline

Accuracy has converged. Price hasn't. Inside the top ten, cost per run ranges from $0.30 (for GLM-5.3 Flash) to $22.29 (for Opus 5) — 75× the price for a model that scores slightly lower. GLM-5.3 Flash scores 87.6 for $0.30 per run. That's frontier-adjacent accuracy for the price of a rounding error, and it's open-weight. Even the first place model overall (GPT-6 Astra) came in at $10.80 per run which was less than half the cost of Opus 5 on this benchmark. The top open-weight model (and second overall) GLM-5.3 cost $3.80 per run.


Accuracy has converged; price hasn't. GLM-5.3 Flash posts a higher overall score than Opus 5 at roughly one seventy-fifth of the cost per completed run.

For a team choosing where to run production clinical and regulatory workloads, that changes the calculation. Matching the frontier on accuracy no longer requires paying frontier prices — or sending proprietary trial data to a single closed vendor.

Methods and what's next

The scoring is built to be defensible: each task is graded against a real-world reference, judged by independent models from more than one lab rather than the model under test, with costs recorded as verified list prices per run. The full methodology — task construction, judge panels, harness settings, and per-model notes — lives in the appendix of the original post.

The full leaderboard — sortable by task and by cost — is live at clinreg.log10.io.

We plan to open-source the ClinReg benchmark tasks, and we welcome additions and feedback. Please contact ai@log10.io.

Ready to see Everest in action?

Discover how Everest delivers accuracy, automation, and compliance other tools can’t match.

Schedule a demo

Ready to see Everest in action?

Discover how Everest delivers accuracy, automation, and compliance other tools can’t match.

Schedule a demo