Skip to main content

Meet TwIL-LM3-Pro, webAI's Next Step in Local Reasoning

September 30, 2026

Our first-generation TwIL models were downloaded more than half a million times since they were released. Thank you to everyone who downloaded them, put them through their own tests, and shared the work.

Today, we’re releasing TwIL-LM3-Pro. It delivers a 31% relative improvement over the original TwIL-LM3 on our headline formal-logic composite, at just 3.66 billion parameters. Quantized builds run locally on everyday computers. No cloud required.

With the first release, we wanted to show how much reasoning we could put into a model small enough to run where the work happens. Pro takes that work further. It improves the aggregate formal-logic score and the broader benchmark average while keeping the model compact enough for local deployment.

Selected results from webAI’s evaluation. Full results and evaluation details are available in the model card.

For builders, the question is straightforward: can this model do the job, on the hardware you have, with the data you need to keep close? That is the question we keep working on.

What improved

TwIL focuses on formal reasoning: checking whether a conclusion follows from its premises, working out rules from examples, and translating language into representations that other software can use. Those skills matter when you’re building around conditions, exceptions, and decisions that need to be checked.

On our formal-logic suite, Pro’s headline composite rises from 42.2 to 55.4 on a 100-point scale. That is the 31% relative gain over the original model. The strict seven-task composite improves by 46% relative, from 19.7 to 28.8.

Formal tasks use thinking mode and 200 examples per task. F1, accuracy, and composite scores measure different things; they should be read by row.

The improvements show up in individual tasks, too. Strict multiple-choice accuracy moves from 11% to 41%. Entailment accuracy rises from 57.5% to 67%. Lean proof critique goes from 66% to 79.5%.

Pro also leads the public VibeThinker-3B checkpoint on all six reported formal-task scores. Its headline composite is roughly 35% higher than VibeThinker’s, 24% higher than Qwen3.5-4B’s, and 47% higher than LFM2.5-8B-A1B’s in this evaluation.

There are tradeoffs. The original TwIL-LM3 still scores higher on Lean formalization and slightly higher on semantic parsing. Pro’s small headline lead over Qwen3-8B is within sampling noise. We’re publishing the full table so you can choose based on the task you’re building for.

Reasoning beyond the formal suite

A specialist still needs to handle the surrounding work. We tested Pro on a separate set of broader benchmarks to see how it performs outside the formal tasks.

The ten-dataset average rises from 73.4 for the original TwIL-LM3 to 79.0 for Pro. On SVAMP, a math word-problem benchmark, Pro reaches 95%. On MuSR, which tests multistep reasoning, it reaches 64.1%. Both are the highest scores among the models shown in the table below.

Broader benchmark results. Some scores are limited by truncated generations, as noted in the chart. Track B uses thinking mode off for TwIL.

On BIG-Bench Hard’s logic subset, Pro scores 95.4%, compared with 66.3% for the original TwIL-LM3 and 61.1% for public VibeThinker-3B. On GSM8K, it reaches 94.3%, close to the larger Qwen3-8B at 95.7%.

These results don’t make Pro the best model for every kind of reasoning. VibeThinker leads the ten-dataset average, and Qwen3-8B is stronger across several broader tasks. What they show is meaningful progress for a compact model with a specific job to do.

What speed means for a builder

A reasoning model can generate tokens quickly and still take a long time to answer. The length of the reasoning matters. For a builder, the useful measurement is how quickly the system returns an answer you can work with.

Derived answer-throughput estimates from server testing. These are approximate comparisons, not measurements of response time on a personal device.

On the broader suite, Pro averages about 792 output tokens, compared with roughly 1,789 for VibeThinker-3B. Combining average answer length with our separate decode test gives an estimated throughput of 26.7 answers per second for Pro, versus 15.7 for VibeThinker and 4.9 for Qwen3-8B.

These are approximate server comparisons, with different engine versions across some runs and two GPUs for the 120B model. Measure response time on your own hardware.

Pro reasons longer than the original TwIL-LM3, which remains faster by this measure. Roughly one quarter of Pro’s formal-task generations hit the length cap, so give it enough room to finish.

Run it with your data

The recommended Q4 GGUF build is a 2.09 GiB weight file. It can run through llama.cpp on CPU or suitable local GPU hardware; full weights are also available for the Transformers ecosystem. Published benchmark scores were measured on BF16 weights, so we have not yet quantified the accuracy difference for the quantized builds.

Local deployment gives you control over where inference happens. You can work with private data without sending it to an external model service. For formal work where a wrong answer has consequences, pair the model with a verifier or symbolic solver and check the output before acting on it.

Why we keep investing in post-training

Pro starts from IBM’s Granite 4.2 3B foundation. We adapted it through supervised fine-tuning, diverse reasoning traces, checkpoint fusion, weight interpolation, and contrastive reinforcement learning with programmatic feedback.

The point of that work is to teach a model useful expertise while retaining enough of its broader capability to handle the rest of the task. Against its own Granite base, Pro improves the formal-logic composite by 28% relative while keeping the broader ten-dataset average approximately level.

We believe AI is entering a post-training era. Strong foundation models give builders a starting point. The work after that increasingly determines what a model is good at, how reliably it behaves, and whether it fits the system you’re trying to build.

The advantage will belong to teams that can produce capable, personalized intelligence faster and more efficiently, then deliver it on devices people already own. That is what we’re building at webAI.

TwIL is one part of that effort. We want builders to have useful experts they can run close to their data, test for themselves, and combine with other models and tools. Releasing the weights and the evaluation details gives you a way to judge the work on your own terms.

What comes next

We are only beginning to share what’s coming out of our lab. Coming soon: Meridian, our upcoming family of models built for on-device intelligence.

Our most advanced models will be available through the webAI application. Join the waitlist as we continue rolling out general availability.

For now, download TwIL-LM3-Pro, read the model card, and try it on a problem you care about. We’re interested in where it helps and where it still needs work.

Proudly built in Austin, Texas.

Availability and evaluation

TwIL-LM3-Pro is released under the webAI Non-Commercial License v1.0. The repository includes full weights and GGUF builds. The model card documents the evaluation settings, task metrics, and limitations. Formal-task results use 200 examples per objective; broader results use 300 examples per task. Comparisons describe these tests, rather than a universal model ranking.

Model and evaluation details: https://huggingface.co/webAI-Official/TwIL-LM3-Pro