Skip to main content

Looking Under the Hood of webAI Frontline

August 11, 2026

If you've read our other blog posts about webAI Frontline, some part of you might still be thinking: this shouldn't be possible.

Good instinct. For years, it wasn't.

Here's the math that stopped everyone. Take one aircraft type's complete manual set, about 34,000 pages, and prepare it for intelligent retrieval, the kind where a question gets matched against the meaning of every page. The search index alone used to demand roughly 17 gigabytes of memory. An iPad has 12 to 16. Total. For everything.

The index for one manual set was bigger than the entire device. That single equation is why the vast majority of document AI tools live on servers.

This post is about what changed, what's actually running on the device, and what it takes to get one of your own. Because "we built you an expert" is a big claim, and a claim that size should have to show its work.

It reads the page the way you do

Technical documentation isn't mostly text. It's wiring schematics. Exploded parts views. Torque tables. Charts where the value you need is a cell in row nine. A tool that extracts the words and discards everything else has thrown away half your manual, and arguably the half you walk to the terminal for.

Frontline doesn't extract the text. It reads the page, as a page, the way your eyes do. Every figure, every table, every diagram gets embedded and understood alongside the words around it. The retrieval model doing this, ColVec1.1, was built by webAI and ranks first in the world on ViDoRe V3, the benchmark for visual document retrieval, ahead of Nvidia's best embedding model.

So when the answer is a chart, it finds the chart. When it's a schematic, it surfaces the schematic, cited to its page, in front of you. Keyword search cannot see any of this. It never could. That's not an incremental gap between us and CMD-F. It's a category difference.

It specializes, on purpose

Ask what makes a person an expert and it's never "they know everything." It's depth in a domain.

Frontline is built the same way, deliberately. Each Collection is scoped to one domain: one aircraft type, one equipment family, one regulatory regime. Behind it sits a set of documents it answers from, and a Collection maps to a task, not a job title. The person doing the work carries several and switches as the work changes, in under half a second.

You don't hire one generalist who claims to know the whole building. You build a bench of specialists. Focused beats broad, in people and in AI: a tightly scoped expert answers better, drifts less, and is easier to trust, because you know exactly what it read.

It understands, then it reconciles

Here's what happens in the second after you ask.

Your question gets matched against the meaning of every page in the Collection, not the keywords. Ask "how tight" and it finds the torque section that never uses the word "tight." Ask the way you'd ask a colleague, in your own words, out loud if your hands are full.

Then the part a search box has never done for you: when the answer lives in more than one place, a value here, a condition there, a note two documents over, it reads all of them and comes back with one answer. Not a list of places to go look. An answer.

It shows its work

Every answer arrives with the exact source page attached. Tap it, the manual opens, the section is right there. You verify the way you'd verify a colleague's answer: by looking at the source they pointed to.

This is the entire trust model, and it's why the system works well in regulated environments. You're never asked to believe the AI. You're handed the page.

For organizations that go one step further, where policy says generated text can't be shown at all, there's a citations-only mode: same understanding, same retrieval, and the output is the cited source pages themselves. Some of the most regulated maintenance operations in the world start exactly there.

The engineering that made it fit

Back to the math from the top. The full search index is ~17 GB, which is too large to keep entirely in RAM on a 12–16 GB iPad. So webAI built a two-tier memory system that keeps only the actively needed parts in memory and loads the rest as needed. This lets the iPad search the full corpus without shrinking or degrading the index, while still achieving sub second retrieval latency.

The numbers

Claims like these get tested, so we had them tested.

On a public 5,244-page set of Air Force Technical Orders, across 166 evaluation questions, independently validated: Frontline achieved 88.9% answer accuracy. Every cited page surfaced in about a second. The entire run, all 166 questions, happened on one iPad, fully offline, and only used about half a battery.

And one more, from the field rather than the lab: an airline maintenance operation measured up to 16 hours of documentation search inside a single engine change. With the full manual set on the device, at the aircraft, that work fell to as little as 20 minutes.

It runs on the iPads you probably already own

Any M-series iPad runs Frontline. If your organization deployed iPads in the last few years, there's a decent chance the hardware conversation is already over.

For the best experience, we recommend an iPad Pro, M4 or M5, with 12 gigabytes of RAM or more, and storage sized to how many Collections you want on the device. If you're holding an older fleet, that's a pilot, not a blocker: we validate on your hardware before you commit to anything.

Yes, that's a consumer device you can buy today, at a store, carrying your organization's entire technical library. No server racks. No integration project. No infrastructure line item.

What step one actually looks like

Smaller than you're imagining.

The first pilot can be as simple as a handful of Collections and a team of 10-15 people. You pick the group and the documents. We build it and put it on your iPads.

Where the build runs is your call. Our infrastructure, your local Mac Studios, or your cloud, on whatever provider you already use. That portability makes sure that a pilot doesn't wait on procurement.

From there you set the pace, from ten users to thousands. None of it changes what your technicians touch. The asking and the answering happen on the iPad, offline.

The expert gets built the way any expert gets built. It reads everything, then it works alongside your best people until it's earned the floor.

See it against your own documentation. Request a demo. →