IT Brief UK - Technology news for CIOs & IT decision-makers
United Kingdom
Exclusive: ABBYY brings zero-shot document AI to Vantage

Exclusive: ABBYY brings zero-shot document AI to Vantage

Wed, 16th Sep 2026 (Today)
Sean Mitchell
SEAN MITCHELL Publisher

ABBYY is adding zero-shot document extraction to its Vantage platform as the company develops a modular, multi-model architecture designed to move enterprise artificial intelligence projects beyond pilot deployments.

Zero-shot release

"Just this week, we're releasing zero-shot capabilities and more with Phoenix Plus inside our Vantage platform. That's being delivered today to the market," said Max Vermeir, Vice President, AI Strategy, ABBYY.

The capability is intended to extract information from previously unseen document types without requiring customers to create templates or assemble labelled training data for each format. It forms part of Phoenix, ABBYY's portfolio of models optimised for document processing.

Phoenix Core incorporates computer vision, image enhancement and technologies used to identify the structure of a document. Phoenix Plus adds generative capabilities intended to interpret the information within that structure, including zero-shot and few-shot extraction, classification, question answering and data enrichment.

ABBYY is integrating the portfolio with orchestration and model-routing functions that can select different technologies for particular document-processing tasks. Customers can use ABBYY's models or connect models already approved and deployed within their own infrastructure through its bring-your-own-model option.

The company is also working towards making functions from across its product portfolio available as individual components. This would allow customers to deploy selected capabilities in different products and environments rather than adopt an entire platform configuration.

That work includes adding components to FlexiCapture, ABBYY's established document capture product, and revising how FineReader Engine can be delivered in different environments. The roadmap also covers assisted manual-review agents, workflow enhancements and further enrichment features for Vantage.

"I don't think that's actually the question any more: is AI already in the enterprise? It absolutely is. The question that everybody is asking is: is it actually working? Is it actually producing something functional that helps your organisation, and is not just an impressive demo? Because that's a huge difference," added Vermeir.

Model routing

"It's not one model that rules them all. Despite what the more excitable corners of the market would like you to believe, you need a combination, a portfolio of different technologies. Generative AI models matter because documents are messy. They have ambiguity, extensive context and reasoning requirements. But the governance layer is, for me, the most important one. Reliability, auditability and cost control make the difference between a successful pilot and something that truly runs in operational production and delivers a return on investment," said Vermeir.

ABBYY's approach combines deterministic technology, machine learning and probabilistic generative models. Its orchestration layer is designed to route each task to an appropriate model while applying common validation rules, policies and oversight.

This architecture addresses a practical concern for companies operating high-volume processes: using a large general-purpose model for every document and task can produce variable results and unpredictable token costs. Specialised models can instead handle narrowly defined functions, while generative systems are reserved for work that requires interpretation or reasoning.

The platform includes human-in-the-loop review for exceptions and maintains traceability across the processing workflow. ABBYY is positioning those controls as necessary for regulated or operationally sensitive uses, where organisations need to determine how information was extracted and why an automated decision was made.

Vermeir said operational systems require consistent results from the same inputs, integration with existing software and predictable costs. The underlying models may already have sufficient capabilities for many enterprise tasks, but their usefulness depends on access to data that accurately represents how an organisation works.

ABBYY said it has more than 150 specialised models covering different document types. Its wider portfolio is intended to support combinations of smaller task-specific models, machine learning, language models and symbolic reasoning instead of relying on a single system.

Document context

"When I say that enterprises run on documents, it's not nostalgia for paper; it's simply a reality," said Vermeir.

Contracts, invoices, insurance claims, compliance records and shipping documents contain much of the information used to run business processes. Although their storage formats have changed, they remain a primary means through which people and organisations record obligations, evidence, approvals and financial information.

Many companies have previously automated the systems before and after a document-processing stage while retaining manual review in the middle. ABBYY argues that this remaining step requires more than text extraction because systems must also recognise layouts, relationships, context and the meaning of the information presented.

The company describes this as a perception layer between source documents and enterprise AI systems. It converts unstructured information into standardised operational data that software and autonomous agents can use when making decisions or triggering actions.

That requirement becomes more important as organisations introduce agentic AI to plan tasks, co-ordinate workflows and take actions across business systems. An agent operating on incomplete or incorrectly interpreted document data could carry an error into subsequent systems at greater speed and scale.

ABBYY cited an insurance deployment in which document intelligence and AI-assisted workflows were used to extract claims information and validate it against systems of record. Exceptions were passed to a human-review interface, while other cases continued through the automated process. According to the company, the deployment reduced processing cycle times by more than 80%.

The company said its technology has processed more than 200 billion documents, handles billions of pages annually and operates across more than 30 industries. It supports more than 200 languages and has more than 10,000 deployments.

Open standard

"DocLang is an AI-native document standard that allows you to encapsulate all of your business information in a format that is built for AI. It gives you machine-readable information, an open standard and reliable pipelines. It has governance built in, and it also creates the opportunity to reduce token consumption by between 40% and 80% simply by transforming your data into an AI-native standard," said Vermeir.

ABBYY is participating in the DocLang initiative with IBM, NVIDIA, Red Hat and the Linux Foundation. The project is intended to establish a common representation for transferring document information into AI and agentic workflows.

The proposed context layer would sit between existing enterprise content and the models or agents using it. Its purpose is to preserve document structure and business information in a machine-readable format, reducing the need for each organisation to build separate parsers and conversion pipelines.

ABBYY plans to support DocLang exports as part of its product roadmap. The company is also extending the same document-understanding components across Vantage, FlexiCapture and FineReader Engine as it moves towards a more unified portfolio.

The standard is being developed as an open-source project under the Linux Foundation.