Finance

Document AI in Finance: Expanding a multimodal ensemble model for the automated classification of real estate documents

How an IT service provider in the financial sector successfully expanded its existing AI document processing system to include new document classes and consolidated various models into a unified, scalable architecture.
Person at a desk sorting real estate documents into a file (AI-generated)

Initial Situation

A leading IT service provider in the financial sector operates an AI-powered system for the automatic classification of real estate financing documents. The existing system was already capable of reliably identifying and processing around 30 different document classes. However, with the steady increase in the volume of documents uploaded by case workers, a strategic need arose to expand the system to include additional document classes essential to real estate financing.

Without automatic recognition for these specific classes, the responsible staff had to manually identify and assign these documents, which slowed down the processing workflow and tied up valuable resources. The challenge lay not only in integrating the new classes but also in the strategic consolidation of various models developed in parallel into a single, high-performance overall model. Furthermore, development had to take place under strict regulatory requirements and within a restrictive on-premise data center infrastructure, placing the highest demands on data security and the development process.

Our Solution

To meet the requirements for scalability and precision, the existing multimodal ensemble model was expanded to include the new document classes and retrained. The focus was on a robust architecture that processes both text and image information and can be flexibly extended to include new document classes.

  • Multimodal Ensemble Architecture: The technical foundation is the established ensemble model, which combines text and image information. A natural language processing component handles the text extracted from each individual document page via OCR, while a Computer Vision component simultaneously analyzes the visual appearance of the page. The results from both models are fused at the logit level to reach a precise final classification decision that takes both content-related and structural features into account.
  • Structured Data Preparation and Labeling: Due to the varying quality of historical financial documents, systematic data preparation was essential. In an iterative process, representative datasets for the new document classes were analyzed and manually annotated. By creating precise label guides and using professional annotation tools, we ensured high data quality, which served as a reliable foundation for the subsequent model training.
  • Expansion and Consolidation in Model Training: In the first step, the ensemble model was retrained with the two new document classes. Subsequently, the task was to merge several models that had been developed in parallel along with their datasets. These were consolidated into a common master dataset, which served as the basis for training the model from scratch once again. Systematic experiment tracking made the training runs comparable and transparent. The resulting overall model maintains the recognition performance for the previous 30 or so document types.

The final solution was deployed as a Docker-based service on-premise in a Kubernetes-cluster in the customer's data center and integrated into the existing classification service.

Results & Business Impact

By expanding and consolidating the classification model, automated document processing now covers a larger portion of incoming paperwork. The new document classes are classified fully automatically and reliably. This closes a gap in the digital processing workflow for real estate financing.

For case managers, uploaded real estate documents are now pre-classified and can be assigned to the corresponding files without the need to manually determine the document type beforehand. At the same time, merging the various models into a unified architecture has improved system maintainability and created a solid foundation for integrating additional document classes in the future, without compromising performance for existing classes.

Table of contents
Branche
Finance
Thema
Computer Vision
Deep Learning
Document AI
Natural language processing
Dauer
10 months
Share now
Link kopiert!
Dr. Jürgen Stumpp

Ready for the next step?

Let's find out together how we can support your business with customized AI.

Dr. Jürgen Stumpp
Managing Partner | AI Strategy Consultant
+49 170 9374842
Contact us