The growing relevance of artificial intelligence (AI) and machine learning (ML) is increasing the pressure to further optimize the efficiency of AI models, especially when it comes to inference. Inference - the time a model needs to make a prediction or take an action - is coming into focus, as it has a significant impact on the performance and reaction time of a system. Hot AI offers an opportunity to significantly reduce the response times of AI models, thereby increasing the efficiency of inference processes.
In this blog post, we use real data from a real use case to illustrate the savings of Hot AI and how Konfuzio has successfully implemented this technology.
What is container AI and how does it work?
Container AI is an approach based on Containers to provide artificial intelligence efficiently. The complete environment required by an AI application is compactly packaged in a container. Technically speaking AI Container The containers therefore contain the necessary components of an AI application and run in so-called "lightweight" environments. A container AI at Konfuzio therefore functions like a virtual workspace for an AI.
AI models from Konfuzio - in the past vs. today
Previously, the Konfuzio AI models were reloaded with each document and did not run in isolation - this was slow in a direct comparison (see comparison further down in the article). Now, with Container AI, there are three operating modes to make the AI models more flexible and faster to use, depending on requirements:
- On Demand AI - AI takes a few seconds to load the first document and switches off automatically after 10 minutes of inactivity.
- Hot AI - Container is "always-on", which eliminates loading time.
- Autoscaled AI - Hot AI with an autoscaled function is used here. The container is "always-on" and scalable, which eliminates loading time and keeps pace with the increase in data volumes.
All three operating modes offer different advantages, depending on the individual company requirements. For more details click here. The following use case focuses on Hot AI.
What is Hot AI and how does it work?
Hot AI means that these virtual workspaces (containers) always remain switched on and are immediately ready for use (always-online mode). This ensures that the AI starts immediately without losing time for start-up. This eliminates the waiting time.
Advantages of Hot AI
Hot AI not only shortens inference times. Other advantages are
- Security - Each container is a self-contained workspace that functions in strict isolation from other containers and systems. A potential error or attack within a container affects neither other processes nor the overall system. Your data remains in a completely protected virtual environment.
- Speed - With the new Hot AI, all loading times have been almost completely eliminated, allowing Konfuzio to significantly increase the performance of AI services. This leads to faster document processing and an improved user experience.
- Scalability - With an optional "auto-scaled function", Hot AI ensures that the Konfuzio AI models are not only portable, but also highly responsive in real-time scenarios. In the event of a sudden increase in demand, additional containers can also be started immediately without having to put up with the longer delays that occur when the application is first completely reloaded.
- Cost savings - The reduced loading time and more efficient use of computing resources sustainably reduce operating costs. As the AI model requires less computing time, the costs per task performed are reduced.
Use case - Konfuzio Hot AI in use
A typical use case in which these advantages became apparent comes from a customer of the Helm & Nagel GmbH,the company behind Konfuzio.
Initial situation
Since 2022, the customer has been using AI to process compliance documents. These documents are time-critical, as they are essential for meeting compliance requirements. Between 10 and 100 documents are uploaded to the Konfuzio platform every day, resulting in over 12,000 processed documents since the start of the collaboration.
Challenge
Although the customer was satisfied with the Konfuzio platform, he complained about irregular and sometimes long processing times. As the documents were placed in a general queue, this prevented immediate processing. In particular, starting and loading the AI model led to delays, as the model had to be fully loaded before it was ready for document processing. These processing times of up to 25 seconds did not match the requirements of the customer, who needed a fast and efficient processing process.
Fast processing of compliance documents is key to meeting critical compliance requirements. Delays could disrupt production and supply chains and cause serious disruptions. Document processing must therefore be available in real time to ensure smooth processes and avoid penalties.
Objective
The aim was to drastically reduce loading times by using Hot AI. The AI models were to be available immediately in "always online mode" so that documents could be processed immediately after uploading without having to wait in a queue for the AI model to load.
Solution and implementation
The Konfuzio experts implemented three parallel containers to keep the AI model permanently loaded. This approach ensured that the documents could be processed immediately after uploading without losing time for model loading. The loading time was thus reduced to almost zero. This technological solution also allowed multiple documents to be processed in parallel, increasing the efficiency of the entire process.
Result
With Hot AI, loading times are completely eliminated, so the AI model is immediately available and starts 40 to 70 times faster. In addition, Konfuzio processes individual pages three times faster and entire documents four times faster than before.
Comparison - savings in inference time
In this use case, the loading time of the AI model with Hot AI was reduced from 8 to 25 seconds to 0.099 to 0.506 seconds, with 0.099 seconds being the optimum value. This reduced the total processing time of a document to 1 to 8 seconds, resulting in significantly faster workflows and increased customer satisfaction. In direct comparison:
Before
| Distribution of the loading time (seconds) | Distribution of the total time (seconds) | ||
| Percentile | Time | percentile. | Time |
| P1 | 6,721 | P1 | 8,829 |
| P5 | 9,651 | P5 | 11,340 |
| P10 | 10,315 | P10 | 12,169 |
| P25 | 11,426 | P25 | 13,666 |
| P50 | 12,792 | P50 | 15,258 |
| P75 | 14,590 | P75 | 17,315 |
| P90 | 16,465 | P90 | 19,591 |
| P95 | 17,203 | P95 | 21,047 |
| P99 | 19,460 | P99 | 25,843 |
After
| Distribution of the loading time (seconds) | Distribution of the total time (seconds) | |||
| Percentile | Time | Percentile | Time | Improvement |
| P1 | 0,099 | P1 | 1,787 | 79,8 % |
| P5 | 0,108 | P5 | 2,021 | 82,2 % |
| P10 | 0,125 | P10 | 2,131 | 82,5 % |
| P25 | 0,173 | P25 | 2,411 | 82,4 % |
| P50 | 0,202 | P50 | 2,743 | 82,0 % |
| P75 | 0,240 | P75 | 3,382 | 80,5 % |
| P90 | 0,333 | P90 | 4,291 | 78,1 % |
| P95 | 0,404 | P95 | 5,146 | 75,6 % |
| P99 | 0,506 | P99 | 8,077 | 68,7 % |
The tables illustrate the improvement with Hot AI using percentiles that show the processing times more accurately. Before the optimization, 99 % of the documents took up to 25 seconds, while the fastest 1 % took only 8 seconds. After the introduction of the new feature, 99 % of documents process in 8 seconds or less. These changes demonstrate the significant increase in performance and efficiency of the solution.
Conclusion
As the use case illustrates, switching to Hot AI can drastically increase the speed of inference processes. Companies that process large amounts of data in real time benefit directly from reduced waiting times and faster workflows. If companies have other requirements for AI models, there are other operating modes in addition to Hot AI: either an additional autoscaled function can enable the scalability of the systems - or an on-demand function can conserve resources in the long term.
Take now Contact and our experts will advise you in detail on which Container AI is best suited to your requirements. Secure decisive competitive advantages for yourself!
