Event
Bavarian Data Summit 2026
Summit by Bayern Innovativ on real-time data and AI for decisions in production and mobility in Munich.
Bavarian Data Summit 2026A good data foundation for AI is the basis of every successful AI application in a company. Data largely determines how reliably an AI can work and what benefits it brings to your company. It is not about possessing as much data as possible. What is crucial is that the right data is available, current, and of sufficient quality - data quality outweighs data quantity. Specifically, are the details complete, current, consistent, and clearly named? If an item is sometimes called "lamp" and other times "lighting element," or if the price list is two years old, even the best AI cannot work reliably.
Data is the foundation on which every AI application in a company is built. Whether AI creates real value depends on the type, origin, and quality of the data used. It is equally important to first define the appropriate use case.
Andreas Gillhuber (German Data Science Society) gives a brief insight into the topic of data around AI.
Many small and medium-sized enterprises (SMEs) already possess more usable company data than they are aware of. These may be located in the ERP or CRM system, Excel lists, production facilities, service reports, PDF documentation, emails, or shared drives. The experiential knowledge of employees can also be crucial for an AI application.
In this context, Retrieval-Augmented Generation (RAG) becomes interesting. Here, no own AI model is trained with company data. Instead, an AI assistant searches internal documents for information to make available to the language model for its response. It is a low-threshold method to make knowledge accessible, make decisions based on data, or accelerate innovation without having to train expensive AI models yourself.
Which data is needed depends on the specific AI use case. At the same time, the existing data base determines which AI applications can be usefully implemented in your company. It doesn't require an extensive data management concept. More important is a clear thread: What should the AI be able to do, what information does it need, and are these already available?
For starting out, your company neither needs a fully cleansed data stock nor automatically a new data platform or additional software. Incremental, targeted improvements to data quality gradually create a reliable data base that provides tangible benefits to your company and ensures its competitiveness.
A solid data base is worthwhile not only for the use with AI: It makes evaluations more reliable, increases transparency in processes, and improves decisions. If you systematically use data, you can also develop new products and offers more quickly and strengthen competitiveness.
The topic is particularly relevant when...
you want to introduce a specific AI application,
you want to make internal company knowledge usable with AI,
you want to systematically evaluate customer or service data, or
you need to combine information from several systems.
If you are unsure whether your data is sufficient for the planned use, early examination is advisable. You can clarify this based on a specific use case. This allows you to assess which data you are missing and how much effort is necessary for a viable data base.
:quality(85))
German Data Science Society (GDS) e.V
Member of the board
"Data is the great differentiator for AI-supported digitization and digital solutions. BAIOSPHERE brings together users, universities, and providers with specific expertise to make Germany and Europe successful."
:quality(85))
JMU Würzburg
Spokesman of CAIDAS and Head of Data
Science Chair (Computer Science X)
"Data is the new raw material of this world and the basis for the ongoing revolution in AI. They offer many new opportunities, whose potential we should and must utilize without losing sight of the risks."
Event
Summit by Bayern Innovativ on real-time data and AI for decisions in production and mobility in Munich.
Bavarian Data Summit 2026Event
Why AI agents quickly become expensive – and how medium-sized companies can manage costs, quality, and benefits
How AI creates measurable added valueEvent
Webinar on AI Search and New Content Strategy
Visibility in AI Search: Top Position in AI AnswersEvent
Practice forum on scalable AI applications in medium-sized businesses
From AI Investments to Measurable Value Creation:quality(85))
Event
:quality(85))
Event
Step 1 – Start from the use case
Do not start with the question "What data do we have?" but with "What should AI be able to do in our company?" For example, if it should answer service questions, it may need manuals, repair reports, and spare part information. For a sales forecast, sales figures, time periods, and other influencing factors are relevant. First, create a short list of the information that is actually needed for the desired task. This way, you avoid preparing data that is not relevant for the project.
Step 2 – Check what's already available
Then, specifically search for this information. It is worth looking beyond the traditional databases: Relevant company data can be in the ERP system, but also in Excel files, PDFs, emails, SharePoint, a network drive, or with individual employees. Note where the necessary information is located, who can access it, and what is still missing. A simple table is sufficient for an initial overview - your company does not need new software for this.
Step 3 – Test the quality with real examples
Do not evaluate data quality abstractly but based on 20 to 30 typical cases from your company's everyday business. For a service assistant, these could be actual questions from customer service. Do you find a current and clear answer for these questions in the available documents? Do different product names, outdated price lists, duplicate customer entries, or contradictory work instructions appear? It's clear what needs to be improved before a pilot project.
Step 4 – Prepare only the necessary data
Your company does not need to clean up all data first. Focus on the information necessary for the selected use case. Remove outdated versions, standardize important designations, and add missing information. Furthermore, clarify who will keep this data up-to-date in the future and who is allowed to access it. For personal, confidential, or otherwise sensitive information, it should be clarified in advance whether and under what conditions they may be processed with the intended AI solution.
Step 5 – Test small and learn from it
First test the application with a limited data set and real tasks from daily work. If the AI only provides twelve useful answers out of 20 typical questions, review the remaining eight cases: Was information missing? Was it outdated? Could the AI not clearly assign it? Or is the problem not with the data at all? Before your company invests money in new infrastructure or external support, it should first check whether the use case can be sensibly tested with the existing data and systems. The result may also be that the effort for this use case is not worthwhile - in that case, it is more sensible to select another use case than to invest further time and money in data preparation.
Anyone who wants to use AI with company data quickly encounters terms like RAG, Training or Fine-Tuning. Which approach makes sense depends on what the AI should do with the data - and significantly affects effort and costs.
Using company knowledge: If an AI assistant is to answer questions about products, processes, or documentation, it often does not require training its own model. In Retrieval-Augmented Generation (RAG) , the system searches for suitable information in the approved company data and provides it to the language model for the answer. Therefore, it is important that relevant information is current and findable. When selecting a solution, you should ask: How does the AI find the information, and can it name the sources used?
Making predictions: If AI is expected to predict machine failures or sales volumes, suitable historical training data is needed. It's not just about the quantity. Thousands of machine values are not helpful if it's not documented when a defect actually occurred. A smaller, well-documented data set can be significantly more valuable.
Fine-Tuning: not automatically necessary. Fine-tuning adjusts an existing AI model for a specific task. For the pure use of company knowledge, this is often not required. If a provider recommends fine-tuning, you should ask: What specific problem does it solve, and why is access to existing information not sufficient?
External AI services: Even if data is not used for model training, it can still be stored or otherwise processed. For confidential or personal information, you should therefore check the conditions of the specific service, tariff, and contract.
Loading...
Loading...
You can find this and many more pieces of information on the AI compass topic "Data" in our download portal:
Loading...
Loading...
Here's how it works: Turn the arrows to navigate through our ten focus topics. With one click, you will reach the subpage with further information.