---
title: "Preparing Your Data"
description: "Organize folders, file names, formats, and content so that search in data sources reliably finds the right passages"
canonical_url: "https://meingpt.com/en/docs/integrations/data-pools-rag/prepare-data"
language: en
---

# Preparing Your Data

How well MeinGPT finds the right information in a data source depends heavily on how the files are stored, named, and structured. This page describes what you can do about that before connecting a source or while tidying one up. For how to set up a data source, see [Data Sources & RAG](/en/docs/integrations/data-pools-rag); for which areas to pick for a pilot, see the [Guide: Data Quality](/en/docs/integrations/guide-data-quality).

The advice applies to cloud data sources. An Outpost reads some other formats and has its own limits, see [Folders on this computer](/en/docs/integrations/vault/sources/local#what-gets-read-and-what-does-not).

**The essentials**

- Store only the currently valid version, and each file only once.
- Name files so that the name already says what they are about.
- Add a text layer to scanned PDFs beforehand.
- Convert `.doc` and `.ppt` to `.docx` and `.pptx`.
- Structure important documents with real headings.

## What search sees

Documents are split into short passages. For a question, MeinGPT looks for the matching passages, by keywords and by meaning ([Search & Answers](/en/docs/integrations/data-pools-rag/search#how-search-works)). For preparation, that means:

- **The file name is part of the searchable text.** A name like `Scan_0034.pdf` adds nothing to search.
- **The folder path is visible.** It shows where each source comes from, and the AI can narrow a search to individual folders.
- **Search works on text.** Scans without text recognition are found poorly or not at all, and images inside documents only through their description ([Images in a data pool](/en/docs/integrations/data-pool-images)).
- **Search does not evaluate custom metadata**, such as SharePoint columns or tags. The file's modification date, however, it does use.

## Storage and naming

### Folder structure

- Name folders by subject, such as customer, product, or process, not by person.
- Keep the structure flat, at most three to four levels. Example: `Sales / Product data sheets / Pumps`.

### File names

- Define a pattern, for example `DocumentType_Topic_Year`: `TravelExpensePolicy_Domestic_2025.pdf`, `DataSheet_Pump-KX200.pdf`.
- Use the terms colleagues actually ask about. Internal abbreviations are fine in addition, not instead.
- Separate words with `_` or `-`.

### Versions and legacy files

When old and current versions sit side by side, MeinGPT can cite an outdated one. So:

- Ideally store only the currently valid version in the data source.
- If you need older versions, put them in a subfolder called `Archive` or `Old`. MeinGPT also recognizes `Archiv`, `Alt`, `Veraltet`, and `Historie`. Content there is ranked slightly lower in search, but it can still be found when nothing else answers the question.
- When several versions are in the same folder, mark them unambiguously with a year or version number in the name, for example `PriceList_2024.pdf` and `PriceList_2025.pdf`, or `Manual_v3.pdf`. MeinGPT then recognizes them as belonging together and prefers the newest, unless the question names a specific version.
- Avoid suffixes like "final", "new", or "corrected". MeinGPT cannot tell an order from those.

You can also mark individual files as **Authoritative** or **Outdated** by hand, see [Marking files as authoritative or outdated](/en/docs/integrations/knowledge-insights-glossary#marking-files-as-authoritative-or-outdated).

### No duplicates

Duplicate files are processed twice and crowd other hits out of the results. Store each file in one place only. On an old drive, your IT team can find duplicates with a tool that compares files by content, such as the free "dupeGuru". The same applies to folders connected in several data sources, see [Narrowing down data](/en/docs/integrations/data-pools-rag#narrow-the-data-scope).

## Preparing file formats

Which formats and limits apply is listed under [Formats & Access](/en/docs/integrations/data-pools-rag/formats-and-access). This section covers what you can do beforehand.

### Making scanned PDFs searchable (text recognition)

To search, a scan is initially just an image. MeinGPT does recognize text in scans automatically, but only up to 50 pages per file. Running text recognition beforehand gives more reliable results.

**Check whether a PDF contains text:** Open the PDF and search for a visible word with Ctrl+F. If the word isn't found or you can't select any text, it's a pure scan.

To add a text layer to scans:

- **Single, short documents:** Upload the scan in chat and have MeinGPT transcribe it (Prompt 1 below). Check the result, paste it into Word, and save it as `.docx`.
- **Adobe Acrobat (Pro):** Tool "Scan & OCR" → "Recognize text" → "In this file", then save. Acrobat also offers batch processing for several files.
- **Many files:** Your IT team can use the free tool "OCRmyPDF". It adds a text layer to whole folders automatically.
- **Scans with more than 50 pages:** Add a text layer with an OCR tool before storing them. Splitting into parts of at most 50 pages alone is not reliable for some scans (see "Splitting very large documents" for text files).

### Converting old Office formats

A cloud data source does not take in `.doc`, `.ppt`, and `.json` files or OneNote.

- **Single file:** Open it in Word or PowerPoint and choose File → Save As → "Word Document" (`.docx`) or "PowerPoint Presentation" (`.pptx`).
- **Many files:** Your IT team can convert them in bulk with the free LibreOffice, for example with `soffice --headless --convert-to docx *.doc`.
- **OneNote:** Export sections as PDF or Word files (File → Export). For a regular automated export, see the [OneNote workaround](/en/docs/integrations/data-pools-rag/formats-and-access#limits-per-file).

### Removing password protection

MeinGPT cannot read encrypted or password-protected PDFs. In Adobe Acrobat: File → Properties → Security → "No Security", then save. This requires that you know the password and are authorized to remove it.

### Splitting very large documents

For very large files, roughly from several hundred pages of running text, meaning search covers only a selection across the whole file, and keyword search only the beginning. The file then shows the note **Partially indexed**.

Split collective documents by chapter or topic into separate files and give each file a descriptive name (see "File names"). You can do this in Adobe Acrobat (Organize Pages → Split) or with the free PDF24 ("Split PDF" function).

### Structuring Excel tables

A table is not searched row by row. What is captured is its structure (sheet names, title rows, column headers, sample rows) and the first data rows, each with its column headers. This only works for clearly structured tables and up to certain sizes.

- Use **one table per sheet**, with the column headers in the first row.
- Give columns and sheets descriptive names ("Delivery date" instead of "DD").
- Mind the limits: at most 2,000 rows per sheet, 5,000 rows per file, and 64 columns. For very wide tables or long cells, it is fewer.

Larger tables are better suited to analysis in chat (upload the file and have it analyzed) or in the [Code Sandbox](/en/docs/platform/code-sandbox) than to search in a data source.

### Storing emails with attachments

For stored emails (`.msg`, `.eml`), attachments are not reliably captured. Additionally store important attachments as separate files. In Outlook: click the attachment → "Save As" or "Save All Attachments". More under [Emails and attachments](/en/docs/integrations/data-pools-rag/formats-and-access#emails-and-attachments).

## Making content easy to read

This part is mainly worthwhile for frequently asked documents, such as policies, manuals, or product information.

### Structure with real headings

Documents are split into sections at headings. Real headings lead to sensibly delimited sections. In Word, use the "Heading 1", "Heading 2", etc. styles (Home tab → Styles), not just bold or larger text.

### Paragraphs that make sense on their own

MeinGPT finds individual short passages. If a passage only says "It is €0.30", the reference is missing.

- Name the subject in the sentence: "The mileage allowance for business trips is €0.30" instead of "It is €0.30".
- Spell out technical terms, or explain them once per section.
- MeinGPT can revise existing documents accordingly (Prompt 2 below). Check the result for content.

### Add content from images and charts as text

Images in Word and PowerPoint files are described automatically, but important statements are easier to find reliably as text. Add a short caption or description below important graphics. MeinGPT can write it (Prompt 3 below). How images reach the index and how to lay them out: [Images in a data pool](/en/docs/integrations/data-pool-images#best-practices).

## Moving to a new drive

MeinGPT takes the files' modification date into account, for example in the time-range filter in chat. If all files carry the same date after a move, old and current documents can no longer be told apart.

- For moves to SharePoint or OneDrive, use Microsoft's migration tools (SharePoint Migration Tool or Migration Manager in the SharePoint admin center). These carry over the file dates.
- After the move, spot-check the "Modified" column in the SharePoint library.

## Prompts for MeinGPT

These prompts let MeinGPT help with preparation. Check each result before you store it.

**Prompt 1: Transcribe a scan**

```text
Transcribe the text of this scanned document completely and verbatim. Keep headings, lists, and tables in their structure. Do not summarize and do not add anything. Mark illegible passages with [illegible].
```

**Prompt 2: Make sections understandable on their own**

```text
Revise this document so that every paragraph is understandable without the rest of the document. Replace references like "it", "this", or "see above" with the specific term. Structure the text with clear headings. Do not change any content, numbers, or statements.
```

**Prompt 3: Describe an image or chart**

```text
Describe this chart in 2–4 sentences as a caption for a technical document. Name all values, labels, and the central message that can be read from it.
```

## Further reading

### [Formats & Access](/en/docs/integrations/data-pools-rag/formats-and-access)

Supported formats, limits per file, access control.

### [Sync Status & Troubleshooting](/en/docs/integrations/data-pools-rag/troubleshooting)

Check which files are not searchable and why.

### [Search Settings](/en/docs/integrations/data-pool-search-settings)

Use Try it out to check what a question finds in the source.
