Preparing Your Data

Organize folders, file names, formats, and content so that search in data sources reliably finds the right passages

How well MeinGPT finds the right information in a data source depends heavily on how the files are stored, named, and structured. This page describes what you can do about that before connecting a source or while tidying one up. For how to set up a data source, see Data Sources & RAG; for which areas to pick for a pilot, see the Guide: Data Quality.

The advice applies to cloud data sources. An Outpost reads some other formats and has its own limits, see Folders on this computer.

The essentials

  • Store only the currently valid version, and each file only once.
  • Name files so that the name already says what they are about.
  • Add a text layer to scanned PDFs beforehand.
  • Convert .doc and .ppt to .docx and .pptx.
  • Structure important documents with real headings.

What search sees

Documents are split into short passages. For a question, MeinGPT looks for the matching passages, by keywords and by meaning (Search & Answers). For preparation, that means:

  • The file name is part of the searchable text. A name like Scan_0034.pdf adds nothing to search.
  • The folder path is visible. It shows where each source comes from, and the AI can narrow a search to individual folders.
  • Search works on text. Scans without text recognition are found poorly or not at all, and images inside documents only through their description (Images in a data pool).
  • Search does not evaluate custom metadata, such as SharePoint columns or tags. The file's modification date, however, it does use.

Storage and naming

Folder structure

  • Name folders by subject, such as customer, product, or process, not by person.
  • Keep the structure flat, at most three to four levels. Example: Sales / Product data sheets / Pumps.

File names

  • Define a pattern, for example DocumentType_Topic_Year: TravelExpensePolicy_Domestic_2025.pdf, DataSheet_Pump-KX200.pdf.
  • Use the terms colleagues actually ask about. Internal abbreviations are fine in addition, not instead.
  • Separate words with _ or -.

Versions and legacy files

When old and current versions sit side by side, MeinGPT can cite an outdated one. So:

  • Ideally store only the currently valid version in the data source.
  • If you need older versions, put them in a subfolder called Archive or Old. MeinGPT also recognizes Archiv, Alt, Veraltet, and Historie. Content there is ranked slightly lower in search, but it can still be found when nothing else answers the question.
  • When several versions are in the same folder, mark them unambiguously with a year or version number in the name, for example PriceList_2024.pdf and PriceList_2025.pdf, or Manual_v3.pdf. MeinGPT then recognizes them as belonging together and prefers the newest, unless the question names a specific version.
  • Avoid suffixes like "final", "new", or "corrected". MeinGPT cannot tell an order from those.

You can also mark individual files as Authoritative or Outdated by hand, see Marking files as authoritative or outdated.

No duplicates

Duplicate files are processed twice and crowd other hits out of the results. Store each file in one place only. On an old drive, your IT team can find duplicates with a tool that compares files by content, such as the free "dupeGuru". The same applies to folders connected in several data sources, see Narrowing down data.

Preparing file formats

Which formats and limits apply is listed under Formats & Access. This section covers what you can do beforehand.

Making scanned PDFs searchable (text recognition)

To search, a scan is initially just an image. MeinGPT does recognize text in scans automatically, but only up to 50 pages per file. Running text recognition beforehand gives more reliable results.

Check whether a PDF contains text: Open the PDF and search for a visible word with Ctrl+F. If the word isn't found or you can't select any text, it's a pure scan.

To add a text layer to scans:

  • Single, short documents: Upload the scan in chat and have MeinGPT transcribe it (Prompt 1 below). Check the result, paste it into Word, and save it as .docx.
  • Adobe Acrobat (Pro): Tool "Scan & OCR" → "Recognize text" → "In this file", then save. Acrobat also offers batch processing for several files.
  • Many files: Your IT team can use the free tool "OCRmyPDF". It adds a text layer to whole folders automatically.
  • Scans with more than 50 pages: Add a text layer with an OCR tool before storing them. Splitting into parts of at most 50 pages alone is not reliable for some scans (see "Splitting very large documents" for text files).
Converting old Office formats

A cloud data source does not take in .doc, .ppt, and .json files or OneNote.

  • Single file: Open it in Word or PowerPoint and choose File → Save As → "Word Document" (.docx) or "PowerPoint Presentation" (.pptx).
  • Many files: Your IT team can convert them in bulk with the free LibreOffice, for example with soffice --headless --convert-to docx *.doc.
  • OneNote: Export sections as PDF or Word files (File → Export). For a regular automated export, see the OneNote workaround.
Removing password protection

MeinGPT cannot read encrypted or password-protected PDFs. In Adobe Acrobat: File → Properties → Security → "No Security", then save. This requires that you know the password and are authorized to remove it.

Splitting very large documents

For very large files, roughly from several hundred pages of running text, meaning search covers only a selection across the whole file, and keyword search only the beginning. The file then shows the note Partially indexed.

Split collective documents by chapter or topic into separate files and give each file a descriptive name (see "File names"). You can do this in Adobe Acrobat (Organize Pages → Split) or with the free PDF24 ("Split PDF" function).

Structuring Excel tables

A table is not searched row by row. What is captured is its structure (sheet names, title rows, column headers, sample rows) and the first data rows, each with its column headers. This only works for clearly structured tables and up to certain sizes.

  • Use one table per sheet, with the column headers in the first row.
  • Give columns and sheets descriptive names ("Delivery date" instead of "DD").
  • Mind the limits: at most 2,000 rows per sheet, 5,000 rows per file, and 64 columns. For very wide tables or long cells, it is fewer.

Larger tables are better suited to analysis in chat (upload the file and have it analyzed) or in the Code Sandbox than to search in a data source.

Storing emails with attachments

For stored emails (.msg, .eml), attachments are not reliably captured. Additionally store important attachments as separate files. In Outlook: click the attachment → "Save As" or "Save All Attachments". More under Emails and attachments.

Making content easy to read

This part is mainly worthwhile for frequently asked documents, such as policies, manuals, or product information.

Structure with real headings

Documents are split into sections at headings. Real headings lead to sensibly delimited sections. In Word, use the "Heading 1", "Heading 2", etc. styles (Home tab → Styles), not just bold or larger text.

Paragraphs that make sense on their own

MeinGPT finds individual short passages. If a passage only says "It is €0.30", the reference is missing.

  • Name the subject in the sentence: "The mileage allowance for business trips is €0.30" instead of "It is €0.30".
  • Spell out technical terms, or explain them once per section.
  • MeinGPT can revise existing documents accordingly (Prompt 2 below). Check the result for content.

Add content from images and charts as text

Images in Word and PowerPoint files are described automatically, but important statements are easier to find reliably as text. Add a short caption or description below important graphics. MeinGPT can write it (Prompt 3 below). How images reach the index and how to lay them out: Images in a data pool.

Moving to a new drive

MeinGPT takes the files' modification date into account, for example in the time-range filter in chat. If all files carry the same date after a move, old and current documents can no longer be told apart.

  • For moves to SharePoint or OneDrive, use Microsoft's migration tools (SharePoint Migration Tool or Migration Manager in the SharePoint admin center). These carry over the file dates.
  • After the move, spot-check the "Modified" column in the SharePoint library.

Prompts for MeinGPT

These prompts let MeinGPT help with preparation. Check each result before you store it.

Prompt 1: Transcribe a scan

Transcribe the text of this scanned document completely and verbatim. Keep headings, lists, and tables in their structure. Do not summarize and do not add anything. Mark illegible passages with [illegible].

Prompt 2: Make sections understandable on their own

Revise this document so that every paragraph is understandable without the rest of the document. Replace references like "it", "this", or "see above" with the specific term. Structure the text with clear headings. Do not change any content, numbers, or statements.

Prompt 3: Describe an image or chart

Describe this chart in 2–4 sentences as a caption for a technical document. Name all values, labels, and the central message that can be read from it.

Further reading

Was this page helpful?