Data sources

Which sources an Outpost connects, which meinGPT connects directly, and how a connection is made.

A data source gets its name and permissions in meinGPT — that is where who may find it is decided. Folders on an Outpost machine are created there and only adopted here; how much a folder releases is decided by the Outpost too. Where its documents live is decided afterwards — and for some sources the answer is "on a machine at your site", which is exactly what the Outpost is for.

Who fetches the files?

SourcesConfigured where
meinGPT directlySharePoint, OneDrive, Google Drive, Confluence, S3, WebDAV, web crawlerEntirely in meinGPT. No machine of your own needed.
Only via an OutpostFolders on a machine — including SMB shares once mounted; databases and local MCP servers when the separate services preview is enabledData source in meinGPT, the link in the Outpost

Services are not generally enabled

Databases and MCP servers only appear when the controlled services preview is enabled for your workspace. Without that enablement, the Outpost only shows file-based shares.

In the default case you do not need an Outpost

If your documents live in SharePoint or Google Drive, meinGPT connects them directly. An Outpost pays off when source files and the search index must stay in your network, or when systems cannot be reached from the cloud at all — a network drive and a database on your own network cannot. You decide per folder on the Outpost what a request may transmit.

How an Outpost source comes about

Two halves in two places, in this order:

  1. In the Outpost: share the folder — pick the path and decide how much it releases.
  2. In meinGPT: adopt the reported folder under Shares and name it. That is where who may find it is decided.

There is no button in the Outpost that creates a data source. If you are looking for one, step 1 is missing.

The dialog in the Outpost knows exactly one source type: a folder path on this machine — a permanently mounted SMB share counts too. S3 storage, WebDAV and web crawler are connected exclusively in meinGPT directly. If such a service is not reachable from the internet — an internal MinIO or a Nextcloud, say — that needs its own connection method (IP allowlisting, Enterprise Connection Network or VPN) instead of an Outpost — see Guide: On-premise.

The sources in detail

Needs an Outpost

Directly in meinGPT, no Outpost needed

What is the same for every source

  1. The Outpost fetches the files into its working area on this machine.
  2. It extracts text, splits it into passages, encodes them and stores them in the search index.
  3. Later runs only process what changed.
  4. What it cannot read, it skips — and records it by file extension in the folder's row under Shares.

Point 4 is the one to keep an eye on: a skipped file cannot be found, and the search tells nobody. → Console & operations

Permissions are created in meinGPT, not in the file system

File, NTFS, AD and POSIX permissions are not carried into the index. Anyone allowed to see a data source in meinGPT finds everything in it. How you split things into data sources is your permission model.

Exception: for local NTFS folders, there is a beta feature, Windows user permissions (Beta), that honors each user's real NTFS permissions — see A folder on this machine for details.

Was this page helpful?