Integration · Website crawler

Connect a website as a knowledge source

Crawl selected public pages into a Cloud Vault data space and make their content searchable.

The website crawler is a knowledge source in meinGPT. It follows pages within a defined scope, extracts their main content and synchronizes it into a Cloud Vault data space for search in chats and assistants.

An administrator enters a start URL and limits depth, page count and included or excluded paths. The crawler can use browser rendering for JavaScript-heavy pages. Later runs refresh content according to the configured freshness window.

Technical frame
Website crawler
meinGPT
Source
Website crawler
Capability
Crawl, index and search selected website content
Connection
Start URL and crawl scope; optional headers, user agent or proxy
Data residency
Cloud Vault source; no dedicated Outpost required

Source, permissions, and data residency are documented separately for each connector.

Setup in the docs

Typical workflows

  • Search public documentation
  • Summarize page collections
  • Limit the crawl

What the integration makes possible

01

Search public documentation

Use approved product, help or policy pages as current assistant context.

02

Summarize page collections

Condense related pages into a briefing with source URLs.

03

Limit the crawl

Restrict indexing to relevant sections and exclude paths that do not belong in the data space.

Review the prompt and result openly

Prompt

Summarize the setup requirements from our public help center and cite the pages you used.

Result

Setup requires an organization account, an administrator and an approved data source. The answer uses /help/setup and /help/data-sources.

Does this connection fit your use case?

We clarify the source, permissions, and connection required for the concrete use case. If the existing integration fits, you can move straight to setup afterwards.

Security and operations

Define a narrow crawl scope and review which public pages enter the data space. Custom headers and proxies should be used only when their access implications are understood.

Known limitations

  1. 01Login walls, robots rules, missing links and unsupported rendering can keep pages out of the index.

Frequently asked questions

Yes, when browser rendering is selected, subject to the page and crawl configuration.