Integration · Website crawler
Connect a website as a knowledge source
Crawl selected public pages into a Cloud Vault data space and make their content searchable.
The website crawler is a knowledge source in meinGPT. It follows pages within a defined scope, extracts their main content and synchronizes it into a Cloud Vault data space for search in chats and assistants.
An administrator enters a start URL and limits depth, page count and included or excluded paths. The crawler can use browser rendering for JavaScript-heavy pages. Later runs refresh content according to the configured freshness window.
- Source
- Website crawler
- Capability
- Crawl, index and search selected website content
- Connection
- Start URL and crawl scope; optional headers, user agent or proxy
- Data residency
- Cloud Vault source; no dedicated Outpost required
Source, permissions, and data residency are documented separately for each connector.
Typical workflows
- Search public documentation
- Summarize page collections
- Limit the crawl
What the integration makes possible
01
Search public documentation
Use approved product, help or policy pages as current assistant context.
02
Summarize page collections
Condense related pages into a briefing with source URLs.
03
Limit the crawl
Restrict indexing to relevant sections and exclude paths that do not belong in the data space.
Review the prompt and result openly
Summarize the setup requirements from our public help center and cite the pages you used.
Result
Setup requires an organization account, an administrator and an approved data source. The answer uses /help/setup and /help/data-sources.
Does this connection fit your use case?
We clarify the source, permissions, and connection required for the concrete use case. If the existing integration fits, you can move straight to setup afterwards.
Security and operations
Define a narrow crawl scope and review which public pages enter the data space. Custom headers and proxies should be used only when their access implications are understood.
Known limitations
- 01Login walls, robots rules, missing links and unsupported rendering can keep pages out of the index.
Frequently asked questions
Yes, when browser rendering is selected, subject to the page and crawl configuration.