Log in

Add your website

Add pages from any website so your bot can answer questions from them: your help centre, docs, product pages or blog.

Who can do this: team owners and admins.

Steps

  1. Open your bot, go to the Data sources tab and click New data source.
  2. Choose Website and click Create data source.
    The Website data source card selected on the Choose a data source page
  3. On the Content tab, click the button for how you want to add pages (on a small screen, these are under Add):
    • A single URL: enter one page address and click Add.
    • Add bulk URLs: paste several addresses, one per line, and click Add.
    • Add sitemap: enter your sitemap address (for example https://example.com/sitemap.xml) and click Add URLs from sitemap. Every page in the sitemap is added.
    • Crawl: Chat Thing follows links from a starting page to find pages for you. See crawl a website.

    The Content tab of a new website data source, with A single URL, Add bulk URLs, Add sitemap and Crawl buttons above an empty page table
  4. Check the pages listed in the table. To remove one, delete its row.
  5. Optional but recommended: set a CSS selector so only the main content of each page is read. See choose which part of each page is read.
  6. Click Synchronise.

Crawl a website

  1. On the Content tab, click Crawl (on a small screen, Add > Crawl).
  2. Enter the address to start from, such as https://example.com.
  3. Choose a Strategy for which links to follow: Same hostname (the default), Same domain or Same origin. Each option shows examples of what it matches.
  4. To stay within one section, turn on Only include urls which match the path of the url being crawled?. For example, crawling https://example.com/help then only adds pages under /help.
    The Crawl website dialog with a starting URL, the Same hostname strategy, examples of matching URLs and the path-matching toggle
  5. Click Start crawl. Crawling can take several minutes on a large or slow site.
  6. When it finishes, click Review pages.
    A finished crawl showing its start time, duration and pages discovered, with a Review pages button
  7. Turn off the toggle next to any page you don't want, then click Add N URLs (the button shows how many pages are selected).

A crawl adds up to 600 pages. For bigger sites, use your sitemap, or crawl each section separately.

Choose which part of each page is read

By default Chat Thing reads each page's whole body and removes the header and footer. Menus, sidebars and cookie banners can still get through, which adds noise and uses storage tokens. A CSS selector fixes this.

  1. Open the data source's Scraping settings tab.
  2. Under Content selector, enter a CSS selector for the element that holds your main content, such as main or article.
  3. Under CSS excludes, click Add exclude rule for each part you want removed, such as nav or .cookie-banner.
    The Scraping settings tab with the CSS selector field set to body and header and footer listed under CSS excludes
  4. Click Save, then sync the data source again.

Not sure which selector to use? See CSS selectors for common platforms. Test on a data source with one URL before syncing a whole site. The selector applies to every page in the data source, so if parts of your site use different layouts, put them in separate data sources.

Add a password-protected site

If your site uses HTTP Basic Auth (the browser pop-up asking for a username and password):

  1. Open the Scraping settings tab.
  2. Under Basic auth, enter the Username and Password.
    The Basic auth section of Scraping settings with Username and Password fields and a Save button
  3. Click Save and sync the data source.

This only works with Basic Auth. Pages behind a login form or single sign-on can't be read.

Add a website from an AI assistant

The MCP tools discover_pages and add_data_source do the same job from Claude, ChatGPT or another AI assistant, with one extra: URL patterns. When adding pages from a discovery, includePatterns and excludePatterns filter pages by URL path, for example /blog/** or !/legal/* (* matches within one path segment, ** across segments). The dashboard doesn't have URL pattern fields.

Check it works

  1. After the sync, the Status column shows Finished for each page.
  2. On the data source card menu, click Documents and open a few documents. They should contain your page content without menus or footers.
  3. Ask your bot a question that one of the pages answers.

Troubleshooting

A page synced but has no content

Chat Thing reads the HTML your server sends and doesn't run JavaScript. Pages that build their content in the browser (some single-page apps) come through almost empty and are skipped. Also check your CSS selector exists on that page.

"We were unable to crawl this sitemap"

Your site may be blocking automated requests, or the sitemap address is wrong. Open the address in your browser to check, then allow our requests in your firewall or bot-protection settings, or add the pages with Add bulk URLs instead.

New pages on my site aren't in the bot

Re-syncing a website data source refreshes the pages already in it; it doesn't crawl the site or re-read the sitemap again. Add new pages with A single URL, Add bulk URLs, Add sitemap or Crawl (only new pages are added), then sync.

A page was removed from my site but the bot still uses it

Re-syncing doesn't remove pages that have gone from your site: if a page now returns an error, its last good content is kept. Delete the page's row on the Content tab and sync again.

Last updated