curl --request POST \
--url https://{host}/v2/workspaces/slug:storage/webhooks/v1/knowledge_bases/{knowledgeBaseId}/web_sources \
--header 'Authorization: Bearer <token>' \
--header 'Content-Type: application/json' \
--data '
{
"url": "<string>"
}
'const options = {
method: 'POST',
headers: {Authorization: 'Bearer <token>', 'Content-Type': 'application/json'},
body: JSON.stringify({url: '<string>'})
};
fetch('https://{host}/v2/workspaces/slug:storage/webhooks/v1/knowledge_bases/{knowledgeBaseId}/web_sources', options)
.then(res => res.json())
.then(res => console.log(res))
.catch(err => console.error(err));import requests
url = "https://{host}/v2/workspaces/slug:storage/webhooks/v1/knowledge_bases/{knowledgeBaseId}/web_sources"
payload = { "url": "<string>" }
headers = {
"Authorization": "Bearer <token>",
"Content-Type": "application/json"
}
response = requests.post(url, json=payload, headers=headers)
print(response.text){
"id": "<string>",
"object": "knowledge_base.web_source",
"knowledge_base_id": "<string>",
"url": "<string>",
"created_at": 123,
"crawler_options": {},
"last_crawl_at": 123,
"next_crawl_at": 123,
"metrics": {
"pages_total": 1,
"pages_indexed": 1,
"pages_failed": 1,
"pages_pending": 1,
"pages_skipped": 1,
"last_run_status": "running",
"last_run_started_at": 123,
"last_run_completed_at": 123
},
"updated_at": 123
}Create a recurring crawl seed
Registers a start URL + per-domain crawler options. The first
crawl is scheduled immediately unless the knowledge base is
paused (crawl_settings.paused - creating a seed never resumes
the crawler). Discovered pages arrive as documents
(origin: web_source, web_source_id set), deduped against
manual documents by source_key. The engine-global crawl
settings (webpages_limit, periodicity, paused, plus
decorative depth) are rejected here with 400 - set them via
the KB PATCH crawl_settings. The URL is the idempotency key:
re-POSTing an already-registered URL returns the existing seed
with 200 rather than creating a duplicate. Requires editor+.
curl --request POST \
--url https://{host}/v2/workspaces/slug:storage/webhooks/v1/knowledge_bases/{knowledgeBaseId}/web_sources \
--header 'Authorization: Bearer <token>' \
--header 'Content-Type: application/json' \
--data '
{
"url": "<string>"
}
'const options = {
method: 'POST',
headers: {Authorization: 'Bearer <token>', 'Content-Type': 'application/json'},
body: JSON.stringify({url: '<string>'})
};
fetch('https://{host}/v2/workspaces/slug:storage/webhooks/v1/knowledge_bases/{knowledgeBaseId}/web_sources', options)
.then(res => res.json())
.then(res => console.log(res))
.catch(err => console.error(err));import requests
url = "https://{host}/v2/workspaces/slug:storage/webhooks/v1/knowledge_bases/{knowledgeBaseId}/web_sources"
payload = { "url": "<string>" }
headers = {
"Authorization": "Bearer <token>",
"Content-Type": "application/json"
}
response = requests.post(url, json=payload, headers=headers)
print(response.text){
"id": "<string>",
"object": "knowledge_base.web_source",
"knowledge_base_id": "<string>",
"url": "<string>",
"created_at": 123,
"crawler_options": {},
"last_crawl_at": 123,
"next_crawl_at": 123,
"metrics": {
"pages_total": 1,
"pages_indexed": 1,
"pages_failed": 1,
"pages_pending": 1,
"pages_skipped": 1,
"last_run_status": "running",
"last_run_started_at": 123,
"last_run_completed_at": 123
},
"updated_at": 123
}Authorizations
User session JWT or instance API key (iak_*). Send as Authorization: Bearer <token>.
Path Parameters
Knowledge base id. Legacy physical prefix vs_ (the kb_ rename is deferred).
128^vs_[A-Za-z0-9-]+$Response
Existing seed returned (the URL was already registered).
Recurring web crawl seed. A seed is a generator of documents,
not a document: pages it discovers materialize as
knowledge_base.document rows (source_type: web_page,
origin: web_source, web_source_id set) deduped in the same
source_key space as manual documents. A seed carries only its
URL and per-domain crawler_options - the engine-global crawl
settings (webpages_limit, periodicity, paused) live on the
knowledge base's crawl_settings object (one crawler per KB).
128^seed_[A-Za-z0-9-]+$knowledge_base.web_source ^vs_[A-Za-z0-9-]+$Seed URL the crawler starts from.
Provider-specific crawl configuration passed through to the crawler (e.g. include/exclude URL patterns, auth headers, render mode). Validated by the crawler, not by this API. Projects PER DOMAIN on the store-level crawler: sources on the same hostname share these options.
Reserved. Currently always null - there is no per-seed
crawl-run ledger yet.
Reserved. Currently always null (no per-seed scheduled-run
computation yet).
Per-seed page counts, DB-derived from documents attributed
to this seed via parent_web_source_ids (multi-attribution,
§F-5). Independent of crawler reachability - these are
always available even when the crawler is unreachable.
Hide child attributes
Hide child attributes
x >= 0x >= 0x >= 0x >= 0Reserved. Currently always 0 (no per-seed "skipped" status).
x >= 0Reserved. Currently always null (no per-seed run ledger).
running, completed, failed, partial, null Reserved. Currently always null.
Reserved. Currently always null.
Was this page helpful?