The Algolia Crawler connector lets your AI agent act in each user's Algolia Crawler account. Each user connects their own Algolia Crawler username and password once, and Scalekit sends it with every call, so your agent never handles credentials. It comes with 20 tools.
- Tools
- 20
- What they doRead · write · destructive
- 9 · 9 · 29 read9 write2 destructive
- Users sign in with
- Username and password
Setup
Install the SDK
Terminal window npm install @scalekit-sdk/node dotenvTerminal window pip install scalekit-sdk-python python-dotenvSet your credentials
Add your Scalekit credentials to your
.envfile. Find values in app.scalekit.com > Developers > API Credentials..env SCALEKIT_ENVIRONMENT_URL=<your-environment-url>SCALEKIT_CLIENT_ID=<your-client-id>SCALEKIT_CLIENT_SECRET=<your-client-secret>Create the Algolia Crawler connection
In AgentKit > Connections, create an Algolia Crawler connection. The name you give it is the
connection_nameyour code passes. See Configure connections.
Tools
Pass the exact name toexecute_toolalgoliacrawler_get_config_versionRetrieve one saved version of a crawler's configuration.Read-onlyGet Config Version
Retrieve one saved version of a crawler's configuration. Returns the version number, creation time, author id, and the full configuration object for that version (start URLs, actions, index prefix, schedule, rate limit, and other crawler settings). Use this to inspect or restore an earlier configuration. Use list_config_versions to find version numbers and get_crawler for the current configuration. Requires a crawler id from list_crawlers and a version number from list_config_versions.
Inputs
idstringrequired- Unique ID (UUID) of the crawler. Use list_crawlers to find it. Example: e0f6db8a-24f5-4092-83a4-1b2c6cb6d809.
versionintegerrequired- Version number of the crawler configuration to retrieve. Version 1 is the initial configuration used when the crawler was created. Use list_config_versions to find valid numbers. Example: 3.
algoliacrawler_get_crawl_run_fileRetrieve the downloadable log file for a single crawl run of a crawler.Read-onlyGet Crawl Run File
Retrieve the downloadable log file for a single crawl run of a crawler. Returns a JSON object with a file string holding the log file content for that run. Use this to debug or audit one crawl run. Use list_crawl_runs first to find the log id, and get_url_stats for aggregate counts instead of per-run logs. Requires a crawler id from list_crawlers and a log id from list_crawl_runs.
Inputs
idstringrequired- Unique ID (UUID) of the crawler. Use list_crawlers to find it. Example: e0f6db8a-24f5-4092-83a4-1b2c6cb6d809.
logIdstringrequired- Unique ID (UUID) of the crawler log (crawl run) to download. Use list_crawl_runs to find it. Example: a2ebb507-ef64-4b6b-9d84-ef66baaa7a80.
algoliacrawler_get_crawlerGet the details of one Algolia crawler by id, optionally including its full configuration.Read-onlyGet Crawler
Get the details of one Algolia crawler by id, optionally including its full configuration. Returns the crawler name, created and updated timestamps, whether it is running, reindexing, or blocked (with the blocking error and blocking task id when blocked), and the last reindex start and end times; with the configuration option set, also returns the configuration. Use this for the state or configuration of a known crawler. Use list_crawlers to find ids and get_task_status to follow an asynchronous task. Requires a crawler id from list_crawlers or create_crawler.
Inputs
idstringrequired- Crawler ID (UUID) of the crawler to retrieve. Get it from the list crawlers tool or from the response of the create crawler tool. Example: e0f6db8a-24f5-4092-83a4-1b2c6cb6d809.
withConfigboolean- Set to true to include the crawler's full configuration in the response. Leave empty or false to return only status information such as name, timestamps, and running state. Example: true.
algoliacrawler_get_task_statusRetrieve the status of a crawler task, showing whether it is still pending or has completed.Read-onlyGet Task Status
Retrieve the status of a crawler task, showing whether it is still pending or has completed. Returns an object with a pending boolean. Use this to poll a task started by another crawler tool, such as a run or reindex, until pending is false. Use cancel_task to unblock a crawler whose task failed. Requires a crawler id from list_crawlers and a task id returned by the tool that started the task.
Inputs
idstringrequired- Unique ID (UUID) of the crawler. Use list_crawlers to find it. Example: e0f6db8a-24f5-4092-83a4-1b2c6cb6d809.
taskIDstringrequired- Unique ID (UUID) of the task to check. It is returned when a crawler action such as run, reindex, or crawl URLs is started. Example: 98458796-b7bb-4703-8b1b-785c1080b110.
algoliacrawler_get_url_statsRetrieve crawl statistics for a crawler, broken down by URL status.Read-onlyGet URL Stats
Retrieve crawl statistics for a crawler, broken down by URL status. Returns the total count of crawled URLs and a data array where each entry has a status (DONE, SKIPPED, or FAILED), a category (fetch, extraction, indexing, or success), a reason, a readable explanation, and the number of URLs with that status. Use this to see why URLs were skipped or failed. Use list_crawl_runs for per-run history and get_crawler for configuration. Requires a crawler id from list_crawlers.
Inputs
idstringrequired- Unique ID (UUID) of the crawler. Use list_crawlers to find it. Example: e0f6db8a-24f5-4092-83a4-1b2c6cb6d809.
algoliacrawler_list_config_versionsList the saved configuration versions of a crawler, including who authored each change.Read-onlyList Config Versions
List the saved configuration versions of a crawler, including who authored each change. Returns a paginated response with the current page, items per page, total count, and an items array of version number, creation time, and author id. Every configuration update adds a new version. Use this to find a version number before calling get_config_version. The list does not include the configuration itself. Requires a crawler id from list_crawlers.
Inputs
idstringrequired- Unique ID (UUID) of the crawler. Use list_crawlers to find it. Example: e0f6db8a-24f5-4092-83a4-1b2c6cb6d809.
itemsPerPageinteger- Number of configuration versions to return per page, from 1 to 100. Defaults to 20. Use together with page to walk through large result sets. Example: 20.default
20 pageinteger- Page of results to retrieve, from 1 to 100. Defaults to 1. Compare the total in the response with itemsPerPage to decide whether more pages exist. Example: 1.default
1
algoliacrawler_list_crawl_runsList the recorded crawl runs (crawler logs) for a crawler, optionally filtered by date range or URL status.Read-onlyList Crawl Runs
List the recorded crawl runs (crawler logs) for a crawler, optionally filtered by date range or URL status. Returns a logs array and a meta object with the total number of matching records. Each log has its id, config id, reindex id, start and completion times, file sizes, expiry, status, access count, and counts of done, skipped, and failed URLs. Use this to find a log id before calling get_crawl_run_file or delete_crawl_runs. Use get_url_stats for aggregate URL status counts across the crawler. Requires a crawler id from list_crawlers.
Inputs
idstringrequired- Unique ID (UUID) of the crawler. Use list_crawlers to find it. Example: e0f6db8a-24f5-4092-83a4-1b2c6cb6d809.
fromstring- Only return runs on or after this date, as a Unix timestamp in seconds, passed as a string. Leave empty for no lower bound. Example: 1762264044.
limitinteger- Maximum number of runs to return, from 1 to 1000. Defaults to 10. Use together with offset to page through results. Example: 10.default
10 offsetinteger- Number of runs to skip before returning results, used for pagination together with limit. Leave empty to start from the first run. Example: 10.
orderstring- Sort direction of the results, either ASC or DESC. Leave empty to use the server default order. Example: DESC.one of
ASCDESC statusstring- Filter runs by crawled URL status. Must be one of DONE, SKIPPED, or FAILED. Leave empty to return runs regardless of status. Example: FAILED.one of
DONESKIPPEDFAILED untilstring- Only return runs on or before this date, as a Unix timestamp in seconds, passed as a string. Leave empty for no upper bound. Example: 1762264044.
algoliacrawler_list_crawlersList the Algolia crawlers on the account, optionally filtered by crawler name or Algolia application ID.Read-onlyList Crawlers
List the Algolia crawlers on the account, optionally filtered by crawler name or Algolia application ID. Returns a paginated response with the current page, items per page, total count, and an items array of crawler ids and names. Use this to find a crawler id before calling a tool that needs one. The list only carries id and name, so use the get-crawler tool for configuration details.
Inputs
itemsPerPageinteger- Number of crawlers to return per page, from 1 to 100. Defaults to 20. Use together with page to walk through large result sets.default
20 namestring- Crawler name used to filter the response, up to 64 characters. Leave empty to list crawlers regardless of name. Example: test-crawler.
pageinteger- Page of results to retrieve, from 1 to 100. Defaults to 1. Compare the total in the response with itemsPerPage to decide whether more pages exist.default
1
algoliacrawler_list_domainsList the domains registered for crawling on the account, optionally filtered by Algolia application ID.Read-onlyList Domains
List the domains registered for crawling on the account, optionally filtered by Algolia application ID. Returns a paginated response with the current page, items per page, total count, and an items array of domain name, Algolia application ID, and whether the domain is validated. Use this to check which domains crawlers are allowed to run against, since crawlers only run if their URLs match a registered domain. Use list_crawlers to list crawlers instead.
Inputs
itemsPerPageinteger- Number of domains to return per page, from 1 to 100. Defaults to 20. Use together with page to walk through large result sets. Example: 20.default
20 pageinteger- Page of results to retrieve, from 1 to 100. Defaults to 1. Compare the total in the response with itemsPerPage to decide whether more pages exist. Example: 1.default
1
algoliacrawler_cancel_taskCancel a blocking task on a crawler so its schedule can resume.WriteCancel Task
Cancel a blocking task on a crawler so its schedule can resume. Returns an empty successful response when the task is cancelled. Use this when a task ran into an error and is blocking the crawler's schedule. Use get_task_status to check whether a task is still pending before cancelling. Requires a crawler id from list_crawlers and the id of the blocking task.
Inputs
idstringrequired- Unique ID (UUID) of the crawler. Use list_crawlers to find it. Example: e0f6db8a-24f5-4092-83a4-1b2c6cb6d809.
taskIDstringrequired- Unique ID (UUID) of the blocking task to cancel. Tasks that ran into an error block the crawler schedule until cancelled. Example: 98458796-b7bb-4703-8b1b-785c1080b110.
algoliacrawler_crawl_urlsCrawl the specified URLs, extract records from them, and add the records to the crawler's index.WriteCrawl URLs
Crawl the specified URLs, extract records from them, and add the records to the crawler's index. Returns a taskId acknowledging the action. Use this to refresh specific pages without a full crawl; use start_reindex for a full crawl and test_url to preview extraction without indexing. If a crawl is already running, the records go to a temporary index. This operation is rate limited to 500 requests every 24 hours. Requires a crawler id from list_crawlers.
Inputs
idstringrequired- Crawler ID (UUID) of the crawler to crawl URLs with. Get it from the list crawlers tool or from the response of the create crawler tool. Example: e0f6db8a-24f5-4092-83a4-1b2c6cb6d809.
urlsarrayrequired- List of full URLs to crawl, each including the scheme. Provide at least one. Example: ["https://www.algolia.com/products/crawler/"].
saveboolean- Whether to add the URLs to the crawler configuration's extra URLs list. Set true to always save them, false to never save them. If left empty, the URLs are saved only if they were not indexed during the last reindex. Example: true.
algoliacrawler_create_crawlerCreate a new Algolia crawler from a name and a full configuration.WriteCreate Crawler
Create a new Algolia crawler from a name and a full configuration. Returns the id (UUID) of the new crawler. Use this to set up a crawler that does not exist yet. Use list_crawlers first to check for an existing crawler with the same name, and update_crawler or update_crawler_config to change one that exists. To start crawling afterwards, use start_reindex.
Inputs
configobjectrequired- Full crawler configuration object. A configuration must name the Algolia application the crawler writes to, a rate limit from 1 to 100 that sets how many concurrent tasks run per second (start low, for example 2 to 4), and a non-empty list of actions (up to 30). Each action names the target index and carries a record extractor, which is an object marked as a function whose source is a JavaScript function string that turns a crawled page into Algolia records; an action can also limit which URLs it applies to with path patterns. Add start URLs or sitemaps so the crawler knows where to begin. The Algolia API key is optional on creation (the Crawler generates one if omitted) and must never be an Admin API key. Minimal working example: {"appId": "YourApplicationID", "apiKey": "YourCrawlerIndexingApiKey", "rateLimit": 4, "startUrls": ["https://www.example.com"], "actions": [{"indexName": "example_index", "pathsToMatch": ["https://www.example.com/**"], "recordExtractor": {"__type": "function", "source": "({ url, $ }) => [{ objectID: url.href, title: $('head title').text() }]"}}]}
namestringrequired- Name for the new crawler, up to 64 characters. Use a short descriptive label so it is easy to find with the list crawlers tool. Example: test-crawler.
algoliacrawler_pause_crawlerPause the specified crawler.WritePause Crawler
Pause the specified crawler. Returns a taskId acknowledging the action. Use this to stop a crawler temporarily; use run_crawler to resume it. Use delete_crawler to remove a crawler that is no longer needed. Requires a crawler id from list_crawlers.
Inputs
idstringrequired- Crawler ID (UUID) of the crawler to pause. Get it from the list crawlers tool or from the response of the create crawler tool. Example: e0f6db8a-24f5-4092-83a4-1b2c6cb6d809.
algoliacrawler_run_crawlerUnpause the specified crawler so it resumes work.WriteUnpause Crawler
Unpause the specified crawler so it resumes work. Returns a taskId acknowledging the action. Use this to undo pause_crawler: previously ongoing crawls resume, otherwise the crawler waits for its next scheduled run. Use start_reindex to begin a new crawl right away. Requires a crawler id from list_crawlers.
Inputs
idstringrequired- Crawler ID (UUID) of the crawler to unpause. Get it from the list crawlers tool or from the response of the create crawler tool. Example: e0f6db8a-24f5-4092-83a4-1b2c6cb6d809.
algoliacrawler_start_reindexStart or resume a crawl of the crawler's configured URLs.WriteStart Reindex
Start or resume a crawl of the crawler's configured URLs. Returns a taskId acknowledging the action. Use this to run a full crawl now instead of waiting for the schedule. Use crawl_urls to crawl only specific URLs, and run_crawler only to unpause a paused crawler. Requires a crawler id from list_crawlers.
Inputs
idstringrequired- Crawler ID (UUID) of the crawler to start a crawl for. Get it from the list crawlers tool or from the response of the create crawler tool. Example: e0f6db8a-24f5-4092-83a4-1b2c6cb6d809.
algoliacrawler_test_urlTest crawl one URL with the crawler's configuration, optionally with configuration overrides, and show the records it would extract.WriteTest URL
Test crawl one URL with the crawler's configuration, optionally with configuration overrides, and show the records it would extract. Returns the test start and end dates, extraction logs, the extracted records grouped by index name, links found on the page, any external data, and an error if one occurred. Use this to preview a configuration change before saving it with update_crawler_config. It does not add records to the index; use crawl_urls to actually index URLs. Requires a crawler id from list_crawlers.
Inputs
idstringrequired- Crawler ID (UUID) of the crawler to test with. Get it from the list crawlers tool or from the response of the create crawler tool. Example: e0f6db8a-24f5-4092-83a4-1b2c6cb6d809.
urlstringrequired- Full URL of the web page to test crawl, including the scheme. It should match the crawler's configured start URLs or path patterns. Example: https://www.algolia.com/blog.
configobject- Optional configuration overrides to try during the test, as a JSON object. Only top-level configuration properties can be overridden: to change something nested, such as a record extractor inside actions, send the complete top-level property (the whole actions list). The published API schema lists the application ID, rate limit, and actions as required for a configuration, so if a partial override is rejected, send a complete configuration. Leave empty to test with the crawler's saved configuration. Example overriding the actions: {"actions": [{"indexName": "example_index", "pathsToMatch": ["https://www.example.com/**"], "recordExtractor": {"__type": "function", "source": "({ url, $ }) => [{ objectID: url.href, title: $('head title').text() }]"}}]}
algoliacrawler_update_crawlerRename a crawler or replace its entire configuration with a new one.WriteUpdate Crawler
Rename a crawler or replace its entire configuration with a new one. Returns a taskId for the asynchronous update. Use this to rename a crawler, or to overwrite the whole configuration when versioning is not needed. Use update_crawler_config instead to change individual top-level settings with a new configuration version recorded. Requires a crawler id from list_crawlers.
- Idempotent
Inputs
idstringrequired- Crawler ID (UUID) of the crawler to update. Get it from the list crawlers tool or from the response of the create crawler tool. Example: e0f6db8a-24f5-4092-83a4-1b2c6cb6d809.
configobject- Complete replacement configuration for the crawler. Because this operation replaces the whole configuration, send every setting you want to keep, not only the ones you are changing; anything omitted is dropped. Changes made here are not versioned. A configuration must name the Algolia application the crawler writes to, a rate limit from 1 to 100 that sets how many concurrent tasks run per second (start low, for example 2 to 4), and a non-empty list of actions (up to 30). Each action names the target index and carries a record extractor, which is an object marked as a function whose source is a JavaScript function string that turns a crawled page into Algolia records; an action can also limit which URLs it applies to with path patterns. Add start URLs or sitemaps so the crawler knows where to begin. The Algolia API key is optional on creation (the Crawler generates one if omitted) and must never be an Admin API key. Minimal working example: {"appId": "YourApplicationID", "apiKey": "YourCrawlerIndexingApiKey", "rateLimit": 4, "startUrls": ["https://www.example.com"], "actions": [{"indexName": "example_index", "pathsToMatch": ["https://www.example.com/**"], "recordExtractor": {"__type": "function", "source": "({ url, $ }) => [{ objectID: url.href, title: $('head title').text() }]"}}]}
namestring- New name for the crawler, up to 64 characters. Provide this to rename the crawler; leave empty to keep the current name. Example: test-crawler.
algoliacrawler_update_crawler_configUpdate top-level settings of a crawler's configuration, creating a new configuration version.WriteUpdate Crawler Configuration
Update top-level settings of a crawler's configuration, creating a new configuration version. Returns a taskId for the asynchronous update. Use this for versioned configuration changes. Use update_crawler to rename a crawler or replace its whole configuration without versioning, and list_config_versions to review earlier versions. Requires a crawler id from list_crawlers.
Inputs
configobjectrequired- Configuration properties to update, as a JSON object. Only top-level configuration properties can be changed: to change something nested, such as a record extractor inside actions, send the complete top-level property (the whole actions list). Properties you do not include are left as they are. The published API schema lists the application ID, rate limit, and actions as required for a configuration, so if a partial update is rejected, send the full configuration including them. Each update creates a new configuration version. Example: {"rateLimit": 8, "maxUrls": 5000}
idstringrequired- Crawler ID (UUID) of the crawler to update the configuration of. Get it from the list crawlers tool or from the response of the create crawler tool. Example: e0f6db8a-24f5-4092-83a4-1b2c6cb6d809.
algoliacrawler_delete_crawl_runsDelete the recorded crawl run logs with the given log ids from a crawler.DestructiveDelete Crawl Runs
Delete the recorded crawl run logs with the given log ids from a crawler. This removes the log records only, not the crawler itself. Returns an array of the log ids that were submitted. Use this to clean up old run logs. Use list_crawl_runs to find log ids, and delete_crawler to remove the crawler. Deletion cannot be undone. Requires a crawler id from list_crawlers and one or more log ids from list_crawl_runs.
Inputs
idstringrequired- Unique ID (UUID) of the crawler. Use list_crawlers to find it. Example: e0f6db8a-24f5-4092-83a4-1b2c6cb6d809.
logIdsarrayrequired- List of crawler log (crawl run) UUIDs to delete. Use list_crawl_runs to find them. Example: ["a2ebb507-ef64-4b6b-9d84-ef66baaa7a80"].
algoliacrawler_delete_crawlerDelete the specified Algolia crawler.DestructiveDelete Crawler
Delete the specified Algolia crawler. Returns a taskId acknowledging the deletion. Use this only when the crawler is no longer needed. Use pause_crawler to stop a crawler temporarily instead. Requires a crawler id from list_crawlers.
- Idempotent
Inputs
idstringrequired- Crawler ID (UUID) of the crawler to delete. Get it from the list crawlers tool or from the response of the create crawler tool. Example: e0f6db8a-24f5-4092-83a4-1b2c6cb6d809.
No tools match.