Skip to content

Go API Reference

Convert markdown links to numbered citations.

[Example](https://example.com) becomes Example[1] with [1]: <https://example.com> in the reference list. Images ![alt](url) are preserved unchanged.

Signature:

func GenerateCitations(markdown string) CitationResult

Example:

result := GenerateCitations("value")

Parameters:

Name Type Required Description
Markdown string Yes The markdown

Returns: CitationResult


Create a new crawl engine with the given configuration.

If config is nil, uses CrawlConfig.default(). Returns an error if the configuration is invalid.

Signature:

func CreateEngine(config CrawlConfig) (CrawlEngineHandle, error)

Example:

result, err := CreateEngine(CrawlConfig{})
if err != nil {
return err
}

Parameters:

Name Type Required Description
Config *CrawlConfig No The configuration options

Returns: CrawlEngineHandle

Errors: Returns error.


Scrape a single URL, returning extracted page data.

Signature:

func Scrape(engine CrawlEngineHandle, url string) (ScrapeResult, error)

Example:

result, err := Scrape(CrawlEngineHandle{}, "value")
if err != nil {
return err
}

Parameters:

Name Type Required Description
Engine CrawlEngineHandle Yes The crawl engine handle
Url string Yes The URL to fetch

Returns: ScrapeResult

Errors: Returns error.


Crawl a website starting from url, following links up to the configured depth.

Signature:

func Crawl(engine CrawlEngineHandle, url string) (CrawlResult, error)

Example:

result, err := Crawl(CrawlEngineHandle{}, "value")
if err != nil {
return err
}

Parameters:

Name Type Required Description
Engine CrawlEngineHandle Yes The crawl engine handle
Url string Yes The URL to fetch

Returns: CrawlResult

Errors: Returns error.


Discover all pages on a website by following links and sitemaps.

Signature:

func MapUrls(engine CrawlEngineHandle, url string) (MapResult, error)

Example:

result, err := MapUrls(CrawlEngineHandle{}, "value")
if err != nil {
return err
}

Parameters:

Name Type Required Description
Engine CrawlEngineHandle Yes The crawl engine handle
Url string Yes The URL to fetch

Returns: MapResult

Errors: Returns error.


Execute browser actions on a single page.

Signature:

func Interact(engine CrawlEngineHandle, url string, actions []PageAction) (InteractionResult, error)

Example:

result, err := Interact(CrawlEngineHandle{}, "value", nil)
if err != nil {
return err
}

Parameters:

Name Type Required Description
Engine CrawlEngineHandle Yes The crawl engine handle
Url string Yes The URL to fetch
Actions \[\]PageAction Yes The actions

Returns: InteractionResult

Errors: Returns error.


Scrape multiple URLs concurrently.

Signature:

func BatchScrape(engine CrawlEngineHandle, urls []string) (BatchScrapeResults, error)

Example:

result, err := BatchScrape(CrawlEngineHandle{}, nil)
if err != nil {
return err
}

Parameters:

Name Type Required Description
Engine CrawlEngineHandle Yes The crawl engine handle
Urls \[\]string Yes The urls

Returns: BatchScrapeResults

Errors: Returns error.


Crawl multiple seed URLs concurrently, each following links to configured depth.

Signature:

func BatchCrawl(engine CrawlEngineHandle, urls []string) (BatchCrawlResults, error)

Example:

result, err := BatchCrawl(CrawlEngineHandle{}, nil)
if err != nil {
return err
}

Parameters:

Name Type Required Description
Engine CrawlEngineHandle Yes The crawl engine handle
Urls \[\]string Yes The urls

Returns: BatchCrawlResults

Errors: Returns error.


Result from a single page action execution.

Field Type Default Description
ActionIndex int Zero-based index of the action in the sequence.
ActionType string The type of action that was executed.
Success bool Whether the action completed successfully.
Data *interface{} nil Action-specific return data (screenshot bytes, JS return value, scraped HTML).
Error *string nil Error message if the action failed.

Article metadata extracted from article:* Open Graph tags.

Field Type Default Description
PublishedTime *string nil The article publication time.
ModifiedTime *string nil The article modification time.
Author *string nil The article author.
Section *string nil The article section.
Tags \[\]string nil The article tags.

Result from a single URL in a batch crawl operation.

Field Type Default Description
Url string The seed URL that was crawled.
Result *CrawlResult nil The crawl result, if successful.
Error *string nil The error message, if the crawl failed.

Aggregate result of a batch crawl, exposing per-URL results plus precomputed counts.

The counts are derived once at construction so every binding language can read them as plain integer fields without re-iterating the results vector.

Field Type Default Description
Results \[\]BatchCrawlResult nil Per-URL crawl results, in the order seed URLs were submitted.
TotalCount int Total number of seed URLs in the batch (equal to results.len()).
CompletedCount int Number of seed URLs whose crawl succeeded (error is nil).
FailedCount int Number of seed URLs whose crawl failed (error is Some).

Request to begin a multi-URL streaming crawl.

Wraps a set of seed URLs for delivery through the streaming-adapter binding surface. Required as a struct because alef’s streaming adapter requires a named request type — primitives are not supported.

Field Type Default Description
Urls \[\]string nil The seed URLs to crawl. Each URL is followed independently up to the engine’s configured depth.

Result from a single URL in a batch scrape operation.

Field Type Default Description
Url string The URL that was scraped.
Result *ScrapeResult nil The scrape result, if successful.
Error *string nil The error message, if the scrape failed.

Aggregate result of a batch scrape, exposing per-URL results plus precomputed counts.

The counts are derived once at construction so every binding language can read them as plain integer fields without re-iterating the results vector.

Field Type Default Description
Results \[\]BatchScrapeResult nil Per-URL scrape results, in the order URLs were submitted.
TotalCount int Total number of URLs in the batch (equal to results.len()).
CompletedCount int Number of URLs whose scrape succeeded (error is nil).
FailedCount int Number of URLs whose scrape failed (error is Some).

Browser fallback configuration.

Field Type Default Description
Mode BrowserMode BrowserMode.Auto When to use the headless browser fallback.
Backend BrowserBackend BrowserBackend.Chromiumoxide Browser backend used to render JavaScript-heavy pages.
Endpoint *string nil CDP WebSocket endpoint for connecting to an external browser instance.
Timeout time.Duration 30000ms Timeout for browser page load and rendering (in milliseconds when serialized).
Wait BrowserWait BrowserWait.NetworkIdle Wait strategy after browser navigation.
WaitSelector *string nil CSS selector to wait for when wait is Selector.
ExtraWait *time.Duration nil Extra time to wait after the wait condition is met.
Proxy *ProxyConfig nil Proxy for browser fetches. Overrides CrawlConfig.proxy when set. Native backend supports http/https only (no SOCKS5).
BlockUrlPatterns \[\]string nil URL patterns to block before the network request fires. Supports * wildcards. Useful for skipping ads/analytics/large images. Honored by BrowserBackend.Native; chromiumoxide ignores this field today.
EvalScript *string nil JavaScript snippet evaluated after navigation completes. Scraping captures the native backend result in ScrapeResult.browser.eval_result. Interactions run this script before page actions on both browser backends but do not include the script result in InteractionResult.
RobotsUserAgent *string nil User-agent used when fetching robots.txt. Defaults to BrowserConfig.user_agent (or crawlberg’s default) if unset. Native only.
CaptureNetworkEvents bool false Capture the full network event stream into the result. Default false (only the document event is captured). Native only.
SessionAffinity bool true Enable session affinity: reuse chromiumoxide Pages for same-domain requests so cookies + fingerprint + solved challenges persist. Default: true. When false, each request gets a fresh Page.

Signature:

func (o *BrowserConfig) Default() BrowserConfig

Example:

result := BrowserConfig.Default()

Returns: BrowserConfig


Browser-specific extras populated when the native browser backend was used.

Available on ScrapeResult.browser when BrowserBackend.Native handled the request.

Field Type Default Description
EvalResult *interface{} nil Return value of BrowserConfig.eval_script, if provided.
NetworkEvents \[\]ResponseMeta nil Network events captured during page navigation (only populated when BrowserConfig.capture_network_events is true).
Cookies \[\]CookieInfo nil All non-expired cookies present in the browser’s cookie jar after navigation completes (includes both prior cookies and server Set-Cookie).

A single numbered reference in a citation list — produced by the citation extractor when content uses inline [N]-style markers.

Field Type Default Description
Index int 1-based reference number as it appears in the source text.
Url string Resolved absolute URL for this reference.
Text string Human-readable anchor text or title for the reference.

Result of citation conversion.

Field Type Default Description
Content string Markdown with links replaced by numbered citations.
References \[\]CitationReference nil Numbered reference list: (index, url, text).

Content extraction and conversion configuration.

Controls how HTML is converted to the output format. Uses html-to-markdown-rs as the conversion engine for all formats (markdown, plain text, djot).

Field Type Default Description
OutputFormat string "markdown" Output format: "markdown" (default), "plain", "djot".
PreprocessingPreset string "standard" Preprocessing aggressiveness: "minimal", "standard" (default), "aggressive". - Minimal: only scripts/styles removed. - Standard: also removes nav, nav-hinted headers/footers/asides, forms. - Aggressive: removes all footers/asides unconditionally.
RemoveNavigation bool true Remove navigation elements (nav, breadcrumbs, menus). Default: true.
RemoveForms bool true Remove form elements. Default: true.
StripTags \[\]string nil HTML tag names to strip (render children only, remove the tag wrapper). Default: \["noscript"\].
PreserveTags \[\]string nil HTML tag names to preserve as raw HTML in output.
ExcludeSelectors \[\]string nil CSS selectors for elements to exclude entirely (element + all content). Unlike strip_tags (which removes the wrapper but keeps children), excluded elements and all descendants are dropped. Supports CSS selectors: .class, #id, \[attribute\], compound selectors. Example: \[".cookie-banner", "#ad-container", "\[role='complementary'\]"\]
SkipImages bool false Skip image elements in output. Default: false.
MaxDepth *int nil Max DOM traversal depth. Prevents stack overflow on deeply nested HTML.
Wrap bool false Enable line wrapping. Default: false.
WrapWidth int 80 Wrap width when wrap is enabled. Default: 80.
IncludeDocumentStructure bool true Include document structure tree in output. Default: true.

Signature:

func (o *ContentConfig) Default() ContentConfig

Example:

result := ContentConfig.Default()

Returns: ContentConfig


Information about an HTTP cookie received from a response.

Field Type Default Description
Name string The cookie name.
Value string The cookie value.
Domain *string nil The cookie domain, if specified.
Path *string nil The cookie path, if specified.

Configuration for crawl, scrape, and map operations.

Field Type Default Description
MaxDepth *int nil Maximum crawl depth (number of link hops from the start URL).
MaxPages *int nil Maximum number of pages to crawl.
MaxLinksPerPage *int nil Maximum links enqueued from a single page. Defaults to 10000. Bounds the work one hostile or pathological page can create; links past the cap are dropped and a warning is logged.
MaxConcurrent *int nil Maximum number of concurrent requests.
RespectRobotsTxt bool false Whether to respect robots.txt directives.
SoftHttpErrors bool false When true, HTTP-level error responses (404 NotFound, 403 Forbidden, WAF blocks) are surfaced as ScrapeResult records with the matching status_code rather than raised as CrawlError. Default false preserves the historical throw-on-error contract for direct fetches. Independently of this flag, 404s reached at the end of a redirect chain are always surfaced softly — the user opted into redirect-following, so receiving a 404 there is part of the normal flow rather than an unexpected error.
UserAgent *string nil Custom user-agent string.
StayOnDomain bool false Whether to restrict crawling to the same domain.
AllowSubdomains bool false Whether to allow subdomains when stay_on_domain is true.
IncludePaths \[\]string nil Regex patterns for paths to include during crawling.
ExcludePaths \[\]string nil Regex patterns for paths to exclude during crawling.
CustomHeaders map\[string\]string nil Custom HTTP headers to send with each request.
RequestTimeout time.Duration 30000ms Timeout for individual HTTP requests (in milliseconds when serialized).
RateLimitMs *uint64 nil Per-domain rate limit in milliseconds. When set, enforces a minimum delay between requests to the same domain. Defaults to 200ms when nil.
MaxRedirects int 10 Maximum number of redirects to follow.
RetryCount int 0 Number of retry attempts for failed requests.
RetryCodes \[\]uint16 nil HTTP status codes that should trigger a retry.
CookiesEnabled bool false Whether to enable cookie handling.
Auth *AuthConfig nil Authentication configuration.
MaxBodySize *int nil Maximum response body size in bytes.
RemoveTags \[\]string nil CSS selectors for tags to remove from HTML before processing.
Content ContentConfig Content extraction and conversion configuration.
MapLimit *int nil Maximum number of URLs to return from a map operation.
MapSearch *string nil Search filter for map results (case-insensitive substring match on URLs).
DownloadAssets bool false Whether to download assets (CSS, JS, images, etc.) from the page.
AssetTypes \[\]AssetCategory nil Filter for asset categories to download.
MaxAssetSize *int nil Maximum size in bytes for individual asset downloads.
Browser BrowserConfig Browser configuration.
Proxy *ProxyConfig nil Proxy configuration for HTTP requests.
UserAgents \[\]string nil List of user-agent strings for rotation. If non-empty, overrides user_agent.
CaptureScreenshot bool false Whether to capture a screenshot when using the browser. Only supported by scrape() with BrowserBackend.Chromiumoxide and BrowserMode.Always or Stealth. A screenshot is 100–500 KB of PNG per page, so crawl() does not carry screenshots in CrawlPageResult/CrawlResult at all — a multi-thousand-page crawl holding one per page in memory is not a safe default. Setting this with any other configuration (a different backend, BrowserMode.Auto/Never, or during crawl()) has no effect and logs a warning rather than silently doing nothing.
FollowDocumentUrls bool false Re-enqueue discovered LinkType.Document URLs into the crawl frontier so the crawl follows links from document pages (PDFs, etc.) as it would from HTML pages. Default: false (documents terminate at materialisation).
DocumentUrlDepth *uint32 nil Maximum document-depth (from the seed URL through document links only) when follow_document_urls is true. nil means inherit max_depth. Independent of max_depth: a document URL is enqueued only if BOTH the outer max_depth and (if set) document_url_depth permit it.
DownloadDocuments bool true Whether to download non-HTML documents (PDF, DOCX, images, code, etc.) instead of skipping them. Defaults to true — unlike download_assets and capture_screenshot, which default to false.
DocumentMaxSize *int 52428800 Maximum size in bytes for document downloads. Defaults to 50 MB.
DocumentMimeTypes \[\]string nil Allowlist of MIME types to download. If empty, uses built-in defaults.
DocumentOutputDir *string nil Directory to stream downloaded document bytes into instead of holding them in memory on DownloadedDocument.content. When set, content is left empty and DownloadedDocument.content_path is populated with <dir>/<content_hash>.<ext>. nil (default) preserves today’s in-memory-only behavior. Has no effect on wasm32, which has no filesystem — use document_content_encoding there instead.
DocumentContentEncoding *DocumentContentEncoding nil Opt-in encoding that duplicates DownloadedDocument.content into a serializable field for language bindings that need the bytes in-memory (content itself is alef(skip)ed). nil (default) means no encoding is produced. Independent of document_output_dir — set both to get a file on disk and an in-memory copy.
WarcOutput *string nil Path to write WARC output. If nil, WARC output is disabled.
BrowserProfile *string nil Named browser profile for persistent sessions (cookies, localStorage). Chromiumoxide backend only. The native backend runs an in-process JavaScript engine with no Chrome process and therefore no profile directory, so this is ignored there and logs a warning. It is also ignored — with a warning — when a shared browser pool is in use (the pool launches before any per-crawl config exists) or when connecting to an external CDP endpoint whose process crawlberg does not own.
SaveBrowserProfile bool false Whether to save changes back to the browser profile on exit.
Ssrf SsrfPolicy SSRF policy for outbound network requests. Default: deny private networks, allow http/https only, max 5 redirects. deny_private, allowlist and max_redirects are exposed to all language bindings. scheme_allowlist stays Rust-only — see SsrfPolicy. wasm32 (including Node.js): deny_private does not stop hostname-based requests. There is no DNS resolution on this target, so only a literal IP host is checked against the policy — a domain name is always permitted, regardless of deny_private. Under Node, where fetch enforces no CORS, this means a service embedding the wasm binding can be driven to internal hosts by domain name even with deny_private = true. Enforce egress restrictions at the network layer for that deployment target; do not rely on this field. See crawlberg.net.validate_url.
SsrfDenyPrivateExplicit *bool nil Pins SsrfPolicy.deny_private to a caller-chosen value, bypassing the CRAWLBERG_ALLOW_PRIVATE_NETWORK operator override entirely for this config. ssrf.deny_private is a plain, always-serialized bool: several alef-generated bindings construct SsrfPolicy.default() (hardcoding deny_private: true) whenever their caller never touches SSRF settings at all, so true on that field alone cannot distinguish “the caller wants private networks denied” from “the binding’s own structural default landed on true”. The environment variable exists precisely to resolve that ambiguity in the common case by treating any true as inconclusive and deferring to the operator. Set this field when that default-deferral is wrong for your call — e.g. a test that must prove deny_private: true still denies even while the operator has set CRAWLBERG_ALLOW_PRIVATE_NETWORK suite-wide for every other call. nil (default) preserves today’s behavior: the environment variable may still flip ssrf.deny_private to false. Some(value) pins ssrf.deny_private to value and the environment variable is not consulted for this config.

Signature:

func (o *CrawlConfig) Default() CrawlConfig

Example:

result := CrawlConfig.Default()

Returns: CrawlConfig

Validate the configuration, returning an error if any values are invalid.

Signature:

func (o *CrawlConfig) Validate() error

Example:

if err := instance.Validate(); err != nil {
return err
}

Returns: No return value.

Errors: Returns error.


Opaque handle to a configured crawl engine.

Constructed via create_engine with an optional CrawlConfig. Default implementations for all pluggable components are used internally.


The result of crawling a single page during a crawl operation.

Field Type Default Description
Url string The original URL of the page.
NormalizedUrl string The normalized URL of the page.
StatusCode uint16 The HTTP status code of the response.
ContentType string The Content-Type header value.
Html string The HTML body of the response.
BodySize int The size of the response body in bytes.
Metadata PageMetadata Extracted metadata from the page.
Links \[\]LinkInfo nil Links found on the page.
Images \[\]ImageInfo nil Images found on the page.
Feeds \[\]FeedInfo nil Feed links found on the page.
JsonLd \[\]JsonLdEntry nil JSON-LD entries found on the page.
Depth int The depth of this page from the start URL.
StayedOnDomain bool Whether this page is on the same domain as the start URL.
WasSkipped bool Whether this page was skipped (binary or PDF content).
IsPdf bool Whether the content is a PDF.
DetectedCharset *string nil The detected character set encoding.
Markdown *MarkdownResult nil Markdown conversion of the page content.
ExtractedData *interface{} nil Structured data extracted by LLM. Populated when extraction is configured.
ExtractionMeta *ExtractionMeta nil Metadata about the LLM extraction pass (cost, tokens, model).
DownloadedDocument *DownloadedDocument nil Downloaded non-HTML document (PDF, DOCX, image, code, etc.).
BrowserUsed bool Whether the browser fallback was used to fetch this page.

The result of a multi-page crawl operation.

Field Type Default Description
Pages \[\]CrawlPageResult nil The list of crawled pages.
FinalUrl string The final URL after following redirects.
RedirectCount int The number of redirects followed.
WasSkipped bool Whether any page was skipped during crawling.
Error *string nil An error message, if the crawl encountered an issue.
Cookies \[\]CookieInfo nil Cookies collected during the crawl.
StayedOnDomain bool Whether all crawled pages stayed on the same domain as the start URL.
BrowserUsed bool Whether the browser fallback was used for any page in this crawl.

Returns the count of unique normalized URLs encountered during crawling.

Computed from pages (not the deprecated normalized_urls field) so it is correct across every binding that reconstructs CrawlResult from pages alone. In streaming mode pages is empty, so this returns 0 on the opaque-handle (C/Go/C#/Zig/Dart) path where it previously counted streamed pages — a known, accepted cost of making the other ten binding families correct.

Signature:

func (o *CrawlResult) UniqueNormalizedUrls() int

Example:

result := instance.UniqueNormalizedUrls()

Returns: int


Request to begin a single-URL streaming crawl.

Wraps a single seed URL for delivery through the streaming-adapter binding surface. Required as a struct because alef’s streaming adapter requires a named request type — primitives are not supported.

Field Type Default Description
Url string The seed URL to crawl.

A downloaded asset from a page.

Field Type Default Description
Url string The original URL of the asset.
ContentHash string The SHA-256 content hash of the asset.
MimeType *string nil The MIME type from the Content-Type header.
Size int The size of the asset in bytes.
AssetCategory AssetCategory AssetCategory.Image The category of the asset.
HtmlTag *string nil The HTML tag that referenced this asset (e.g., “link”, “script”, “img”).

A downloaded non-HTML document (PDF, DOCX, image, code file, etc.).

When the crawler encounters non-HTML content and download_documents is enabled, it downloads the raw bytes and populates this struct instead of skipping the resource.

Field Type Default Description
Url string The URL the document was fetched from.
MimeType string The MIME type from the Content-Type header.
Size int Size of the document in bytes.
Filename *string nil Filename extracted from Content-Disposition or URL path.
ContentHash string SHA-256 hex digest of the content.
Headers map\[string\]string nil Selected response headers.
Truncated bool True when content (or the file at content_path) was truncated to document_max_size; size still reports the original, untruncated length.
ContentPath *string nil Filesystem path the document was streamed to when document_output_dir was set. content is empty in memory when this is populated.
ContentBase64 *string nil Base64-encoded copy of content, populated only when document_content_encoding was set to Base64.

Metadata about an LLM extraction pass.

Field Type Default Description
Cost *float64 nil Estimated cost of the LLM call in USD.
PromptTokens *uint64 nil Number of prompt (input) tokens consumed.
CompletionTokens *uint64 nil Number of completion (output) tokens generated.
Model *string nil The model identifier used for extraction.
ChunksProcessed int Number of content chunks sent to the LLM.

Information about a favicon or icon link.

Field Type Default Description
Url string The icon URL.
Rel string The rel attribute (e.g., “icon”, “apple-touch-icon”).
Sizes *string nil The sizes attribute, if present.
MimeType *string nil The MIME type, if present.

Information about a feed link found on a page.

Field Type Default Description
Url string The feed URL.
Title *string nil The feed title, if present.
FeedType FeedType FeedType.Rss The type of feed.

A heading element extracted from the page.

Field Type Default Description
Level uint8 The heading level (1-6).
Text string The heading text content.

An hreflang alternate link entry.

Field Type Default Description
Lang string The language code (e.g., “en”, “fr”, “x-default”).
Url string The URL for this language variant.

Information about an image found on a page.

Field Type Default Description
Url string The image URL.
Alt *string nil The alt text, if present.
Width *uint32 nil The width attribute, if present and parseable.
Height *uint32 nil The height attribute, if present and parseable.
Source ImageSource ImageSource.Img The source of the image reference.

Result of executing a sequence of page interaction actions.

Field Type Default Description
ActionResults \[\]ActionResult nil Results from each executed action.
FinalHtml string Final page HTML after all actions completed.
FinalUrl string Final page URL (may have changed due to navigation).
ScreenshotBase64 *string nil Base64-encoded PNG screenshot taken after all actions. Populated only when a PageAction.Screenshot action actually ran, so callers that never request a screenshot do not pay the encoding cost.

A JSON-LD structured data entry found on a page.

Field Type Default Description
SchemaType string The @type value from the JSON-LD object.
Name *string nil The name value, if present.
Raw string The raw JSON-LD string.

Information about a link found on a page.

Field Type Default Description
Url string The resolved URL of the link.
Text string The visible text of the link.
LinkType LinkType LinkType.Internal The classification of the link.
Rel *string nil The rel attribute value, if present.
Nofollow bool Whether the link has rel="nofollow".

The result of a map operation, containing discovered URLs.

Field Type Default Description
Urls \[\]SitemapUrl nil The list of discovered URLs.

Rich markdown conversion result from HTML processing.

Field Type Default Description
Content string Converted markdown text.
DocumentStructure *interface{} nil Structured document tree with semantic nodes.
Tables \[\]interface{} nil Extracted tables with structured cell data.
Warnings \[\]string nil Non-fatal processing warnings.
Citations bool Whether citation conversion was applied and produced at least one reference. true when the markdown contained inline links that were converted to numbered citation references. The converted content (with \[N\] markers) is available in content; the full reference list is accessible via generate_citations if needed separately.
FitContent *string nil Content-filtered markdown optimized for LLM consumption.

Metadata extracted from an HTML page’s <meta> tags and <title> element.

Field Type Default Description
Title *string nil The page title from the <title> element.
Description *string nil The meta description.
CanonicalUrl *string nil The canonical URL from <link rel="canonical">.
Keywords *string nil Keywords from <meta name="keywords">.
Author *string nil Author from <meta name="author">.
Viewport *string nil Viewport content from <meta name="viewport">.
ThemeColor *string nil Theme color from <meta name="theme-color">.
Generator *string nil Generator from <meta name="generator">.
Robots *string nil Robots content from <meta name="robots">.
HtmlLang *string nil The lang attribute from the <html> element.
HtmlDir *string nil The dir attribute from the <html> element.
OgTitle *string nil Open Graph title.
OgType *string nil Open Graph type.
OgImage *string nil Open Graph image URL.
OgDescription *string nil Open Graph description.
OgUrl *string nil Open Graph URL.
OgSiteName *string nil Open Graph site name.
OgLocale *string nil Open Graph locale.
OgVideo *string nil Open Graph video URL.
OgAudio *string nil Open Graph audio URL.
OgLocaleAlternates *\[\]string nil Open Graph locale alternates.
TwitterCard *string nil Twitter card type.
TwitterTitle *string nil Twitter title.
TwitterDescription *string nil Twitter description.
TwitterImage *string nil Twitter image URL.
TwitterSite *string nil Twitter site handle.
TwitterCreator *string nil Twitter creator handle.
DcTitle *string nil Dublin Core title.
DcCreator *string nil Dublin Core creator.
DcSubject *string nil Dublin Core subject.
DcDescription *string nil Dublin Core description.
DcPublisher *string nil Dublin Core publisher.
DcDate *string nil Dublin Core date.
DcType *string nil Dublin Core type.
DcFormat *string nil Dublin Core format.
DcIdentifier *string nil Dublin Core identifier.
DcLanguage *string nil Dublin Core language.
DcRights *string nil Dublin Core rights.
Article *ArticleMetadata nil Article metadata from article:* Open Graph tags.
Hreflangs *\[\]HreflangEntry nil Hreflang alternate links.
Favicons *\[\]FaviconInfo nil Favicon and icon links.
Headings *\[\]HeadingInfo nil Heading elements (h1-h6).
WordCount *int nil Computed word count of the page body text.

Proxy configuration for HTTP requests.

Field Type Default Description
Url string Proxy URL (e.g. “http://proxy:8080", “socks5://proxy:1080”).
Username *string nil Optional username for proxy authentication.
Password *string nil Optional password for proxy authentication.

Response metadata extracted from HTTP headers.

Field Type Default Description
Etag *string nil The ETag header value.
LastModified *string nil The Last-Modified header value.
CacheControl *string nil The Cache-Control header value.
Server *string nil The Server header value.
XPoweredBy *string nil The X-Powered-By header value.
ContentLanguage *string nil The Content-Language header value.
ContentEncoding *string nil The Content-Encoding header value.

The result of a single-page scrape operation.

Field Type Default Description
StatusCode uint16 The HTTP status code of the response.
FinalUrl string The final URL after following all redirects.
ContentType string The Content-Type header value.
Html string The HTML body of the response.
BodySize int The size of the response body in bytes.
Metadata PageMetadata Extracted metadata from the page.
Links \[\]LinkInfo nil Links found on the page.
Images \[\]ImageInfo nil Images found on the page.
Feeds \[\]FeedInfo nil Feed links found on the page.
JsonLd \[\]JsonLdEntry nil JSON-LD entries found on the page.
IsAllowed bool Whether the URL is allowed by robots.txt.
CrawlDelay *uint64 nil The crawl delay from robots.txt, in seconds.
NoindexDetected bool Whether a noindex directive was detected.
NofollowDetected bool Whether a nofollow directive was detected.
XRobotsTag *string nil The X-Robots-Tag header value, if present.
IsPdf bool Whether the content is a PDF.
WasSkipped bool Whether the page was skipped (binary or PDF content).
DetectedCharset *string nil The detected character set encoding.
AuthHeaderSent bool Whether an authentication header was sent with the request.
ResponseMeta *ResponseMeta nil Response metadata extracted from HTTP headers.
Assets \[\]DownloadedAsset nil Downloaded assets from the page.
JsRenderHint bool Whether the page content suggests JavaScript rendering is needed.
BrowserUsed bool Whether the browser fallback was used to fetch this page.
Markdown *MarkdownResult nil Markdown conversion of the page content.
ExtractedData *interface{} nil Structured data extracted by LLM. Populated when extraction is configured.
ExtractionMeta *ExtractionMeta nil Metadata about the LLM extraction pass (cost, tokens, model).
ScreenshotBase64 *string nil Base64-encoded PNG screenshot of the page. Populated only when CrawlConfig.capture_screenshot was enabled for this request, so callers that never requested a screenshot do not pay the encoding cost.
DownloadedDocument *DownloadedDocument nil Downloaded non-HTML document (PDF, DOCX, image, code, etc.).
Browser *BrowserExtras nil Browser-specific extras (eval result, network events, cookies). Only populated when BrowserBackend.Native was used for this request.

A URL entry from a sitemap.

Field Type Default Description
Url string The URL.
Lastmod *string nil The last modification date, if present.
Changefreq *string nil The change frequency, if present.
Priority *string nil The priority, if present.

SSRF policy configuration.

Field Type Default Description
DenyPrivate bool true If true, reject URLs that resolve to private/metadata IP ranges.
Allowlist \[\]HostMatcher nil Hostnames and IP ranges permitted regardless of deny_private. The allowlist is an override of deny_private, not an intersection with it. Precedence, in order: 1. deny_private == false permits everything; the allowlist is not consulted. 2. A hostname matching an Exact or Suffix entry is permitted immediately, before DNS resolution — so the deny-list is never applied to it. This trusts the host string: a name that resolves into private space is still permitted. 3. A literal or resolved IP inside a Cidr entry is permitted even though it is in the default deny-list. 4. Otherwise the default deny-list decides. An empty allowlist therefore denies nothing by itself — it simply leaves deny_private and the deny-list in sole control.
MaxRedirects uint8 5 Maximum number of HTTP redirects to follow during validation.

Signature:

func (o *SsrfPolicy) Default() SsrfPolicy

Example:

result := SsrfPolicy.Default()

Returns: SsrfPolicy

Create a policy from environment variables.

On native platforms, reads CRAWLBERG_ALLOW_PRIVATE_NETWORK — if set to “1” or “true” (case-insensitive), sets deny_private = false. Otherwise, defaults to deny_private = true.

On wasm32 targets (browser/Node.js), environment variables are not accessible to the compiled module. Defaults to deny_private = false because:

  • Outbound requests in a browser go through the fetch API, which enforces its own network policies.
  • Rust-side SSRF checking is unenforceable and redundant in a wasm32 context.
  • For testing and localhost access, the host’s network sandbox is the enforcing boundary.

Node.js caveat: deny_private (whatever its value) has no effect on hostname-based requests under wasm32. There is no DNS resolution on this target, so validate_url only ever checks a literal IP host; a domain name falls straight through to Ok(()). In a browser this is covered by same-origin/CORS. Node’s fetch enforces no CORS, so a Node service embedding this wasm module can be driven to internal hosts by domain name even though deny_private = true. Do not rely on this policy to stop that in Node — enforce egress restrictions (network policy, firewall, proxy allowlist) outside the process.

Signature:

func (o *SsrfPolicy) FromEnv() SsrfPolicy

Example:

result := SsrfPolicy.FromEnv()

Returns: SsrfPolicy


When to use the headless browser fallback.

Value Description
Auto Automatically detect when JS rendering is needed and fall back to browser.
Always Always use the browser for every request.
Never Never use the browser fallback.
Stealth Always use the browser with all stealth surfaces enabled. Behaves like Always for escalation purposes (every request is routed through the browser tier), but additionally enables: - browser JavaScript stealth patches - native-backend TLS fingerprint spoofing - stealth-aware default user-agent when no explicit UA is set - 1920×1080 viewport override Use this instead of setting the now-removed BrowserConfig.stealth boolean field.

Wait strategy for browser page rendering.

Value Description
NetworkIdle Wait until network activity is idle.
Selector Wait for a specific CSS selector to appear in the DOM.
Fixed Wait for a fixed duration after navigation.

Browser backend used for JavaScript rendering.

Value Description
Chromiumoxide Existing Chromium/CDP backend powered by chromiumoxide.
Native Crawlberg-owned native browser backend derived from Obscura.

Opt-in encoding applied to a downloaded document’s bytes for callers who need the content available in a serializable field rather than reading it from disk.

nil (the CrawlConfig.document_content_encoding default) produces neither — unlike screenshots, base64-encoding a document by default would duplicate an already up-to-document_max_size buffer (50 MB default) in memory per document.

Value Description
Base64 Populate DownloadedDocument.content_base64 with a base64-encoded copy.

Authentication configuration.

Value Description
Basic HTTP Basic authentication. — Fields: Username: string, Password: string
Bearer Bearer token authentication. — Fields: Token: string
Header Custom authentication header. — Fields: Name: string, Value: string

The classification of a link.

Value Description
Internal A link to the same domain.
External A link to a different domain.
Anchor A fragment-only link (e.g., #section).
Document A link to a downloadable document (PDF, DOC, etc.).

The source of an image reference.

Value Description
Img An <img> tag.
PictureSource A <source> tag inside <picture>.
OgImage An og:image meta tag.
TwitterImage A twitter:image meta tag.

The type of a feed (RSS, Atom, or JSON Feed).

Value Description
Rss RSS feed.
Atom Atom feed.
JsonFeed JSON Feed.

The category of a downloaded asset.

Value Description
Document A document file (PDF, DOC, etc.).
Image An image file.
Audio An audio file.
Video A video file.
Font A font file.
Stylesheet A CSS stylesheet.
Script A JavaScript file.
Archive An archive file (ZIP, TAR, etc.).
Data A data file (JSON, XML, CSV, etc.).
Other An unrecognized asset type.

An event emitted during a streaming crawl operation.

Not available on wasm32 targets — streaming requires native concurrency primitives (tokio channels, JoinSet) that are not supported on wasm32.

Delivered to bindings through each target’s native streaming idiom.

Value Description
Page A single page has been crawled. — Fields: Result: CrawlPageResult
Error An error occurred while crawling a URL. — Fields: Url: string, Error: string
Complete The crawl has completed. — Fields: PagesCrawled: int

A single page interaction action.

Actions are serialized with a type tag using camelCase naming, except ExecuteJs which is explicitly renamed to "executeJs".

Value Description
Click Click on an element matching the given CSS selector. — Fields: Selector: string
TypeText Type text into an element matching the given CSS selector. — Fields: Selector: string, Text: string
Press Press a keyboard key (e.g. “Enter”, “Tab”, “Escape”). — Fields: Key: string
Scroll Scroll the page or a specific element. — Fields: Direction: ScrollDirection, Selector: string, Amount: int64
Wait Wait for a duration or for an element to appear. — Fields: Milliseconds: int64, Selector: string
Screenshot Take a screenshot of the current page. — Fields: FullPage: bool
ExecuteJs Execute arbitrary JavaScript in the page context. Safety: The script runs with full page privileges in the browser context. Only execute scripts from trusted sources. — Fields: Script: string
Scrape Scrape the current page HTML.

Direction for a scroll action.

Value Description
Up Scroll upward.
Down Scroll downward.

Hostname/IP allowlist matcher for SSRF policy.

Serializes as an internally-tagged object so each variant is distinguishable on the wire and round-trips losslessly:

{"type": "exact", "value": "api.example.com"}
{"type": "suffix", "value": ".example.com"}
{"type": "cidr", "value": "10.0.0.0/8"}

A bare JSON string is still accepted on deserialization and resolves to Exact, preserving configs written against the previous untagged representation.

Exact: HostMatcher.Exact

Value Description
Exact Exact hostname match (case-insensitive). — Fields: Value: string
Suffix Suffix match: “.xberg.io” matches “api.xberg.io” and “xberg.io”. — Fields: Value: string
Cidr CIDR match: “10.0.0.0/8” matches IP addresses in that range. — Fields: Value: string

Errors that can occur during crawling, scraping, or mapping operations.

Variant Description
NotFound The requested page was not found (HTTP 404).
Unauthorized The request was unauthorized (HTTP 401).
Forbidden The request was forbidden (HTTP 403).
WafBlocked The request was blocked by a WAF or bot protection (HTTP 403 with WAF indicators). vendor is the lowercase identifier of the detected WAF (e.g. “cloudflare”, “datadome”). When the engine cannot identify the vendor, it uses “unknown”. message is the freeform description for logs and human readers. The stable error tag remains forbidden: waf/blocked: MESSAGE so existing log-grep patterns and cross-language bindings continue to work; vendor is surfaced separately for structured consumers.
Timeout The request timed out.
RateLimited The request was rate-limited (HTTP 429).
ServerError A server error occurred (HTTP 5xx).
BadGateway A bad gateway error occurred (HTTP 502).
Gone The resource is permanently gone (HTTP 410).
Connection A connection error occurred.
Dns A DNS resolution error occurred.
Ssl An SSL/TLS error occurred.
DataLoss Data was lost or truncated during transfer.
BrowserError The browser failed to launch, connect, or navigate.
BrowserTimeout The browser page load or rendering timed out.
InvalidConfig The provided configuration is invalid.
Unsupported The requested capability is not supported by the active backend or build.
SsrfPolicyViolation A URL was rejected by SSRF policy (private IP, metadata, disallowed scheme, etc).
Other An unclassified error occurred.

SSRF validation error.

Variant Description
DeniedByPolicy URL denied by SSRF policy: private IP, metadata IP, etc.
NotOnAllowlist Host not on allowlist when an allowlist is configured.
InvalidCidr Allowlist entry is not a parseable CIDR block.
DnsResolutionFailed DNS resolution failed for hostname.
InvalidUrl Invalid URL format.
DisallowedScheme URL scheme not in allowlist (e.g., ftp:// when only http/https allowed).
TooManyRedirects Too many HTTP redirects encountered during validation.