Reported issues for Crawlberg
Pod holds 6 of 6 GitHub reports that passed its relevance review. This can include external user reports, maintainer-confirmed bugs, and concrete feature gaps. Treat them as evidence to inspect, not a count of distinct defects.
Back to Crawlberg.
Most discussed
fix(mcp): the download tool returns the caller's URL with its password
Description
The MCP download tool returns the caller's URL in its result with the user name and password still in it. The CLI download command prints the same URL. Expected: the URL comes back without its userinfo, as it does from the other entry points.
Example: downloading https://alice:hunter2@example.com/file.pdf returns that string, password included.
Steps to reproduce
- Start a local server that serves an HTML page.
- Call the MCP
downloadtool with the page's address…
Read the thread · 2026-09-27 · closed · 1 comment
fix(browser): a proxy address with an upper-case scheme loses its credentials
A proxy address with an upper-case scheme, such as HTTP://user:pass@proxy:8080, loses its configured credentials without any error: the native browser backend recognises the credentials only after a lower-case http:// prefix. The proxy is then used without authentication, so requests fail with a 407 or, on an open proxy, go out unauthenticated.
Where
crates/crawlberg/src/native_browser.rs:155-157
Fix
Parse the proxy address with the URL parser and take the scheme, user and…
Read the thread · 2026-09-26 · open · 0 comments
fix(api): an address with an upper-case scheme is refused
The REST API and the MCP tools refuse a caller-supplied address whose scheme is not lower case, such as HTTP://example.com/ or Https://example.com/. A URL scheme is case-insensitive, and the URL parser accepts these addresses.
Where
crates/crawlberg/src/api/handlers.rs:62crates/crawlberg/src/mcp/tool_result.rs:10
Fix
Parse the address with the URL parser and check the parsed scheme, instead of testing the raw text for an http:// or https:// prefix. Use one check for…
Read the thread · 2026-09-26 · open · 0 comments
fix(api): a REST caller can send backtracking URL patterns that stall the crawl
After #181, a REST client can send include_paths and exclude_paths patterns that use backtracking (look-around or backreferences). A pathological pattern costs about 34 ms per URL before it hits the backtracking limit, and the check runs on the async executor. The crawled site decides the URL count and the remote client decides the patterns, so one crawl request can stall the server's executor, and it logs one warning per URL.
Fix
For REST and MCP callers, either reject patterns that…
Read the thread · 2026-09-26 · open · 0 comments
fix(interact): with an external browser, interact intercepts every tab of the remote browser
After #163, interact enables its request check on the browser session. With browser.endpoint set, that session is the caller's shared browser, so while interact runs, it pauses and checks the requests of every tab, including pages that other clients opened.
Decision needed
Checking the whole browser is what covers popups. Options: keep it and document it; limit the check to targets that interact's own page opened (by opener id); or refuse interact on a shared browser unless the caller…
Read the thread · 2026-09-26 · open · 0 comments
fix(api): the MCP and REST crawl tools lack the robots and path-pattern settings
The CLI and the library let a caller set respect_robots_txt, which after #154 also controls nofollow handling. The MCP crawl tool and the REST crawl endpoint accept no such input, so their callers cannot turn it on or off.
Fix
Add the setting to the MCP tool parameters and the REST request body, with the same default and validation as the CLI.
The REST crawl request (crates/crawlberg/src/api/types.rs, around line 50) also takes include_paths and exclude_paths but neither…
Read the thread · 2026-09-26 · open · 0 comments
Most recent
The remaining reports are on the project's issue tracker.