Reported issues for Kreuzcrawl
Pod holds 6 of 6 GitHub reports that passed its relevance review. This can include external user reports, maintainer-confirmed bugs, and concrete feature gaps. Treat them as evidence to inspect, not a count of distinct defects.
Back to Kreuzcrawl.
Most discussed
fix(mcp): the download tool returns the caller's URL with its password
Description
The MCP download tool returns the caller's URL in its result with the user name and password still in it. The CLI download command prints the same URL. Expected: the URL comes back without its userinfo, as it does from the other entry points.
Example: downloading https://alice:hunter2@example.com/file.pdf returns that string, password included.
Steps to reproduce
- Start a local server that serves an HTML page.
- Call the MCP
downloadtool with the page's address…
Read the thread · 2026-09-27 · closed · 1 comment
fix(http): only a debug assertion keeps userinfo out of the fetch layer
Description
Only a debug assertion stops a URL that still carries userinfo from reaching the fetch layer. In a release build the assertion is compiled out. A future path that keeps userinfo past admission would then send and log it with nothing to stop it: the HTTP client silently turns the userinfo into an Authorization: Basic header for whatever host the URL names. The transport-error test passes with the assertion disabled and the userinfo kept, so it cannot catch that either.…
Read the thread · 2026-09-27 · open · 0 comments
fix(api): an address with an upper-case scheme is refused
The REST API and the MCP tools refuse a caller-supplied address whose scheme is not lower case, such as HTTP://example.com/ or Https://example.com/. A URL scheme is case-insensitive, and the URL parser accepts these addresses.
Where
crates/crawlberg/src/api/handlers.rs:62crates/crawlberg/src/mcp/tool_result.rs:10
Fix
Parse the address with the URL parser and check the parsed scheme, instead of testing the raw text for an http:// or https:// prefix. Use one check for…
Read the thread · 2026-09-26 · open · 0 comments
fix(browser): WebSocket, WebRTC and WebTransport connections and DNS rebinding bypass the SSRF check
In browser mode (crawl, scrape and interact), some traffic does not go through the CDP Fetch check that enforces the SSRF policy:
- A
new WebSocket(...)to a denied address opens a TCP connection. Measured on Chrome 154:Network.setBlockedURLson the page still allowed 1 connect, the browser session does not support it, andwebSocketWillSendHandshakeRequestandwebSocketCreatedfire after the connect. - WebRTC and WebTransport connections are not covered either.
- The check resolves a…
Read the thread · 2026-09-26 · open · 0 comments
fix(browser): browser-mode popups and late requests escape the SSRF check
In browser mode, some requests the page makes reach addresses the SSRF policy refuses:
- A popup (
window.open) that the page opens on load. Fetch interception on the page does not cover a new target. - A request sent after the check ends, during the extra wait or while taking a screenshot. A
fetch700 ms after load reached a denied address.
#163 fixes the first two for interact by checking on the browser session for the whole session.
Fix
Keep the browser-session check on for the…
Read the thread · 2026-09-26 · open · 0 comments
fix(api): the MCP and REST crawl tools lack the robots and path-pattern settings
The CLI and the library let a caller set respect_robots_txt, which after #154 also controls nofollow handling. The MCP crawl tool and the REST crawl endpoint accept no such input, so their callers cannot turn it on or off.
Fix
Add the setting to the MCP tool parameters and the REST request body, with the same default and validation as the CLI.
The REST crawl request (crates/crawlberg/src/api/types.rs, around line 50) also takes include_paths and exclude_paths but neither…
Read the thread · 2026-09-26 · open · 0 comments
Most recent
The remaining reports are on the project's issue tracker.