Reported issues for codearia-sieve
Pod holds 5 of 5 GitHub reports that passed its relevance review. This can include external user reports, maintainer-confirmed bugs, and concrete feature gaps. Treat them as evidence to inspect, not a count of distinct defects.
Back to codearia-sieve.
Most discussed
Title suffixes like … | Blog and … | Site Name end up in state.title
web.dev and developer.chrome.com pages come back with state.title = 'New to the web platform in May | Blog'; academy.codearia.com pages come back with the site name only (Codearia Academy), because the extractor picks the wrong side of <title>codearia-sieve — MCP-сервер | Codearia Academy</title>.
Proposed rule (src/extract.ts): when <title> ends with | X, - X, — X or :: X and X equals og:site_name (or the <title> of the site root), drop the suffix; when…
Read the thread · 2026-09-22 · open · 1 comment
Japanese kitchen units: 大さじ, 小さじ, 分, 個 are not facts
Cookpad recipes (lab 8) yield 0 facts because Japanese units are unknown to src/normalize/numbers.ts.
What to add
大さじ→tbsp,小さじ→tsp,分→min,個→piece,カップ→cup,合(rice) can wait- CJK units glue to the digits and to the particle after them, so the entry goes in the block without the
(?!\p{L})boundary check, like the CJK currency words already there
How to check
- A test in
test/numbers.test.tsnext to'Hebrew and Arabic currencies…':…
Read the thread · 2026-09-22 · open · 1 comment
Empty list or table container inside a full page gives no signal (script-rendered ingredients, fee tables)
Marmiton's ingredient list and USCIS's fee schedule are rendered by script; the surrounding prose is static, so the page passes with no warning and the numbers are simply missing.
empty-without-js today fires when the article container is empty (looksClientRendered in src/extract.ts). A narrower signal is possible: an <ul>/<ol>/<table> whose class or id names ingredients/fees/prices (ingredient, fee, price, schedule) and whose text is empty in the served HTML.
**Proposed…
Read the thread · 2026-09-22 · open · 0 comments
state.meta for SEO audits: html title, h1, meta description, canonical, schema.org types
Two lab scenarios (competitor product pages, documentation freshness) wanted fields Sieve does not expose: the literal <title>, the <h1>, meta[name=description], link[rel=canonical] and the schema.org @types found in JSON-LD.
All of them are read from the untouched tree, the same way dates are (findDates in src/normalize/dates.ts), so the shape is:
state.meta?: { htmlTitle?: string; h1?: string; description?: string; canonical?: string; schemaTypes?: string[] }
```…
[Read the thread](https://github.com/AntonG87/codearia-sieve/issues/4) · 2026-09-22 · open · 0 comments
### Korean and Japanese compound numbers read only their last group (1조 8000억 → 8000억)
`extractFacts([block('GDP는 1조 8000억 달러였다')], 'ko')` returns `8000e8 USD`; the `1조` is lost. Same for Japanese `1兆2000億円`.
**Where**: `scan()` in `src/normalize/numbers.ts` — after a CJK scale (`조`/`억`/`兆`/`億`) matches, look back for a preceding `<number><bigger scale>` group and add them.
**Check**: extend the `'scale and currency words across languages'` test in `test/numbers.test.ts` with the two compound cases; the existing case `8000억 달러` must still pass.
Harder than the Japanese units…
[Read the thread](https://github.com/AntonG87/codearia-sieve/issues/2) · 2026-09-22 · open · 0 comments
## Most recent
The remaining reports are on [the project's issue tracker](https://github.com/AntonG87/codearia-sieve/issues).