Pod

Available as Markdown and JSON. Pod is also available over MCP.

Reported issues for Iris by iris-eval

Pod holds 18 of 22 GitHub reports that passed its relevance review. This can include external user reports, maintainer-confirmed bugs, and concrete feature gaps. Treat them as evidence to inspect, not a count of distinct defects.

Back to Iris by iris-eval.

Most discussed

Storage init: friendlier failure when better-sqlite3 bindings are missing (#369 item 7)

Carried over from #369 item 7; confirmed during the v0.6.0 polish arc.

better-sqlite3 is a static import in src/storage/sqlite-adapter.ts, evaluated when the module graph loads, so no catch in storage init can turn "Could not locate the bindings file" into the README-grade message (npm rebuild better-sqlite3). --self-test already reports the symptom and the README documents the fix; the remaining ask is a lazy load in the storage boot path so the server itself prints the fix on first…

Read the thread · 2026-09-03 · closed · 1 comment

Input validation: nested objects and dashboard routes still strip unknown keys

Follow-ups to the strict top-level tool schemas (fix/uat-args):

  1. Nested objects still strip silently. custom_rules: [{name: 'a', wieght: 5}] parses fine with wieght discarded — a misspelled rule weight silently changes scoring. Same family as the fixed defect, one level down. Decide: recursive strictness (with care for the free-form record fields — metadata, span attributes, rule config — whose arbitrary keys are legitimate) or document the boundary.
  2. **All 14 dashboard-route…

Read the thread · 2026-08-12 · closed · 1 comment

passed semantics: persistence and cross-tool consistency gaps after the critical-rule veto

Follow-ups to the hard-fail semantics change (fix/uat-crit), found during its test-completion:

  1. critical_failures is never persisted. eval_results has no column for it and insertEvalResult (src/storage/sqlite-adapter.ts:263) doesn't write it. The MCP response carries it (computed in-process), but anything reading the eval back — dashboard drill-through, trace summaries, MCP resources — sees passed: false with no stated cause. Worse: the veto suggestion line is only appended…

Read the thread · 2026-08-12 · closed · 1 comment

PII patterns are label-anchored more than the pattern list implies

Three of the 19 patterns only fire with specific labels: ISO-format DOB (Date of birth: 1987-03-15) is MISSED while DOB: 03/15/1987 is caught; unlabeled passport numbers in prose are missed; an unlabeled 12-word BIP39 seed phrase is missed while My seed phrase is: … is caught. Fix passport to its documented modern format, add an ISO alternative to the DOB date part, and qualify the README pattern list to note label-anchoring.


_Found during install-only acceptance testing of the…

Read the thread · 2026-08-12 · closed · 1 comment

Dashboard/API correctness batch

  1. /api/v1/health trace_count contradicts /api/v1/traces (0 vs 253 in demo mode): health uses a windowed summary. Make it an all-time COUNT(*) or rename to traces_last_hour (src/dashboard/routes/health.ts:21).
  2. Rule preview endpoint silently ignores sample text: deploy_rule's description says to use POST /api/v1/rules/custom/preview for "dry-run validation against sample output", but the endpooint only replays stored traces; sampleOutput/output/sample/text keys…

Read the thread · 2026-08-12 · closed · 1 comment

Data retention and permissions for stored eval/trace text

PII the tool just flagged is stored verbatim in plaintext and served to any local process; deleting iris.db leaves the text in iris.db-wal.

  1. Document in the README that iris.db stores raw output text verbatim, including anything no_pii flags.
  2. eval_results is missing from the retention sweep — extend deleteTracesOlderThan (src/storage/sqlite-adapter.ts:622) or add deleteEvalResultsOlderThan called alongside src/index.ts:250; add a --purge flag.
  3. Create IRIS_HOME…

Read the thread · 2026-08-12 · closed · 1 comment

Add OpenTelemetry trace export support

Iris currently stores traces in SQLite. Many teams already have observability stacks (Datadog, Grafana, New Relic) that ingest OpenTelemetry spans.

Goal: Add an optional OTel exporter that converts Iris traces to OTel spans and sends them to any OTLP-compatible endpoint.

Acceptance criteria:

Read the thread · 2026-03-16 · closed · 1 comment

Search index rebuild: dropping the old index holds the event loop (5.2 s at 100,000 traces)

What happens

When the trace search index has to be rebuilt from scratch, the server renames the old index at the start and drops it in the background, before the rebuild indexes anything (reconcileSearchIndex and dropRetiredIndex in src/storage/search-index.ts). The drop is one DROP TABLE statement. With secure_delete on, SQLite rewrites and zeroes every page of the old index, and the statement holds the event loop for its whole length.

That rebuild happens when:

Read the thread · 2026-09-27 · open · 0 comments

Most recent

Trace search: include span attributes and events, for OTLP traces

Trace search (q, #7) reads four fields of the trace row: input, output, and the values inside tool_calls and metadata. It does not read spans.

For a trace sent through POST /v1/traces (OTLP), that leaves text out. src/otel/ingest.ts fills input and output from the first span that carries them (gen_ai.input.messages, gen_ai.output.messages, input.value and the other keys it knows). Everything else stays only in the spans table, as span attributes and events:

Read the thread · 2026-09-26 · closed · 0 comments

Trace search: find words inside Chinese, Japanese and Korean text

Trace search (q on get_traces, GET /api/v1/traces, the dashboard, the Python client; #7) indexes text with SQLite FTS5's unicode61 tokenizer. That tokenizer splits on spaces and punctuation. Chinese and Japanese text have no spaces between words, and Korean runs words together often enough, so a whole run of CJK characters becomes one token. A search for a word inside that run finds nothing.

Example: an output reading 退款已经批准了,请查收邮件 ("the refund has been approved, please check your…

Read the thread · 2026-09-26 · open · 0 comments

Regex sandbox: a busy host kills trivial patterns as backtracking

On a heavily loaded host, the regex sandbox kills trivial custom-rule patterns as though they were backtracking. The rule is then reported as a budget skip, and when the rule is critical the verdict becomes unknown.

Where. src/eval/rules/regex-sandbox.ts. sandboxedRegexTest posts the match to the worker and then blocks in Atomics.wait(view, 0, 0, budgetMs), with REGEX_MATCH_BUDGET_MS = 100. That wall-clock deadline covers the whole round trip:

Read the thread · 2026-09-26 · closed · 0 comments

e2e: a stray server on port 6921 is tested instead of the current build

The Playwright suite starts its dashboard on a hardcoded port, 6921 (tests/e2e/_constants.ts). Outside CI, playwright.config.ts sets reuseExistingServer, so if anything is already listening on 6921 the suite quietly runs against that server instead of the build it just made. The "something" could be an Iris started with npx @iris-eval/mcp-server --dashboard --dashboard-port 6921, an earlier run's dashboard, or another checkout's.

A test of new code then fails, or passes, against an old…

Read the thread · 2026-09-26 · closed · 0 comments

Install succeeds when better-sqlite3 cannot build: make it optional and run on node:sqlite

better-sqlite3 is a hard dependency, so when it cannot be installed the whole npm install @iris-eval/mcp-server fails, even though Iris already runs without it: since 0.15.0 the storage layer loads it lazily and falls back to Node's built-in node:sqlite (Node 22.13+) with a one-line warning.

This matters whenever a better-sqlite3 release ships without prebuilt binaries for a platform. 13.0.3 publishes none for Windows, so npm falls back to compiling with node-gyp, which needs Visual…

Read the thread · 2026-09-26 · closed · 0 comments

Moments: filtered pages in the newest-first order come back short and page by trace offset

GET /api/v1/moments in its default order (sort_by=timestamp) applies the verdict, min_significance and significance_kind filters after it reads a page of traces. As a result:

Read the thread · 2026-09-26 · closed · 0 comments

Build an MCPB bundle with each release: one-click install in Claude Desktop, and the path back into Smithery

An MCPB bundle (.mcpb) is a packaged MCP server that installs in one step. Claude Desktop installs one by double-click, and Smithery lists a local server through one. Iris does not build one, so neither path is open; the Smithery listing was retired in 0.16.0 when its old smithery.yaml form stopped being the supported path.

What ships

Read the thread · 2026-09-26 · closed · 0 comments

An off-topic answer should fail by default when the LLM judge is on: a relevance judge for answers_the_ask

answers_the_ask compares the words of the answer with the words of the ask. At the shipped thresholds that failed correct answers that paraphrase, so since 0.18.0 it advises instead of blocking unless a threshold is set. The result is that an answer on the wrong topic passes at the defaults.

What ships

Read the thread · 2026-09-26 · closed · 0 comments

Record every OpenAI and Anthropic call without the model choosing to: thin wrappers in Python and JavaScript

Today a trace reaches Iris when the agent calls log_trace, when it is posted to POST /api/v1/traces, through OpenTelemetry, or through the Claude Code capture plugin. An application that calls a model provider directly has no one-line way to send every call.

What ships

Read the thread · 2026-09-26 · closed · 0 comments

LangChain and LangGraph integration: Python and JavaScript, proven against a real LangGraph app

Iris should score a LangChain or LangGraph agent without the agent having to call an Iris tool.

What ships

Read the thread · 2026-09-25 · closed · 0 comments

The remaining reports are on the project's issue tracker.