# Reported issues for Iris by iris-eval

Pod holds 18 of 22 GitHub reports that passed its relevance review. This can include external user reports, maintainer-confirmed bugs, and concrete feature gaps. Treat them as evidence to inspect, not a count of distinct defects.

Back to [Iris by iris-eval](/mcp/iris-by-iris-eval).

## Most discussed

### Storage init: friendlier failure when better-sqlite3 bindings are missing (#369 item 7)

Carried over from #369 item 7; confirmed during the v0.6.0 polish arc.

`better-sqlite3` is a static import in `src/storage/sqlite-adapter.ts`, evaluated when the module graph loads, so no catch in storage init can turn "Could not locate the bindings file" into the README-grade message (`npm rebuild better-sqlite3`). `--self-test` already reports the symptom and the README documents the fix; the remaining ask is a lazy load in the storage boot path so the server itself prints the fix on first…

[Read the thread](https://github.com/iris-eval/mcp-server/issues/408) · 2026-09-03 · closed · 1 comment

### Input validation: nested objects and dashboard routes still strip unknown keys

Follow-ups to the strict top-level tool schemas (fix/uat-args):

1. **Nested objects still strip silently.** `custom_rules: [{name: 'a', wieght: 5}]` parses fine with `wieght` discarded — a misspelled rule weight silently changes scoring. Same family as the fixed defect, one level down. Decide: recursive strictness (with care for the free-form record fields — `metadata`, span `attributes`, rule `config` — whose arbitrary keys are legitimate) or document the boundary.
2. **All 14 dashboard-route…

[Read the thread](https://github.com/iris-eval/mcp-server/issues/376) · 2026-08-12 · closed · 1 comment

### `passed` semantics: persistence and cross-tool consistency gaps after the critical-rule veto

Follow-ups to the hard-fail semantics change (fix/uat-crit), found during its test-completion:

1. **`critical_failures` is never persisted.** `eval_results` has no column for it and `insertEvalResult` (`src/storage/sqlite-adapter.ts:263`) doesn't write it. The MCP response carries it (computed in-process), but anything reading the eval back — dashboard drill-through, trace summaries, MCP resources — sees `passed: false` with no stated cause. Worse: the veto suggestion line is only appended…

[Read the thread](https://github.com/iris-eval/mcp-server/issues/375) · 2026-08-12 · closed · 1 comment

### PII patterns are label-anchored more than the pattern list implies

Three of the 19 patterns only fire with specific labels: ISO-format DOB (`Date of birth: 1987-03-15`) is MISSED while `DOB: 03/15/1987` is caught; unlabeled passport numbers in prose are missed; an unlabeled 12-word BIP39 seed phrase is missed while `My seed phrase is: …` is caught. Fix passport to its documented modern format, add an ISO alternative to the DOB date part, and qualify the README pattern list to note label-anchoring.

---
_Found during install-only acceptance testing of the…

[Read the thread](https://github.com/iris-eval/mcp-server/issues/374) · 2026-08-12 · closed · 1 comment

### Dashboard/API correctness batch

1. **`/api/v1/health` `trace_count` contradicts `/api/v1/traces`** (0 vs 253 in demo mode): health uses a windowed summary. Make it an all-time `COUNT(*)` or rename to `traces_last_hour` (`src/dashboard/routes/health.ts:21`).
2. **Rule preview endpoint silently ignores sample text**: `deploy_rule`'s description says to use `POST /api/v1/rules/custom/preview` for "dry-run validation against sample output", but the endpooint only replays stored traces; `sampleOutput`/`output`/`sample`/`text` keys…

[Read the thread](https://github.com/iris-eval/mcp-server/issues/373) · 2026-08-12 · closed · 1 comment

### Data retention and permissions for stored eval/trace text

PII the tool just flagged is stored verbatim in plaintext and served to any local process; deleting `iris.db` leaves the text in `iris.db-wal`.

1. Document in the README that `iris.db` stores raw output text verbatim, including anything `no_pii` flags.
2. `eval_results` is missing from the retention sweep — extend `deleteTracesOlderThan` (`src/storage/sqlite-adapter.ts:622`) or add `deleteEvalResultsOlderThan` called alongside `src/index.ts:250`; add a `--purge` flag.
3. Create `IRIS_HOME`…

[Read the thread](https://github.com/iris-eval/mcp-server/issues/372) · 2026-08-12 · closed · 1 comment

### Add OpenTelemetry trace export support

Iris currently stores traces in SQLite. Many teams already have observability stacks (Datadog, Grafana, New Relic) that ingest OpenTelemetry spans.

**Goal:** Add an optional OTel exporter that converts Iris traces to OTel spans and sends them to any OTLP-compatible endpoint.

**Acceptance criteria:**
- New CLI flag `--otel-endpoint` to configure the OTLP endpoint
- Traces logged via `log_trace` are also exported as OTel spans
- Spans include GenAI semantic conventions (model, token counts,…

[Read the thread](https://github.com/iris-eval/mcp-server/issues/3) · 2026-03-16 · closed · 1 comment

### Search index rebuild: dropping the old index holds the event loop (5.2 s at 100,000 traces)

## What happens

When the trace search index has to be rebuilt from scratch, the server renames the old index at the start and drops it in the background, before the rebuild indexes anything (`reconcileSearchIndex` and `dropRetiredIndex` in `src/storage/search-index.ts`). The drop is one `DROP TABLE` statement. With `secure_delete` on, SQLite rewrites and zeroes every page of the old index, and the statement holds the event loop for its whole length.

That rebuild happens when:

- a file was…

[Read the thread](https://github.com/iris-eval/mcp-server/issues/695) · 2026-09-27 · open · 0 comments

## Most recent

### Trace search: include span attributes and events, for OTLP traces

Trace search (`q`, #7) reads four fields of the trace row: `input`, `output`, and the values inside `tool_calls` and `metadata`. It does not read spans.

For a trace sent through `POST /v1/traces` (OTLP), that leaves text out. `src/otel/ingest.ts` fills `input` and `output` from the first span that carries them (`gen_ai.input.messages`, `gen_ai.output.messages`, `input.value` and the other keys it knows). Everything else stays only in the `spans` table, as span attributes and events:

- the…

[Read the thread](https://github.com/iris-eval/mcp-server/issues/683) · 2026-09-26 · closed · 0 comments

### Trace search: find words inside Chinese, Japanese and Korean text

Trace search (`q` on `get_traces`, `GET /api/v1/traces`, the dashboard, the Python client; #7) indexes text with SQLite FTS5's `unicode61` tokenizer. That tokenizer splits on spaces and punctuation. Chinese and Japanese text have no spaces between words, and Korean runs words together often enough, so a whole run of CJK characters becomes one token. A search for a word inside that run finds nothing.

Example: an output reading `退款已经批准了，请查收邮件` ("the refund has been approved, please check your…

[Read the thread](https://github.com/iris-eval/mcp-server/issues/682) · 2026-09-26 · open · 0 comments

### Regex sandbox: a busy host kills trivial patterns as backtracking

On a heavily loaded host, the regex sandbox kills trivial custom-rule patterns as though they were backtracking. The rule is then reported as a budget skip, and when the rule is critical the verdict becomes unknown.

**Where.** `src/eval/rules/regex-sandbox.ts`. `sandboxedRegexTest` posts the match to the worker and then blocks in `Atomics.wait(view, 0, 0, budgetMs)`, with `REGEX_MATCH_BUDGET_MS = 100`. That wall-clock deadline covers the whole round trip:
- the message reaching the worker…

[Read the thread](https://github.com/iris-eval/mcp-server/issues/677) · 2026-09-26 · closed · 0 comments

### e2e: a stray server on port 6921 is tested instead of the current build

The Playwright suite starts its dashboard on a hardcoded port, 6921 (`tests/e2e/_constants.ts`). Outside CI, `playwright.config.ts` sets `reuseExistingServer`, so if anything is already listening on 6921 the suite quietly runs against that server instead of the build it just made. The "something" could be an Iris started with `npx @iris-eval/mcp-server --dashboard --dashboard-port 6921`, an earlier run's dashboard, or another checkout's.

A test of new code then fails, or passes, against an old…

[Read the thread](https://github.com/iris-eval/mcp-server/issues/665) · 2026-09-26 · closed · 0 comments

### Install succeeds when better-sqlite3 cannot build: make it optional and run on node:sqlite

`better-sqlite3` is a hard dependency, so when it cannot be installed the whole `npm install @iris-eval/mcp-server` fails, even though Iris already runs without it: since 0.15.0 the storage layer loads it lazily and falls back to Node's built-in `node:sqlite` (Node 22.13+) with a one-line warning.

This matters whenever a better-sqlite3 release ships without prebuilt binaries for a platform. 13.0.3 publishes none for Windows, so npm falls back to compiling with node-gyp, which needs Visual…

[Read the thread](https://github.com/iris-eval/mcp-server/issues/661) · 2026-09-26 · closed · 0 comments

### Moments: filtered pages in the newest-first order come back short and page by trace offset

`GET /api/v1/moments` in its default order (`sort_by=timestamp`) applies the `verdict`, `min_significance` and `significance_kind` filters after it reads a page of traces. As a result:

- **Pages come back short.** With a verdict filter, the route reads `limit` traces and drops the ones that don't match, so `?verdict=fail&limit=50` can return 3 moments while hundreds of failures exist further back. With a significance filter, it reads `4 × limit` traces (at most 200), which narrows the gap but…

[Read the thread](https://github.com/iris-eval/mcp-server/issues/657) · 2026-09-26 · closed · 0 comments

### Build an MCPB bundle with each release: one-click install in Claude Desktop, and the path back into Smithery

An MCPB bundle (`.mcpb`) is a packaged MCP server that installs in one step. Claude Desktop installs one by double-click, and Smithery lists a local server through one. Iris does not build one, so neither path is open; the Smithery listing was retired in 0.16.0 when its old `smithery.yaml` form stopped being the supported path.

**What ships**

- `release.yml` builds `iris-eval.mcpb` from the published package and attaches it to the GitHub release, signed like the SBOMs.
- `server.json` names…

[Read the thread](https://github.com/iris-eval/mcp-server/issues/650) · 2026-09-26 · closed · 0 comments

### An off-topic answer should fail by default when the LLM judge is on: a relevance judge for answers_the_ask

`answers_the_ask` compares the words of the answer with the words of the ask. At the shipped thresholds that failed correct answers that paraphrase, so since 0.18.0 it advises instead of blocking unless a threshold is set. The result is that an answer on the wrong topic passes at the defaults.

**What ships**

- A `relevance` template for the LLM judge (the judge has accuracy, correctness, faithfulness, helpfulness, safety and task_completed today, and no relevance template).
- When the judge…

[Read the thread](https://github.com/iris-eval/mcp-server/issues/649) · 2026-09-26 · closed · 0 comments

### Record every OpenAI and Anthropic call without the model choosing to: thin wrappers in Python and JavaScript

Today a trace reaches Iris when the agent calls `log_trace`, when it is posted to `POST /api/v1/traces`, through OpenTelemetry, or through the Claude Code capture plugin. An application that calls a model provider directly has no one-line way to send every call.

**What ships**

- `wrap_openai(client)` and `wrap_anthropic(client)` in the Python client (`iris-eval`), and the same pair in a small JavaScript package. Each call becomes a standard OpenTelemetry `gen_ai.*` span, sent to the OTLP…

[Read the thread](https://github.com/iris-eval/mcp-server/issues/648) · 2026-09-26 · closed · 0 comments

### LangChain and LangGraph integration: Python and JavaScript, proven against a real LangGraph app

Iris should score a LangChain or LangGraph agent without the agent having to call an Iris tool.

**What ships**

- **Python:** a LangChain callback handler in the `iris-eval` client. It sends each run — input, output, tool calls, token usage and latency — to `POST /api/v1/traces` with `evaluate: true`, so the run gets a verdict.
- **JavaScript:** the same handler for LangChain.js, published as `@iris-eval/langchain`. The package in `packages/langchain` today is not on npm and is rebuilt for…

[Read the thread](https://github.com/iris-eval/mcp-server/issues/642) · 2026-09-25 · closed · 0 comments

The remaining reports are on [the project's issue tracker](https://github.com/iris-eval/mcp-server/issues).
