# Reported issues for cern-opendata-mcp-server

Pod holds 15 of 15 GitHub reports that passed its relevance review. This can include external user reports, maintainer-confirmed bugs, and concrete feature gaps. Treat them as evidence to inspect, not a count of distinct defects.

Back to [cern-opendata-mcp-server](/mcp/cern-opendata-mcp-server).

## Most discussed

### feat(cern_opendata_search_records): category, keyword and LHCb magnet-polarity/stripping filters

### Use case

The portal classifies 66,345 datasets, nearly all simulated (CMS, DELPHI, ATLAS), by physics process in `categories` (`Higgs Physics` › `Standard Model`, `Exotica` › `Dark Matter`, ...) and tags LHCb datasets with magnet polarity and stripping stream and version. `cern_opendata_search_records` exposes none of them: `query` reaches them only through unlisted field forms (`categories.secondary:"Top physics"`), and the tool drops their facets, so a caller can neither discover the…

[Read the thread](https://github.com/cyanheads/cern-opendata-mcp-server/issues/12) · 2026-10-01 · closed · 3 comments

### bug(cern_opendata_get_records): citation drops the authors when no collaboration is set

### Server version

0.1.0

### mcp-ts-core version

0.13.10

### Runtime

Node.js

### Runtime version

Node 26.5.0

### Transport

HTTP (Streamable HTTP)

### Description

`citation.text` takes its creator from `collaboration.name` only, so a record that names authors but no collaboration gets a citation with no creator, opening with a bare `(2014).`. The portal's own "Cite as" for the same record lists the authors. Affects `cern_opendata_get_records` and the `cern-opendata://record/{recid}`…

[Read the thread](https://github.com/cyanheads/cern-opendata-mcp-server/issues/3) · 2026-10-01 · closed · 2 comments

### bug(cern_opendata_list_files): on-demand records report "This record has no files"

### Server version

0.1.0

### mcp-ts-core version

0.13.10

### Runtime

Node.js

### Runtime version

Node 26.5.0

### Transport

HTTP (Streamable HTTP)

### Description

For records whose availability is `ondemand`, the portal's record API returns no `_files`, `files` or `_file_indices`, though `distribution.number_files` and `distribution.size` are set. `cern_opendata_list_files` then answers "This record has no files.", which reads as an empty dataset rather than one whose files sit on…

[Read the thread](https://github.com/cyanheads/cern-opendata-mcp-server/issues/2) · 2026-10-01 · closed · 2 comments

### bug(cern_opendata_list_files): indexed files without a key make the whole record unreadable

### Server version

0.1.1

### mcp-ts-core version

0.13.10

### Runtime

Node.js

### Runtime version

Node 26.5.0

### Transport

stdio

### Description

`cern_opendata_list_files` rejects a whole record when one indexed file has no `key`. Every file in `atlas-160006`'s eight file indexes (309 files, the JetSet2 release) carries only `availability`, `checksum`, `filename`, `size` and `uri`, and the index entries state no `number_files` and an empty `availability`. `toCompactFile` requires…

[Read the thread](https://github.com/cyanheads/cern-opendata-mcp-server/issues/17) · 2026-10-02 · closed · 1 comment

### bug(cern_opendata_get_validated_runs): a CMS collision dataset with no list is told lists exist for CMS collision data only

### Server version

0.1.1

### mcp-ts-core version

0.13.10

### Runtime

Node.js

### Runtime version

Node 26.5.0

### Transport

stdio

### Description

`cern_opendata_get_validated_runs` explains every dataset that links no good-run list the same way: lists "exist for CMS collision data only, so simulated, non-CMS and non-collision records have none". For a CMS collision dataset that is false. Record 93950 (`/ZeroBias/Run2017E-v1/RAW`) is CMS collision data and has no list because no…

[Read the thread](https://github.com/cyanheads/cern-opendata-mcp-server/issues/16) · 2026-10-01 · closed · 1 comment

### bug(cern_opendata_search_trigger_paths): paths without the HLT_ prefix (AlCa_, DST_, output modules) are unreachable

### Server version

0.1.1

### mcp-ts-core version

0.13.10

### Runtime

Node.js

### Runtime version

Node 26.5.0

### Transport

stdio

### Description

`cern_opendata_search_trigger_paths` prepends `HLT_` to every `path` and its input pattern requires the prefix, but 365 of the portal's 4,135 CMS trigger path records (as of 2026-10-01) are named otherwise: 79 `AlCa_` and 52 `DST_` paths, 4 `DQM_`, 214 output modules (`…Output`), and 16 others such as `HLTriggerFinalPath`. None of them can…

[Read the thread](https://github.com/cyanheads/cern-opendata-mcp-server/issues/15) · 2026-10-01 · closed · 1 comment

### bug(htmlToText): a bare < in portal HTML drops text up to the next >

### Server version

0.1.1

### mcp-ts-core version

0.13.10

### Runtime

Node.js

### Runtime version

Node 26.5.0

### Transport

stdio

### Description

`htmlToText` in `src/services/cern-opendata/text.ts` treats every `<` as the start of a tag, so a bare `<` in portal HTML drops all text up to the next `>`. Selection cuts written as `|eta| < 2.4` are common in CMS dataset methodology. Record 5202 loses most of a sentence in `content[]`, while `structuredContent`, which keeps the HTML as…

[Read the thread](https://github.com/cyanheads/cern-opendata-mcp-server/issues/14) · 2026-10-01 · closed · 1 comment

### feat(cern_opendata_get_records): return variable dictionaries, physics categories, pile-up and LHCb run conditions

### Use case

`cern_opendata_get_records` and `cern-opendata://record/{recid}` drop record metadata that answers common questions directly: what the variables in a derived ML sample mean, which physics process a simulated sample belongs to, which pile-up sample it was mixed with, and which magnet polarity and stripping an LHCb dataset uses.

Related: #11, #12

### Proposed behavior

Add these Record fields from the search hits the lookup already reads, as received and absent when the record…

[Read the thread](https://github.com/cyanheads/cern-opendata-mcp-server/issues/13) · 2026-10-01 · closed · 1 comment

## Most recent

### feat(cern_opendata_get_records): cache lookup hits so deferred and body_offset walks don't re-read the portal

### Use case

Related: #11

With the 64,000-byte response budget, a `cern_opendata_get_records` call returns what fits and lists the rest under `deferred`, and a long documentation body is read in 30,000-character slices with `body_offset`. Every follow-up call looks its ids up again, so the walk re-downloads hits the previous call already read:

| Walk | Calls | Upstream bytes |
|:--|--:|--:|
| 20 `stripping21-bhadron-*beauty2charmline` slugs, following `deferred` to the end | 20 | ~9.5 MB…

[Read the thread](https://github.com/cyanheads/cern-opendata-mcp-server/issues/18) · 2026-10-02 · open · 0 comments

### bug(cern_opendata_get_records): no response budget across records, and doc bodies past 30,000 characters are unreachable

### Server version

0.1.1

### mcp-ts-core version

0.13.10

### Runtime

Node.js

### Runtime version

Node 26.5.0

### Transport

HTTP (Streamable HTTP)

### Description

`cern_opendata_get_records` cuts each documentation body at 30,000 characters but never bounds the response: 20 long LHCb stripping pages return 628,420 bytes of structuredContent and 622,551 of content[] text, and ordinary records add up too (a CMS collision dataset costs a median 8 KB). The rest of a cut body is…

[Read the thread](https://github.com/cyanheads/cern-opendata-mcp-server/issues/11) · 2026-10-01 · closed · 1 comment

### bug(cern_opendata_list_files): records with the largest manifests cannot be listed within the call budget

### Server version

0.1.1

### mcp-ts-core version

0.13.10

### Runtime

Node.js

### Runtime version

Node 26.5.0

### Transport

HTTP (Streamable HTTP)

### Description

`cern_opendata_list_files` reads the whole record (`GET /api/records/{recid}`) even when `index` names one file index, and cuts each attempt at 30 s within the 50 s call budget. Record 24464 is 16.3 MB (70 indexes, 32,618 files); once the portal needs over 30 s to send it, the retry restarts the download with about 19 s left…

[Read the thread](https://github.com/cyanheads/cern-opendata-mcp-server/issues/10) · 2026-10-01 · closed · 1 comment

### bug(cern_opendata_search_records): malformed query answered with portal 500 surfaces as a generic upstream error

### Server version

0.1.1

### mcp-ts-core version

0.13.10

### Runtime

Node.js

### Runtime version

Node 26.5.0

### Transport

HTTP (Streamable HTTP)

### Description

The portal answers some malformed `query` strings with HTTP 500 instead of its 400 `The syntax of the search query is invalid.`, varying per request: `(foo` drew 500 on 6 of 12 identical requests, `title:(` on 5 of 8. The service retries a 500, so when all three attempts get one the caller receives a generic…

[Read the thread](https://github.com/cyanheads/cern-opendata-mcp-server/issues/9) · 2026-10-01 · closed · 1 comment

### bug(cern_opendata_search_trigger_paths): titles naming several datasets leave the suffix in path; notices ignore year

### Server version

0.1.1

### mcp-ts-core version

0.13.10

### Runtime

Node.js

### Runtime version

Node 26.5.0

### Transport

HTTP (Streamable HTTP)

### Description

218 CMS trigger records from 2011–2013 name two or three primary datasets in their title, such as `High-Level Trigger path information HLT_Mu17_Mu8 (DoubleMu, DoubleMuParked datasets)`. The title parser knows only the singular ` ({Primary} dataset)` suffix, so `path` keeps the whole suffix and `dataset` is absent. The…

[Read the thread](https://github.com/cyanheads/cern-opendata-mcp-server/issues/8) · 2026-10-01 · closed · 1 comment

### bug(cern_opendata_search_records): sort mostrecent returns oldest first, title_desc returns A-Z

### Server version

0.1.1

### mcp-ts-core version

0.13.10

### Runtime

Node.js

### Runtime version

Node 26.5.0

### Transport

HTTP (Streamable HTTP)

### Description

`cern_opendata_search_records` describes `sort: "mostrecent"` as newest first, but it returns the oldest records first; `title_desc` (Z-A) returns A-Z, the same order as `title`. The portal's REST API takes a sort's direction only from a `-` prefix and ignores the `default_order` its sort options declare, so `mostrecent`…

[Read the thread](https://github.com/cyanheads/cern-opendata-mcp-server/issues/7) · 2026-10-01 · closed · 1 comment

### bug(recid): records with prefixed recids (atlas-160006, cms-93956) are unreachable by recid

### Server version

0.1.1

### mcp-ts-core version

0.13.10

### Runtime

Node.js

### Runtime version

Node 26.5.0

### Transport

HTTP (Streamable HTTP)

### Description

The portal now mints recids as `{experiment}-{number}` (its release validator requires the form for new records); seven records carry one as of 2026-10-01, including `atlas-160006` (JetSet2) and `cms-93956`. `cern_opendata_search_records` returns that `recid`, but every recid input accepts only 1–12 digits:…

[Read the thread](https://github.com/cyanheads/cern-opendata-mcp-server/issues/6) · 2026-10-01 · closed · 1 comment

The remaining reports are on [the project's issue tracker](https://github.com/cyanheads/cern-opendata-mcp-server/issues).
