Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
15 changes: 15 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -13,8 +13,23 @@ Refactors, CI, and formatting land in the git history, not here.

## [Unreleased]

### Changed

- **Serving IDC v25** (`idc-index` 0.13.0 / `idc-index-data` 25.0.0, previously v24): 179
collections, 26 analysis results, 1,044,191 series, 99.9 TB. `GET /v3/version` and the MCP
`get_idc_version` now report `v25`. The release adds a `provenance` column to
`analysis_results_index` — a struct naming who contributed the data to IDC, who provided the
source material, who performed de-identification, and who produced the DICOM representation —
and drops `gcs_bucket_1` from `prior_versions_index`.

### Fixed

- `get_table_schema` / `GET /v3/tables/{table}` now describe struct (`RECORD`) columns by their
full field list — `analysis_results_index.provenance` and the nested
`collections_index.sources` — instead of a bare `RECORD` that named no field a caller could
select. Writing `SELECT provenance.data_contributor` in `run_sql` no longer requires guessing
the field names.

- Range filters now validate their bounds instead of failing or silently matching nothing:
numeric attributes require a number (previously an internal error), and `StudyDate` /
`SeriesDate` require a real date as `"YYYY-MM-DD"` (`"YYYYMMDD"` is normalized). Invalid
Expand Down
2 changes: 1 addition & 1 deletion dev/deployment.md
Original file line number Diff line number Diff line change
Expand Up @@ -159,7 +159,7 @@ python3 dev/smoke_openapi_examples.py "$URL"

The image bakes whatever `idc-index-data` resolves at build time. To publish a new IDC version,
**rebuild and redeploy** (steps 2–3). For reproducibility, pin the version in
[pyproject.toml](../pyproject.toml) (e.g. `idc-index==0.12.4`, which pulls a specific
[pyproject.toml](../pyproject.toml) (e.g. `idc-index==0.13.0`, which pulls a specific
`idc-index-data`) so a rebuild is deterministic; bump the pin to move IDC versions. The running
version is always reported at `/v3/version`.

Expand Down
15 changes: 14 additions & 1 deletion docs/user-guide.md
Original file line number Diff line number Diff line change
Expand Up @@ -169,6 +169,19 @@ GROUP BY 1 ORDER BY slides DESC
> `list_contains(col, 'value')`, not `=` or `LIKE`. If a query is invalid, the error response
> carries DuckDB's own message (including its "Did you mean …?" suggestions), so fix and retry.

> **Struct columns:** a few columns are *structs*, shown in the schema as their full field list
> (`STRUCT(field TYPE, …)`) rather than a plain type. Reach a field with dot notation —
> `SELECT provenance.data_contributor FROM analysis_results_index` — where `provenance` records
> who contributed the data to IDC, who provided the source material, who performed
> de-identification, and who produced the DICOM representation. A struct type ending in `[]` is
> a *list* of structs, so unnest it first:
>
> ```sql
> SELECT s.provenance.dicom_conversion_by AS converted_by, count(*) AS collections
> FROM (SELECT unnest(sources) AS s FROM collections_index)
> GROUP BY 1 ORDER BY collections DESC
> ```

> **Still BigQuery-only:** a handful of things remain outside these indices — *per-individual-segment*
> detail (each segment rather than the series-level `DISTINCT`-aggregated code lists in
> `seg_index`), DICOM SR quantitative/qualitative measurements (radiomics), and private DICOM
Expand Down Expand Up @@ -229,7 +242,7 @@ uv run idc-api # http://127.0.0.1:8000 — Swagger UI at /v3/docs

| Method & path | Purpose |
|---|---|
| `GET /v3/version` | IDC data release served (e.g. `v24`) + pinned index version, **and** this server's own software version (`api_version`, plus `build` if the deploy stamped one) |
| `GET /v3/version` | IDC data release served (e.g. `v25`) + pinned index version, **and** this server's own software version (`api_version`, plus `build` if the deploy stamped one) |
| `GET /v3/stats` | Headline totals (collections, patients, studies, series, size_TB) |
| `GET /v3/collections` | List collections (datasets) |
| `GET /v3/collections/{id}` | Collection detail: counts, modalities, license breakdown |
Expand Down
11 changes: 9 additions & 2 deletions pyproject.toml
Original file line number Diff line number Diff line change
Expand Up @@ -9,8 +9,15 @@ authors = [{ name = "Imaging Data Commons" }]
keywords = ["IDC", "DICOM", "cancer imaging", "MCP", "REST", "idc-index"]

dependencies = [
# Query engine + data
"idc-index>=0.12.5",
# Query engine + data.
# idc-index is pinned EXACTLY, unlike every other dependency here, because it transitively
# pins idc-index-data (0.13.0 -> 25.0.0) — which *is* the IDC data release this server
# advertises at /v3/version and documents in the changelog. The production image installs
# from this file alone (the Dockerfile copies pyproject.toml and src, never uv.lock), so a
# lower bound would let a rebuild of this very commit bake a later IDC release than the one
# it claims to serve. Bump this pin to move IDC versions — see dev/deployment.md
# § Updating for a new IDC release.
"idc-index==0.13.0",
"duckdb>=1.5.5",
"pyarrow>=25.0.0",
# Shared models / settings
Expand Down
2 changes: 1 addition & 1 deletion src/idc_api/core/models.py
Original file line number Diff line number Diff line change
Expand Up @@ -10,7 +10,7 @@


class VersionInfo(BaseModel):
idc_version: str = Field(..., description="IDC data release served, e.g. 'v24'.")
idc_version: str = Field(..., description="IDC data release served, e.g. 'v25'.")
idc_index_data_version: str = Field(..., description="Pinned idc-index-data package version.")
api_version: str = Field(
...,
Expand Down
53 changes: 49 additions & 4 deletions src/idc_api/core/schema.py
Original file line number Diff line number Diff line change
Expand Up @@ -19,7 +19,9 @@

import hashlib
from functools import lru_cache
from pathlib import Path

import duckdb
import idc_index_data

# Exposed SQL table name -> key in idc_index_data.INDEX_METADATA.
Expand Down Expand Up @@ -126,13 +128,51 @@ def include_token(included: list[str]) -> str:
return f"sub-{digest}"


def _column_type(c: dict) -> str:
@lru_cache(maxsize=None)
def _parquet_column_types(table: str) -> dict[str, str]:
"""``{column: DuckDB type}`` read from a table's local Parquet footer, or ``{}`` if there
isn't one.

Only used to expand ``RECORD`` columns. The upstream schema JSON describes a struct as a
bare ``RECORD`` with no ``fields``, so the field names exist nowhere else — and an agent
told only ``RECORD`` has to guess them, which is exactly what this server tells it not to
do. The Parquet footer is the one local source for them.

Returns ``{}`` for specialized indices (their ``parquet_filepath`` is ``None`` until the
index is fetched) so schema discovery keeps working un-expanded, as it must for a table
whose data this build never downloaded.
"""
try:
path = idc_index_data.INDEX_METADATA[metadata_key(table)].get("parquet_filepath")
if not path or not Path(str(path)).exists():
return {}
with duckdb.connect() as con:
rows = con.execute(
"SELECT column_name, column_type FROM (DESCRIBE SELECT * FROM read_parquet(?))",
[str(path)],
).fetchall()
return dict(rows)
except Exception: # pragma: no cover - schema discovery must not fail on a read hiccup
return {}


def _column_type(c: dict, struct_types: dict[str, str] | None = None) -> str:
"""Render a schema-JSON column type, folding the BigQuery-style ``mode`` field in:
``{type: STRING, mode: REPEATED}`` is an *array* column — ``STRING[]`` in DuckDB terms.
Dropping the mode (as we used to) advertises arrays as plain strings, which steers SQL
callers into predicates like ``col = 'x'`` / ``col LIKE ...`` that the engine rejects;
match array elements with ``list_contains(col, 'x')`` instead."""
match array elements with ``list_contains(col, 'x')`` instead.

``RECORD`` columns are expanded to the full DuckDB struct type from ``struct_types``
(e.g. ``STRUCT(data_contributor VARCHAR, ...)``), since ``RECORD`` alone names no field to
select. That type already encodes repetition as a trailing ``[]``, so the ``mode`` is not
applied on top of it. Without a local Parquet to read, it stays ``RECORD``/``RECORD[]``.
"""
t = c.get("type", "")
if t == "RECORD" and struct_types:
expanded = struct_types.get(c.get("name", ""))
if expanded:
return expanded
return f"{t}[]" if c.get("mode") == "REPEATED" else t


Expand Down Expand Up @@ -161,13 +201,18 @@ def table_schema(table: str) -> dict:
repointed inward via ``TABLE_DESCRIPTION_OVERRIDES``)."""
meta = idc_index_data.INDEX_METADATA[metadata_key(table)]
schema = meta.get("schema", {}) or {}
cols = schema.get("columns", [])
# Read the Parquet only when a struct column actually needs expanding.
struct_types = (
_parquet_column_types(table) if any(c.get("type") == "RECORD" for c in cols) else {}
)
columns = [
{
"name": c["name"],
"type": _column_type(c),
"type": _column_type(c, struct_types),
"description": c.get("description", "") or "",
}
for c in schema.get("columns", [])
for c in cols
]
return {
"name": table,
Expand Down
6 changes: 3 additions & 3 deletions src/idc_api/core/services/citations.py
Original file line number Diff line number Diff line change
Expand Up @@ -28,10 +28,10 @@
# Batch resolution. Every IDC *dataset* DOI is DataCite-registered (TCIA 10.7937, Zenodo
# 10.5281), and DataCite's list endpoint honours the same content negotiation as a single-DOI
# resolve — so one request returns N formatted citations instead of N round-trips. This matters:
# an unfiltered cohort spans every DOI in the archive (237 at IDC v24), which one-at-a-time meant
# 237 serial network calls holding a worker for minutes.
# an unfiltered cohort spans every DOI in the archive (242 at IDC v25), which one-at-a-time meant
# 242 serial network calls holding a worker for minutes.
_DATACITE_DOIS_URL = "https://api.datacite.org/dois"
# All 237 DOIs in a single OR-query URL gets an HTTP 414 from DataCite; 50 keeps the URL near
# All 242 DOIs in a single OR-query URL gets an HTTP 414 from DataCite; 50 keeps the URL near
# 2 KB and collapses the whole archive into five requests.
_BATCH_CHUNK = 50
# Turtle is excluded on purpose: concatenated RDF can't be split back into per-DOI entries.
Expand Down
8 changes: 6 additions & 2 deletions src/idc_api/mcp/server.py
Original file line number Diff line number Diff line change
Expand Up @@ -239,7 +239,7 @@ def _filters(terms: dict | None, ranges: dict | None) -> CohortFilters:
@mcp.tool()
@guard
def get_idc_version() -> dict:
"""Return the IDC data release served (e.g. 'v24') and pinned idc-index version, plus this
"""Return the IDC data release served (e.g. 'v25') and pinned idc-index version, plus this
server's own software version (`api_version`, and `build` if the deploy stamped one). Call
this to confirm which IDC version your answers are based on — and which build of the server
produced them."""
Expand Down Expand Up @@ -568,7 +568,11 @@ def get_licenses(terms: dict | None = None, ranges: dict | None = None) -> dict:
`index JOIN seg_index ON seg_index.segmented_SeriesInstanceUID = index.SeriesInstanceUID`,
filtered with `list_contains(seg_index.SegmentedPropertyType_CodeMeanings, 'X')`. Columns typed
`STRING[]` (e.g. the `*_CodeMeanings` columns) are arrays — match elements with
`list_contains(col, 'value')`, not `=` or `LIKE`. Call `get_table_schema(table)` for
`list_contains(col, 'value')`, not `=` or `LIKE`. Columns shown as `STRUCT(field TYPE, …)` are
structs — reach a field with dot notation (`provenance.data_contributor` on
`analysis_results_index`, naming who contributed/de-identified/converted the data); a trailing
`[]` means a list of structs, so unnest first (`SELECT unnest(sources) AS s FROM
collections_index` then `s.provenance.data_contributor`). Call `get_table_schema(table)` for
exact columns. Still BigQuery-only: per-individual-segment detail, SR radiomics measurements,
and private DICOM elements — point the user to `idc-index` + BigQuery for those.

Expand Down
8 changes: 4 additions & 4 deletions src/idc_api/rest/app.py
Original file line number Diff line number Diff line change
Expand Up @@ -372,8 +372,8 @@ def health():
"illustrative": {
"summary": "Illustrative — not live values",
"value": {
"idc_version": "v24",
"idc_index_data_version": "24.0.0",
"idc_version": "v25",
"idc_index_data_version": "25.0.0",
"api_version": "3.0.0",
"build": "a1b2c3d",
},
Expand All @@ -385,7 +385,7 @@ def health():
},
)
def version():
"""Report the IDC data release served (e.g. `v24`) and the pinned idc-index-data
"""Report the IDC data release served (e.g. `v25`) and the pinned idc-index-data
version, plus this server's own software `api_version` (and `build` stamp, if the deploy
set one). Use it to confirm which IDC version — and which build of this server —
produced a given result."""
Expand All @@ -404,7 +404,7 @@ def version():
"illustrative": {
"summary": "Illustrative — not live values",
"value": {
"idc_version": "v24",
"idc_version": "v25",
"collections": 187,
"analysis_results": 42,
"patients": 68000,
Expand Down
6 changes: 3 additions & 3 deletions tests/test_citations.py
Original file line number Diff line number Diff line change
@@ -1,7 +1,7 @@
"""Citation resolution: batched where possible, complete regardless.

An unfiltered cohort spans every DOI in IDC (237 at v24). One content-negotiation request each
meant 237 serial round-trips holding a worker for minutes, so DataCite's list endpoint — which
An unfiltered cohort spans every DOI in IDC (242 at v25). One content-negotiation request each
meant 242 serial round-trips holding a worker for minutes, so DataCite's list endpoint — which
honours the same content negotiation and covers every IDC dataset DOI (TCIA 10.7937, Zenodo
10.5281) — resolves them in chunks instead.
"""
Expand Down Expand Up @@ -106,7 +106,7 @@ def test_turtle_is_not_batched(svc):


def test_chunking_keeps_the_url_short(svc):
"""237 DOIs in a single OR-query URL gets an HTTP 414 from DataCite, so chunk them."""
"""242 DOIs in a single OR-query URL gets an HTTP 414 from DataCite, so chunk them."""
many = [f"10.7937/fake-{i}" for i in range(120)]
rec = _Recorder(batch_text="")
CitationsService._resolve(svc(rec), many, "text/x-bibliography", "apa", 30.0)
Expand Down
30 changes: 30 additions & 0 deletions tests/test_parity.py
Original file line number Diff line number Diff line change
Expand Up @@ -91,3 +91,33 @@ async def test_clinical_parity(ctx, client, parse_mcp):
rest_rows = client.get(f"/v3/clinical/tables/{table}/rows?max_rows=5").json()
mcp_rows = parse_mcp(await mcp.call_tool("get_clinical_table", {"table": table, "max_rows": 5}))
assert core_rows == rest_rows == mcp_rows


async def test_table_schema_parity_including_struct_columns(ctx, client, parse_mcp):
"""get_table_schema agrees across core, REST, and MCP — struct columns included.

Worth pinning in both adapters rather than only in the core service: an expanded struct
renders as a long type string, and `collections_index.sources` embeds a double-quoted
identifier (`"Access" VARCHAR`). That is exactly the kind of payload a serialization
change could mangle on one surface and not the other.
"""
for table in ("analysis_results_index", "collections_index"):
core = ctx.query.get_table_schema(table).model_dump(mode="json")
rest = client.get(f"/v3/tables/{table}").json()
mcp_out = parse_mcp(await mcp.call_tool("get_table_schema", {"table": table}))
assert core == rest == mcp_out, f"{table} schema differs across surfaces"

rest_cols = {
c["name"]: c["type"] for c in client.get("/v3/tables/collections_index").json()["columns"]
}
# the embedded quotes survive the round trip through JSON on the REST surface
assert '"Access" VARCHAR' in rest_cols["sources"]

mcp_cols = {
c["name"]: c["type"]
for c in parse_mcp(
await mcp.call_tool("get_table_schema", {"table": "analysis_results_index"})
)["columns"]
}
assert mcp_cols["provenance"].startswith("STRUCT(")
assert "data_contributor" in mcp_cols["provenance"]
Loading
Loading