Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
51 changes: 51 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,56 @@
# Changelog

## Unreleased

### Added

- The **LLM helper app** (`examples/llm-helper/`): a dedicated function host for the AI nodes of the
agent-orchestration experiment, on the official Anthropic SDK (default model `claude-opus-5-5`).
`llm.chat` answers once, with optional JSON-schema structured output returned parsed as `data`;
`llm.stream` relays the model's token batches over the multi-shot reply contract, each batch
forwarded the moment it is produced and never gathered; `llm.health` reports whether a call could
be sent, with no network traffic. The contract, the `llm.*` configuration keys, the error contract
and the backend seam (AWS Bedrock through IAM is the planned second route) are in its README.
- One shared contract file, `tests/vectors/llm-helper-vectors.json`, byte-identical in mercury-nodejs:
63 cases (request validation, the exact SDK call, replies, errors, streaming) run against a fake
of the SDK in each pack, so the two helpers cannot drift apart. No token is spent, no credential
needed. A separate test fails if a token batch is held back.
- A reply that carries nothing usable is an error, never an empty success: a refusal with no text, an
empty reply, or a structured reply cut off before it is valid JSON is a 422 that names the
`stop_reason`, the tokens spent and what to raise. A stream that ends before its first token fails
with a 422 instead of an empty 200.
- Server-side refusal fallbacks (`fallbacks: "default"`) on the models that support them, switchable
with `llm.fallbacks=off`; `request_id` in every reply and terminal event; usage on the trace
record (`llm_model`, `llm_stop_reason`, `llm_input_tokens`, `llm_output_tokens`, `llm_request_id`).
- `llm.log.batches` (off by default): log each streamed batch's number, size and arrival time, never its
text, to tell the hop that holds batches back from the one that forwards them. The certification drive
used it to show that every batch the helper forwarded reached the engine edge as its own frame. It
also showed that the cadence of progressive rendering is the API's and differs by model: Haiku 4.5
streams continuously, Opus 5.5 in bursts about every 600 ms (documented in the README).

### Changed

- The AI nodes moved out of the demo. `examples/demo-app/demo_app.py` is the minimal polyglot demo
again, with no LLM code; `llm.chat` and `llm.stream` live in `examples/llm-helper/llm_helper.py` on
the same default port (8086), with the same request surface, so a graph that names those routes (the
`support-triage` graph) works unchanged.
- **READ - each example app now lives in its own folder**, with its own README and its own `resources/`:
`examples/demo-app/` and `examples/llm-helper/`. The demo moved from `examples/demo_app.py`, and its
sample configuration from `examples/resources/application.yml` to
`examples/demo-app/resources/application.yml`. Start it with
`mercury-serve examples/demo-app/demo_app.py`; the routes and the default port are unchanged.
- **READ - the contract is stricter than the demo nodes were.** A `params` key outside `provider`,
`model`, `max_tokens`, `timeout_ms`, `effort` and `stop_sequences` is a 400 (the demo forwarded any
key to the provider SDK); a `schema` on `llm.stream` is a 400 (it was ignored); on `llm.chat`,
`params.timeout_ms` bounds the whole call, retries included; provider errors read
`LLM provider error - {status} {type}: {message} (request_id …)`.
- The `llm` extra is `anthropic>=1,<2` only.

### Removed

- The Gemini provider (`google-genai`) and its selection: the helper serves Claude only, so
`-Dllm.provider=gemini` is now a 400 that names what is served.

## Version 4.12.15, 9/22/2026

The lock-step round with the engines: the pack moves from 4.12.1 to 4.12.15, the number the Java engine and the
Expand Down
2 changes: 1 addition & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -138,7 +138,7 @@ The same conventions as the engines, so a polyglot installation stays uniform:
Configuration lives in the `resources` folder, mirroring the engines:
`resources/application.yml` (or `.yaml` / `.properties`) in the working directory or next
to the application file, or an explicit `--config` path — see
[`examples/resources/application.yml`](examples/resources/application.yml) for a worked
[`examples/demo-app/resources/application.yml`](examples/demo-app/resources/application.yml) for a worked
sample. Values support `${ENV_VAR:default}` substitution. Runtime parameter overrides use the
same `-D` syntax as the Java engine and the Rust port — checked first on every read
(`AppConfig.set(key, value)` does the same programmatically, the `f:setConfig` analog):
Expand Down
2 changes: 1 addition & 1 deletion docs/guides/config-logging-actuators.md
Original file line number Diff line number Diff line change
Expand Up @@ -29,7 +29,7 @@ mercury-serve app.py -Drest.server.port=8090 -Dlog.format=compact
```

See the worked sample
[`examples/resources/application.yml`](https://github.com/Accenture/mercury-python/blob/main/examples/resources/application.yml)
[`examples/demo-app/resources/application.yml`](https://github.com/Accenture/mercury-python/blob/main/examples/demo-app/resources/application.yml)
and the full key table in the [Configuration Reference](configuration-reference.md).

## Logging — one aggregation, three presentations
Expand Down
75 changes: 75 additions & 0 deletions docs/test-reports/llm-helper-certification.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,75 @@
---
title: Test Report — The LLM helper in the Python host
summary: This pack's view of the live LLM helper certification of 2026-10-01 - the Python helper
(llm.chat, llm.stream, llm.health on the Anthropic SDK) behind the Java and Rust engines with real
Claude calls - the token-free proof, the live results, and what the Python SDK needed.
layer: reference
audience: [developer, architect]
keywords: [llm helper, llm.chat, llm.stream, claude, anthropic, streaming, certification, test report]
---

# Test Report — The LLM helper in the Python host

*This pack's view of the live certification of 2026-10-01 (UTC). The full report — all four engine and host
pairs, the three layers, the progressive-rendering evidence and the findings — is the
[engine report](https://accenture.github.io/mercury-composable/test-reports/llm-helper-certification/), which the
Rust engine's docs carry too. The helper is
[`examples/llm-helper`](https://github.com/Accenture/mercury-python/tree/main/examples/llm-helper); its README holds the contract.*

## Verified without a credential

- **86 helper tests, no token spent** (the whole suite: 187). `tests/test_llm_helper.py` runs a fake of the Python SDK
that builds the SDK's own error objects, so the status codes and messages are the real ones.
- **One contract, shared with the Node.js twin.** `tests/vectors/llm-helper-vectors.json` is byte-identical in both packs (SHA-256 `1f4823d9…8fa0`, pinned by
a test in each) and holds 63 cases: 22 request validations, 29 chat outcomes and 12 stream outcomes, each fixing the exact SDK
call, the reply or the error, and the segments. A change that drifts one helper away from the other fails a case.
- **The tests can fail.** A helper mutated to gather the token batches and send them at the end fails the sentinel test (the
fake model refuses to produce batch *k* until the caller holds batches 0 to *k*-1) and the vector for tokens delivered before a
mid-stream error; a helper whose default model is changed fails 13 cases. The unmutated helper passes all of them.
- **Static checks:** `ruff check .` clean, `basedpyright` 0 errors on the whole repository, `pytest -q` 187 passed.

## Verified live

Behind the Java engine and behind the Rust engine, the Python helper answered every scenario: 40 results per pair and no
failed check, across a streaming service (Layer 1), an Event Script flow (Layer 2) and two graphs (Layer 3), on
`claude-opus-5-5` and, where a request named it, `claude-haiku-4-5`.

| Progressive stream | Helper batches | Edge frames | Offset, median / spread | Longest gap |
|---|---|---|---|---|
| Java → Python, Opus 5.5 | 67 | 67 | 12 / 10 ms | 1201 / 1206 ms |
| Java → Python, Haiku 4.5 | 101 | 101 | 7 / 4 ms | 191 / 191 ms |
| Rust → Python, Opus 5.5 | 50 | 50 | 4 / 9 ms | 922 / 922 ms |
| Rust → Python, Haiku 4.5 | 82 | 82 | 3 / 5 ms | 228 / 227 ms |

Every batch the helper forwarded reached the engine's HTTP edge as its own frame, within a few milliseconds and without drift:
nothing is gathered or sent once. The long gaps are the API's own pacing, the same at both ends; Haiku 4.5 streams continuously
(a batch about every 25 ms) while Opus 5.5 arrives in bursts about every 600 ms. Credential states, measured on the real SDK:
no credential gave 503 through Java on all three layers and through Rust on the graph, and an invalid key gave Anthropic's 401 with its request id when this pack's client called the helper directly. The same message came from both helpers, through every layer. All traces that touched this helper rebuild
as one connected tree.

## What the Python SDK needed

- **No credential is not an API error.** `anthropic` 1.x raises a bare `TypeError` ("Could not resolve authentication method…")
before anything is sent, so the helper maps it by its text to a 503 `LLM provider credential missing - set ANTHROPIC_API_KEY in the
environment`. A missing key is also what `llm.health` reports, from the client's own attributes and with no network traffic.
- **The SDK takes no sampling parameters.** `temperature`, `top_p` and `top_k` are gone from its signatures, and the current models
reject them, so the contract refuses them (`params` outside `provider`, `model`, `max_tokens`, `timeout_ms`, `effort` and
`stop_sequences` is a 400).
- **A stream's request id is in the response headers.** The final message of a stream carries none (`_request_id` is empty), so the helper
reads `stream.response.headers["request-id"]`; every reply and terminal event carries it.
- **`fallbacks="default"` is a typed value** on the beta namespace, and the API accepted it on Opus 5.5. The helper sends it only for
the models that support it and never on a route that cannot.
- **The deadline is `asyncio.wait_for`** around the whole call, SDK retries included; the abandoned call is cancelled, which a test
asserts. A stream has no total deadline: `timeout_ms` is its idle allowance.

## Reproduce

```bash
pip install -e '.[dev]'
pytest tests/test_llm_helper.py # token-free
export ANTHROPIC_API_KEY=... # live: the credential reaches the helper only
mercury-serve examples/llm-helper/llm_helper.py -Dlog.format=compact -Dllm.log.batches=true
```

The commands for the engines, the deploy folder and the `curl` calls for each layer are in the
[engine report](https://accenture.github.io/mercury-composable/test-reports/llm-helper-certification/#reproduce).
109 changes: 109 additions & 0 deletions examples/demo-app/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,109 @@
# Demo app

The minimal polyglot function host: seven small functions that show what a Mercury engine can reach
over Event-over-HTTP and what a function host does with a call. Run it to watch a Python function
answer an engine, and copy it as the starting point for your own app. It never calls a model and needs
no credential; the LLM routes are a separate app, [`llm-helper`](../llm-helper/README.md).

```text
caller ── REST ──> Java or Rust engine ── Event-over-HTTP ──> demo app
flow, graph or service POST /api/event this app
```

The Node.js twin ([mercury-nodejs `examples/demo-app`](https://github.com/Accenture/mercury-nodejs))
offers the same routes, `hello.node` in place of `hello.python` and without `hello.sync.chain` (a
JavaScript handler has no sync flavour), with the same behaviour.

## Run it

```bash
pip install mercury-composable
mercury-serve examples/demo-app/demo_app.py
```

The app listens on port 8086 (`rest.server.port` in `resources/application.yml`; `-Drest.server.port=8090`
overrides it). The [LLM helper](../llm-helper/README.md) defaults to 8086 as well, so give one of the two
another port when you run both on one machine. Map a route from the engine's `event-over-http.yaml`:

```yaml
event:
http:
- route: 'hello.python'
target: 'http://127.0.0.1:8086/api/event'
```

Then call it from a flow or a graph task like any local function, or ad hoc from the package itself, as
the [getting-started guide](../../docs/guides/getting-started.md) shows:

```python
import asyncio
from mercury_composable import PostOffice

async def main():
async with PostOffice(endpoint="http://127.0.0.1:8086/api/event") as po:
reply = await po.request("hello.python", body={"text": "polyglot"}, timeout_ms=5000)
print(reply.get_status(), reply.body)

asyncio.run(main())
```

```text
200 {'text': 'POLYGLOT', 'language': 'python'}
```

## The functions

| Route | Visibility | What it shows |
|---|---|---|
| `hello.python` | public | The hello world: uppercases `text` and answers `{text, language}`. A body without `text` is a 400 `missing 'text'`, the portable error contract. It annotates the call's trace with `language`. |
| `hello.declarative` | public | Echoes the body and headers it received. The target of the engines' declarative Event-over-HTTP demos. |
| `hello.chain` | public | Local composition: calls the private `demo.suffix.helper` through the in-process bus and returns its reply. |
| `hello.sync.chain` | public | The same composition from a plain `def` handler (the `requests` and NumPy world), through the sync bridge. It blocks its own worker thread, never the event loop. |
| `hello.tokens` | public | Streaming: paced messages over the multi-shot reply contract. See below. |
| `demo.suffix.helper` | private | Appends `!` to `text`. In-app only: the HTTP host answers 403 for a private route. |
| `demo.health` | private | A health check speaking the engines' interface contract (`type=info` and `type=health`). `/health` includes it through `mandatory.health.dependencies`. |

### Streaming

`hello.tokens` sends an introductory message at once, then `count` messages, each after `delay`
milliseconds, then the terminal event. Two optional headers set the pace: `delay` (default 500, clamped
to 50 to 5000) and `count` (default 5, clamped to 1 to 100). The terminal's trailing metadata echoes
`count`, `language`, `trace_id` and `my_correlation_id`, so a calling engine's edge shows both
continuity dimensions, the distributed trace and the business correlation id, end to end.

An engine consumes it progressively (`accept: text/event-stream` on the outbound event) and can render
it out its own HTTP edge, so this is the smallest engine-to-wrapper streaming demonstration. The Java
`lambda-example` and the Rust `hello-world` example drive it that way.

## Configuration

`resources/application.yml`, or `-Dkey=value` at run time.

| Key | Default | Meaning |
|---|---|---|
| `application.name` | `demo-app` | The identity in logs and on the actuator endpoints. |
| `info.app.description` | `Mercury Composable polyglot demo` | The description `/info` reports. |
| `rest.server.port` | `8086` | The Event API and actuator port. |
| `log.format` | `text` | `text`, `json` (pretty-printed) or `compact` (single-line JSONL). |
| `log.level` | `INFO` | The log level; the `LOG_LEVEL` environment variable wins. |
| `mandatory.health.dependencies` | `demo.health` | The health-check routes `/health` requires. |
| `show.env.variables`, `show.application.properties` | `LOG_LEVEL`; `application.name, rest.server.port` | The opt-in lists behind `/env`. Nothing is ever dumped wholesale. |
| `otel.forwarding` | `false` | Opt-in OpenTelemetry export. See below. |

To export the app's trace spans, run with `-Dotel.forwarding=true`. The endpoint and credential come from
the environment (`OTLP_API_ENDPOINT`, `OTLP_AUTH_HEADER`, `OTLP_TOKEN`, and optionally
`OTLP_SERVICE_NAME`), never from the file, and header names are logged, never values. The
[OpenTelemetry certification report](../../docs/test-reports/otel-dynatrace-certification.md) records a
run of this host against a live backend.

## What it does not do

It calls no model, holds no credential and keeps no state between calls. It is a demonstration of the
host's contracts, not a template for business logic.

## Tests

The demo has no unit tests of its own. The mechanisms it demonstrates are covered in `tests/`: the local
bus, private routes and the sync bridge in `tests/test_bus.py`, streaming in `tests/test_event_stream.py`.
`hello.tokens` was driven through both engines in the
[progressive-rendering interop report](../../docs/test-reports/progressive-rendering-interop.md).
Loading
Loading