Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
12 changes: 12 additions & 0 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -97,6 +97,18 @@ jobs:
# once the issue is resolved it should be able to be re-floated
# https://github.com/Azure/azure-cli/issues/32980.
# This can be removed once https://github.com/Azure/azure-cli/issues/32869 is supported.
- name: Set Python 3.13 (Windows)
if: matrix.os-name == 'Windows' && matrix.test-category == 'AzureServiceBus'
uses: actions/setup-python@v5
with:
python-version: '3.13'
- name: Install pinned Azure CLI on Python 3.13 (Windows)
if: matrix.os-name == 'Windows' && matrix.test-category == 'AzureServiceBus'
run: |
python -m pip install --upgrade pip
python -m pip install --user "azure-cli==2.64.0"
$userScripts = python -c "import sysconfig; print(sysconfig.get_path('scripts', 'nt_user'))"
echo $userScripts >> $Env:GITHUB_PATH
- name: Set Python 3.13 (Linux)
if: matrix.os-name == 'Linux' && matrix.test-category == 'AzureServiceBus'
uses: actions/setup-python@v5
Expand Down
1 change: 1 addition & 0 deletions docs/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -28,6 +28,7 @@ This section points to sources that explain why ServiceControl is designed the w
- [Retries over Azure Storage Queues transport](retries-asq-transport.md) — transport-specific retry handling
- [Data versioning design](data-versioning-design.md) — the cache-versioning invariant for API responses
- [Event log design](eventlog-design.md) — what the event log is and what it records
- [Platform health API](platform-health.md) — how ServicePulse reads internal health independently from customer custom checks

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
- [Platform health API](platform-health.md) — how ServicePulse reads internal health independently from customer custom checks
- [Platform health API](platform-health.md) — how ServicePulse can read internal health independently from customer custom checks

- [Multiple ServiceControl instances communication](multipleservicecontrolinstancescommunication.md) — how primary, audit, and monitoring instances talk to each other
- [Handling unavailable runtime dependencies](handling-unavailable-runtime-dependencies.md) — how instances react when a dependency is unavailable
- [Telemetry](telemetry.md) — telemetry configuration and emitted metrics
Expand Down
77 changes: 77 additions & 0 deletions docs/platform-health.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,77 @@
# Platform Health API

ServiceControl exposes `GET /api/platform-health` for the ServiceControl-owned data on ServicePulse's Platform Health page. The API root advertises its URL in `platform_health`. The response uses the existing snake_case JSON convention and omits unknown nullable fields.

The public motivation is [ServiceControl #5860](https://github.com/Particular/ServiceControl/issues/5860). The consumer data requirements were checked against [ServicePulse's Platform Health store](https://github.com/Particular/ServicePulse/blob/e2688743d23fe5a4b835d4cff128ad12da87d34c/src/Frontend/src/stores/PlatformHealthStore.ts) and [platform model](https://github.com/Particular/ServicePulse/blob/e2688743d23fe5a4b835d4cff128ad12da87d34c/src/Frontend/src/resources/PlatformModel.ts).

## Response

The existing `status`, `severity`, and `alerts` fields remain, with additive `instances` and `license` sections.

### Instances

`instances` contains the primary followed by every distinct configured remote, even when no check has reported or a remote cannot be reached. Remotes are ordered by stable ID, not by the order that their requests complete.

| Field | Meaning and source |
| --- | --- |
| `id` | Existing URL-derived ServiceControl instance ID; independent of display name and row position |
| `name` | Configured instance name; a never-observed remote falls back to its URI hostname |
| `kind`, `role` | `error` / `primary-error`, `error` / `remote-error`, `audit` / `remote-audit`, or `unknown` / `remote-unknown` |
| `api_url` | Request-facing primary URL, honoring forwarded scheme, host and prefix; configured remote URL with its virtual directory preserved |
| `version` | Installed local version or remote `X-Particular-Version`; absent when unknown, never replaced with the primary's version |
| `host_id` | Actual reporting host identity from the local NServiceBus host or remote configuration; absent on older remotes |
| `health` | `healthy` for reachable instances without an associated failure, `degraded` for reachable instances with failures, `unavailable` for failed probes |
| `observed_at` | UTC timestamp for the current refresh, from the injected clock |
| `metadata_observed_at` | Timestamp of the last successful metadata observation; differs from `observed_at` during an outage |
| `health_signals_status` | `reported`, `unreported`, `disabled`, or `ambiguous`; not a guarantee that every possible check has run |
| `last_reported_at` | Latest associated check timestamp, including successful reports; distinct from HTTP observation time |
| `issues` | Associated failed internal checks, with the same fields as root alerts |
| `transport_type`, `error_queue`, `error_log_queue`, `forward_error_messages` | Available transport configuration; a known `false` forwarding setting is preserved |
| `audit_queue`, `audit_log_queue`, `forward_audit_messages` | Available audit transport configuration |
Comment on lines +29 to +30

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Are these multiple fields listed in the same row? They read like the possible values for a field and not the field name themselves.

| `error_retention_period`, `audit_retention_period` | Available retention durations in the existing TimeSpan JSON format, for example `14.00:00:00` |

Primary and audit `/api/configuration` (also `/api/instance-info`) include `instance_type` and `host.host_id`. Primary configuration additionally reports `health_checks_enabled`. Older remotes without `instance_type` are identified only when their retention configuration establishes the type. A never-observed, unreachable remote is explicitly unknown, not assumed to be an audit instance.

Remote probes use the registered named HTTP clients and their query timeout. Non-success status codes, empty or malformed configuration, and connection failures do not produce healthy rows. An outage retains the last successful metadata in memory, clearly dated by `metadata_observed_at`. Other rows and the license section still return. Caller cancellation propagates instead of returning partial success. No recursive platform-health requests are made to other primaries.

### Issues and summary

Each failed check has `id`, `check_id`, `category`, `message`, `reported_at`, `instance_name`, `host`, and `host_id`. An associated issue also has `instance_id`.

Association uses case-insensitive instance name plus reporting host ID. A legacy remote without a host ID can use a name match only when there is one matching inventory row and one reporting host with that name. Ambiguous or unmatched reports remain in root `alerts` without `instance_id`; they are never assigned to several rows. Consumers should retain a place to display those unassigned alerts.

The legacy summary describes captured checks, not the whole browser-visible platform: `status` is `unknown` before any internal report, `healthy` when none are failing, and `unhealthy` when at least one is failing. Its corresponding `severity` values are `unknown`, `none`, and `error`. ServicePulse should use per-instance health for page severity and combine it with its independently observed monitoring state. The legacy summary does not account for monitoring, browser connectivity, license expiry, or available upgrades.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This seems to conflict with the health field above that gives degraded for instances with failures.

How is severity determined? I do not see anything in the table above that gives an indication of issue serverity.


Check state is process-local. Reports older than a check's latest `reported_at` are ignored; a newer successful report clears that failure. Reports do not expire: different checks have different schedules, including one-shot checks. After restart, check observations and last-known remote metadata are initially empty. `healthy` therefore means reachable without a known associated failure, not proof of complete or fresh check coverage. An unreachable process cannot report its own browser-facing unavailability in a successful response.

### License

`license.availability` is `available` after a successful refresh and `unavailable` when license details cannot be refreshed. An unavailable license never claims to be valid and does not suppress instance health.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is the license something currently checked by the Custom Checks health checks? From what I see in the task adding a license api is out of scope.

If we want to expand scope and add this it could be in a separate PR to keep the work for each new API clear.


The available summary includes `status`, `license_status`, `license_type`, `trial_license`, optional `expiration_date` and `upgrade_protection_expiration`, and `license_extension_url`. It preserves the existing license status values for subscription, trial and upgrade-protection gates. Renewal URLs share the `/api/license` mapping with `clientName=servicepulse`, including MassTransit evaluation/subscription links. `has_mass_transit_connector` reports connector presence. Customer registration, licensed products and endpoint-license metadata are not included.

## ServicePulse integration

The endpoint supplies primary/remote inventory, installed versions, configuration, issues, and the license summary. Updating this endpoint does not update the ServicePulse consumer automatically; the consumer must map `instances` and `license` into its stores and support unknown instance types and unassigned alerts.

ServicePulse continues to own:

- Its running frontend version and ServicePulse row.
- The browser-selected monitoring URL, monitoring requests, and monitoring row.
- Browser-to-primary connectivity failures, including when this endpoint cannot be reached.
- Release-feed requests, latest-version comparison, release links, upgrade badges, and outdated-only navigation state. Installed version and license validity are not a guarantee that an upgrade path is supported.
- The customer-check fetch for the support export. Export combines this response, browser-owned rows, and the existing custom-check results. Customer checks never affect platform health.

Keep the legacy consumer fallback for supported ServiceControl versions without the advertised capability. Do not interpret `401`, `403`, a timeout, or a failed response as an absent capability. The existing custom-check API, classification, notifications, and integration events remain unchanged. Audit health still arrives through the current custom-check reporting transport; this increment does not remove that dependency or introduce replacement events.

## Access

The endpoint retains `error:customchecks:view`, granted by the existing reader, writer and admin roles. No permission or authentication behavior is changed. With authentication disabled it is anonymous. With authentication and RBAC enabled, anonymous callers receive `401` and authenticated callers without a read role receive `403`.

Known shared-policy limitation: authentication enabled with RBAC disabled currently resolves named permissions to allow-all, so the expanded response, including the license summary, can be accessed anonymously. Fixing that policy is separate work. Container liveness and readiness remain separate at `/health` and `/health/ready`.

## Verification

For manual requests, use [PlatformHealth.http](../src/ServiceControl/PlatformHealth.http). Its authenticated request reads an existing bearer token from `SERVICECONTROL_ACCESS_TOKEN`; do not store credentials in the request file.

`PlatformHealthStateTests` and `PlatformHealthApiTests` cover snapshot ordering, delayed reports, serialization, source projection, identity ambiguity, offline metadata, partial failures, license mapping and cancellation. Remote-client tests cover HTTP status, malformed responses and prefixed URLs. Shared acceptance scenarios exercise the real root/configuration/health responses and preserve custom-check behavior. The multi-instance `When_inspecting_platform_health` scenario exercises real audit check delivery, issue ownership, recovery and unavailable inventory. OIDC acceptance scenarios cover the existing read-role policy.
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,7 @@ namespace ServiceControl.AcceptanceTests.Monitoring.CustomChecks
{
using System;
using System.Linq;
using System.Text.Json;
using System.Threading;
using System.Threading.Tasks;
using AcceptanceTesting;
Expand All @@ -11,6 +12,9 @@ namespace ServiceControl.AcceptanceTests.Monitoring.CustomChecks
using NServiceBus.CustomChecks;
using NUnit.Framework;
using ServiceBus.Management.Infrastructure.Settings;
using ServiceControl.Api.Contracts;
using ServiceControl.Infrastructure;
using ApiSerializerOptions = global::ServiceControl.Infrastructure.WebApi.SerializerOptions;
using CustomCheckView = global::ServiceControl.Contracts.CustomChecks.CustomCheckView;
using CheckStatus = global::ServiceControl.Persistence.Status;

Expand All @@ -28,7 +32,10 @@ public async Task Internal_checks_are_flagged_internal_and_endpoint_checks_are_n

CustomCheckView internalCheck = null;
CustomCheckView endpointCheck = null;
PlatformHealthView platformHealth = null;
RootUrls urls = null;
string wireBody = null;
string healthWireBody = null;

await Define<Context>()
.WithEndpoint<EndpointWithFailingCustomCheck>()
Expand All @@ -47,17 +54,43 @@ await Define<Context>()
wireBody = await raw.Content.ReadAsStringAsync();
}

return internalCheck != null && endpointCheck != null && wireBody != null;
if (internalCheck != null && endpointCheck != null && platformHealth == null)
{
urls = await this.TryGet<RootUrls>("/api");
using var response = await this.GetRaw("/api/platform-health");
healthWireBody = await response.Content.ReadAsStringAsync();
platformHealth = JsonSerializer.Deserialize<PlatformHealthView>(healthWireBody, ApiSerializerOptions.Default);
}

return internalCheck != null && endpointCheck != null && wireBody != null && platformHealth != null;
})
.Run();

using var healthJson = JsonDocument.Parse(healthWireBody);
var instanceJson = healthJson.RootElement.GetProperty("instances")[0];
var instance = platformHealth.Instances.Single(item => item.Role == "primary-error");

using (Assert.EnterMultipleScope())
{
Assert.That(internalCheck, Is.Not.Null, "primary internal checks report at startup; nothing was found");
Assert.That(internalCheck.Internal, Is.True);

Assert.That(endpointCheck, Is.Not.Null);
Assert.That(endpointCheck.Internal, Is.False);
Assert.That(platformHealth.Alerts, Has.None.Matches<PlatformHealthAlert>(alert => alert.CheckId == "MyCustomCheckId"));
Assert.That(urls.PlatformHealth, Does.EndWith("/api/platform-health"));
Assert.That(instance.Id, Is.EqualTo(Settings.InstanceId));
Assert.That(instance.Name, Is.EqualTo(Settings.InstanceName));
Assert.That(instance.HostId, Is.EqualTo(internalCheck.OriginatingEndpoint.HostId));
Assert.That(instance.HealthSignalsStatus, Is.EqualTo("reported"));
Assert.That(instance.Version, Is.EqualTo(ServiceControlVersion.GetFileVersion()));
Assert.That(instance.ApiUrl.TrimEnd('/'), Is.EqualTo(urls.PlatformHealth[..^"/platform-health".Length]));
Assert.That(instance.ErrorQueue, Is.EqualTo(Settings.ErrorQueue));
Assert.That(instance.ErrorRetentionPeriod, Is.EqualTo(Settings.ErrorRetentionPeriod));
Assert.That(instanceJson.GetProperty("health_signals_status").GetString(), Is.EqualTo("reported"));
Assert.That(instanceJson.GetProperty("forward_error_messages").GetBoolean(), Is.EqualTo(Settings.ForwardErrorMessages));
Assert.That(healthJson.RootElement.GetProperty("license").GetProperty("availability").GetString(), Is.EqualTo("available"));
Assert.That(platformHealth.License.LicenseStatus, Is.Not.Null.And.Not.Empty);

// What the wire actually carries:
Assert.That(wireBody, Does.Contain("\"internal\":true"), "internal checks must render internal:true on the wire");
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -82,20 +82,19 @@ await OpenIdConnectAssertions.AssertAuthConfigurationResponse(
expectedRoleBasedAuthorizationEnabled: true);
}

[Test]
public async Task Should_reject_requests_without_bearer_token()
[TestCase("/api/errors")]
[TestCase("/api/platform-health")]
public async Task Should_reject_requests_without_bearer_token(string path)
{
HttpResponseMessage response = null;

_ = await Define<Context>()
.Done(async ctx =>
{
// Use /api/errors which does NOT have [AllowAnonymous] so it should require authentication
// Note: /api is marked [AllowAnonymous] for server-to-server configuration fetching
response = await OpenIdConnectAssertions.SendRequestWithoutAuth(
HttpClient,
HttpMethod.Get,
"/api/errors");
path);
return response != null;
})
.Run();
Expand Down Expand Up @@ -123,22 +122,21 @@ public async Task Should_reject_requests_with_invalid_bearer_token()
OpenIdConnectAssertions.AssertUnauthorized(response);
}

[Test]
public async Task Should_accept_requests_with_valid_bearer_token()
[TestCase("/api/errors")]
[TestCase("/api/platform-health")]
public async Task Should_accept_requests_with_valid_bearer_token(string path)
{
HttpResponseMessage response = null;

_ = await Define<Context>()
.Done(async ctx =>
{
// The "reader" role grants every :view permission, including error:messages:view
// required by /api/errors. Without a role-bearing claim the request would be 403.
var validToken = mockOidcServer.GenerateToken(
additionalClaims: new[] { new Claim("roles", "reader") });
response = await OpenIdConnectAssertions.SendRequestWithBearerToken(
HttpClient,
HttpMethod.Get,
"/api/errors",
path,
validToken);
return response != null;
})
Expand All @@ -147,6 +145,23 @@ public async Task Should_accept_requests_with_valid_bearer_token()
OpenIdConnectAssertions.AssertAuthenticated(response);
}

[Test]
public async Task Should_forbid_platform_health_without_a_read_role()
{
HttpResponseMessage response = null;

await Define<Context>()
.Done(async _ =>
{
response = await OpenIdConnectAssertions.SendRequestWithBearerToken(
HttpClient, HttpMethod.Get, "/api/platform-health", mockOidcServer.GenerateToken());
return response != null;
})
.Run();

OpenIdConnectAssertions.AssertForbidden(response);
}

[Test]
public async Task Should_reject_requests_with_expired_token()
{
Expand Down
Original file line number Diff line number Diff line change
@@ -1,5 +1,6 @@
namespace ServiceControl.AcceptanceTests.WebApi
{
using System;
using System.IO;
using System.IO.Compression;
using System.Net;
Expand Down Expand Up @@ -29,13 +30,17 @@ await Define<Context>()
})
.Run();

using var json = JsonDocument.Parse(configuration);
using (Assert.EnterMultipleScope())
{
Assert.That(configuration, Is.EqualTo(instanceInfo),
"Both routes are the same action, and nothing would notice if they drifted apart");

Assert.That(configuration, Does.Contain(Settings.InstanceName),
"The configuration page names the instance it is describing");
Assert.That(json.RootElement.GetProperty("instance_type").GetString(), Is.EqualTo("error"));
Assert.That(json.RootElement.GetProperty("host").GetProperty("host_id").GetGuid(), Is.Not.EqualTo(Guid.Empty));
Assert.That(json.RootElement.GetProperty("health_checks_enabled").GetBoolean(), Is.False);
}
}

Expand Down
Loading
Loading