Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
42 changes: 42 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,6 +5,48 @@ All notable changes to the Apify Java client are documented in this file.
The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.0.0/), and this project
adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).

## [0.7.0] - 2026-10-02

### Added

- `Build.getImageDigest()`, mirroring the OpenAPI spec's new `Build.imageDigest` field.
- `ApifyApiException` subclasses by HTTP status, matching the reference JS client:
`InvalidRequestError` (400), `UnauthorizedError` (401), `ForbiddenError` (403), `NotFoundError`
(404), `ConflictError` (409), `RateLimitError` (429), `ServerError` (5xx).
- `DatasetClient.createItemsPublicUrl(DatasetListItemsOptions, Long, DownloadItemsFormat)` overload
to set the public URL's serialization format.
- `ScheduleInvoked` model.

### Fixed

- `ScheduleClient.getLog()` was fetching the raw response body as text; the endpoint's response is
a JSON envelope wrapping an array of log entries. It now returns `List<ScheduleInvoked>`.
- `DatasetClient.iterateItems`/the publisher driving it could repeat or silently skip items when
combined with server-side item filters (`clean`/`skipEmpty`/`skipHidden`) or `unwind`: it now
paginates by the `X-Apify-Pagination-Count` scanned-row count the API reports whenever that count
is nonzero (a reported `0` is trusted only when the page also returned nothing, since scanning zero
rows can never produce items; any other `0` falls back to the returned count, as before), not just
by the number of items returned.
- `ApifyClientBuilder`'s `baseUrl`/`publicBaseUrl` doubled the `/v2` suffix when the caller already
included it (e.g. `"https://api.apify.com/v2"` became `".../v2/v2"`).
- `ResourceContext.toSafeId` replaced only the first `/` in a resource id; it now replaces all of
them, and `encodePathSegment` rejects an empty or dot-only (`.`/`..`) path segment instead of
encoding it, matching the reference client's URL path-traversal hardening.
- Request-body compression now skips content types that already carry their own compression
(images, audio, video, archives, office/zip packages, web fonts), matching the reference client.

### Changed

- A 404 on a resource client reached through a run/task/build without an explicit id of its own
(`run.dataset()`, `run.keyValueStore()`, `run.requestQueue()`, `run.log()`, `build.log()`) now
throws `NotFoundError` from `get()`/`delete()`/`log().get()` instead of resolving to an empty
`Optional`/no-op, since the 404 is ambiguous between the parent and the sub-resource being gone.
Resources addressed by an explicit id are unaffected.
- `DatasetClient.getStatistics()` and `TaskClient.getInput()` now throw `NotFoundError` on a 404
(returning `JsonNode` directly) instead of resolving to `Optional.empty()`, for the same reason.
- Bumped `Version.API_SPEC_VERSION` to `v2-2026-10-01T153946Z` and `Version.CLIENT_VERSION` to
`0.7.0`.

## [0.6.5] - 2026-09-29

### Changed
Expand Down
38 changes: 31 additions & 7 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -23,7 +23,7 @@ Maven (Maven Central is a default repository, so no extra configuration is neede
<dependency>
<groupId>com.apify</groupId>
<artifactId>apify-client</artifactId>
<version>0.6.5</version>
<version>0.7.0</version>
</dependency>
```

Expand All @@ -35,7 +35,7 @@ repositories {
}

dependencies {
implementation 'com.apify:apify-client:0.6.5'
implementation 'com.apify:apify-client:0.7.0'
}
```

Expand Down Expand Up @@ -218,15 +218,33 @@ methods directly on top of the JDK `HttpClient`'s own `sendAsync`.

## Fetching single resources

Methods that fetch a single resource complete with an `Optional<T>`: a missing resource is reported
by an empty `Optional` rather than an exception.
Methods that fetch a single resource **addressed by an explicit id** complete with an `Optional<T>`:
a missing resource is reported by an empty `Optional` rather than an exception, and `delete()`
resolves without error if the resource is already gone.

```java
client.actor("apify/hello-world").get()
.thenAccept(actor -> actor.ifPresent(a -> System.out.println(a.getTitle())))
.join();
```

Everywhere else, a 404 throws `NotFoundError` instead: a client chained off a run or build with no
id of its own (e.g. `client.run(id).dataset()`, `client.build(id).log()`), where the missing
resource could be the parent rather than the sub-resource, and a handful of fixed sub-paths where a
missing value is not a meaningful state distinct from "the parent is gone" (`DatasetClient
.getStatistics()`, `ScheduleClient.getLog()`, `TaskClient.getInput()`, `UserClient.monthlyUsage()` /
`limits()`, `WebhookClient.test()`):

```java
try {
client.run("missing-run").dataset().get().join();
} catch (CompletionException e) {
if (e.getCause() instanceof NotFoundError) {
// Either the run or its default dataset does not exist.
}
}
```

## Error handling

Every exception this client throws for a request/transport failure is an unchecked
Expand All @@ -239,7 +257,13 @@ original exception wrapped in an unchecked `CompletionException` (`.get()` wraps
`ApifyApiException`/`ApifyTransportException`, or use `.handle(...)`/`.exceptionally(...)` to react
to it without unwrapping at all:

- `ApifyApiException` — the request reached the API, which answered with a non-success status.
- `ApifyApiException` — the request reached the API, which answered with a non-success status. The
client throws the subclass matching the response's status code, so a `catch` can branch with
`instanceof` instead of comparing `getStatusCode()` by number — `InvalidRequestError` (400),
`UnauthorizedError` (401), `ForbiddenError` (403), `NotFoundError` (404), `ConflictError` (409),
`RateLimitError` (429) or `ServerError` (5xx). Any other status is the plain `ApifyApiException`
base class, which every subclass extends, so an existing `catch (ApifyApiException e)` keeps
working unchanged.
- `ApifyTransportException` — the request never produced an API response at all (connection
failure, DNS, timeout, or a local failure preparing the request/response, e.g. compression).
`isTimeout()` reports whether the underlying cause was specifically a timeout (backed by
Expand Down Expand Up @@ -293,10 +317,10 @@ try {
The public `com.apify.client.Version` class (`import com.apify.client.Version;`) exposes two
constants:

- `Version.CLIENT_VERSION` — the semantic version of this client (`0.6.5`).
- `Version.CLIENT_VERSION` — the semantic version of this client (`0.7.0`).
- `Version.API_SPEC_VERSION` — the version of the [Apify OpenAPI specification](https://docs.apify.com/api/openapi.json)
(its `info.version` field) that this client's endpoints, parameters and models were last generated
and checked against (`v2-2026-09-28T115051Z`). It is a snapshot, not a live compatibility
and checked against (`v2-2026-10-01T153946Z`). It is a snapshot, not a live compatibility
guarantee: the client keeps working against newer, backward-compatible spec revisions, but a
feature added to the API after this snapshot has no corresponding method here yet.

Expand Down
8 changes: 5 additions & 3 deletions docs/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -76,9 +76,11 @@ transitive dependency of this client, so it is already on your classpath.
A few methods return data whose shape is not modelled by this client and is instead exposed as a
Jackson `JsonNode` (or accept an arbitrary `Object` serialized to JSON):

- Read, returning a required `JsonNode` (never absent): `me().monthlyUsage(...)`, `me().limits()`.
- Read, returning `Optional<JsonNode>` (empty when the underlying resource has none): `dataset(id).getStatistics()`,
`task(id).getInput()`, `build(id).getOpenApiDefinition()`.
- Read, returning a required `JsonNode` (never absent; a 404 throws rather than returning empty —
see [error handling](../README.md#error-handling)): `me().monthlyUsage(...)`, `me().limits()`,
`dataset(id).getStatistics()`, `task(id).getInput()`.
- Read, returning `Optional<JsonNode>` (empty when the underlying resource has none):
`build(id).getOpenApiDefinition()`.
- Write: `task(id).updateInput(...)` (itself returning a required `JsonNode`, the updated input)
and `me().updateLimits(...)` accept an arbitrary JSON-serializable value, as do
definition/`update`/`create` arguments generally — a `Map`, a `JsonNode`, or your own POJO.
Expand Down
4 changes: 3 additions & 1 deletion docs/builds.md
Original file line number Diff line number Diff line change
Expand Up @@ -23,7 +23,9 @@ builds) and a single build with `client.build(id)`.
| `log()` | A `LogClient` for the build's log. |

`Build` fields: `getId()`, `getActId()`, `getUserId()`, `getStatus()`, `getStartedAt()`,
`getFinishedAt()`, `getBuildNumber()`, `getMeta()` (`BuildMeta` — `getOrigin()`, `getClientIp()`,
`getFinishedAt()`, `getBuildNumber()`, `getImageDigest()` (`String`, nullable — the built Docker
image manifest's digest, without the `sha256:` prefix; compare two builds' digests to tell whether
their image contents differ), `getMeta()` (`BuildMeta` — `getOrigin()`, `getClientIp()`,
`getUserAgent()`), `getStats()` (`BuildStats` — `getDurationMillis()`/`getRunTimeSecs()` as `Long`,
`getComputeUnits()` as `Double`, `getImageSizeBytes()` as `Long`), `getOptions()` (`BuildOptions` —
`getUseCache()`/`getBetaPackages()` as `Boolean`, `getMemoryMbytes()`/`getDiskMbytes()` as `Long`),
Expand Down
4 changes: 2 additions & 2 deletions docs/misc.md
Original file line number Diff line number Diff line change
Expand Up @@ -78,8 +78,8 @@ Access a build's or run's log directly, or via `client.run(id).log()` / `client.

| Method | Description |
|---|---|
| `get()` / `get(LogOptions)` | The whole log as text. Completes with `Optional<String>`. |
| `stream()` / `stream(LogOptions)` | A live `InputStream` over the log (for redirection). Completes with `InputStream`. |
| `get()` / `get(LogOptions)` | The whole log as text. Completes with `Optional<String>` when addressed by an explicit id (`client.log(id)`); on `run.log()`/`build.log()` (no log id of its own), a 404 is ambiguous and throws `NotFoundError` instead — see [Fetching single resources](../README.md#fetching-single-resources). |
| `stream()` / `stream(LogOptions)` | A live `InputStream` over the log (for redirection). Completes with `InputStream`; throws on any error response, including a 404. |

`LogOptions` fields: `raw(Boolean)`, `download(Boolean)`.

Expand Down
5 changes: 4 additions & 1 deletion docs/schedules.md
Original file line number Diff line number Diff line change
Expand Up @@ -29,7 +29,7 @@ Schedule schedule = client.schedules().create(Map.of(
| Method | Description |
|---|---|
| `get()` / `update(Object)` / `delete()` | CRUD. |
| `getLog()` | The schedule's invocation log. Completes with `Optional<String>`. |
| `getLog()` | Up to the last 1000 entries of the schedule's invocation log. Completes with `List<ScheduleInvoked>`; throws `NotFoundError` if the schedule itself no longer exists. |

```java
Optional<Schedule> s = client.schedule("SCHEDULE_ID").get().join();
Expand All @@ -44,3 +44,6 @@ s.ifPresent(sched -> System.out.println(sched.getCronExpression()));
(`ScheduleNotifications`, exposing `isEmail()`). Any field not covered by a typed getter is still
available via the inherited `getExtra()` (see
[the docs index](README.md#model-fields-and-unmodeled-data-getextra)).

`ScheduleInvoked` (one entry of `getLog()`): `getMessage()`, `getLevel()` (e.g. `INFO`, `ERROR`),
`getCreatedAt()` (`Instant`).
40 changes: 21 additions & 19 deletions docs/storages.md
Original file line number Diff line number Diff line change
Expand Up @@ -38,25 +38,26 @@ any field not covered by a typed getter above is still available via the inherit

| Method | Description |
|---|---|
| `get()` / `update(Object)` / `delete()` | Metadata CRUD. |
| `get()` / `update(Object)` / `delete()` | Metadata CRUD. On a client reached through a run/task with no dataset id of its own (e.g. `run.dataset()`), a 404 is ambiguous (parent vs. sub-resource gone) and throws `NotFoundError` instead of resolving to empty/no-op — see [Fetching single resources](../README.md#fetching-single-resources). |
| `listItems(DatasetListItemsOptions)` | List items. Completes with `PaginationList<JsonNode>`. |
| `listItems(DatasetListItemsOptions, Class<T>)` | List items decoded into `T`. Completes with `PaginationList<T>`. |
| `iterateItems(DatasetListItemsOptions)` / `iterateItems(DatasetListItemsOptions, Long chunkSize)` | A lazy `Flow.Publisher<JsonNode>` over all items; the options' `limit` caps the total yielded (`null`/unset or non-positive = all), the optional `chunkSize` sets the per-request page size (omitted/`null` = server default). |
| `iterateItems(DatasetListItemsOptions, Long chunkSize, Class<T>)` | As above, decoded into `T`. Returns `Flow.Publisher<T>`. For typed iteration at the server-default page size, pass a `null` chunk size: `iterateItems(opts, null, T.class)`. |
| `downloadItems(DownloadItemsFormat, DatasetDownloadOptions)` | Serialized bytes. `DownloadItemsFormat` is one of `JSON`, `JSONL`, `CSV`, `XLSX`, `XML`, `RSS`, `HTML`. Completes with `byte[]`. |
| `pushItems(Object)` | Push a single item or a list of items. Completes with no value (`CompletableFuture<Void>`). |
| `getStatistics()` | Dataset statistics. Completes with `Optional<JsonNode>`. |
| `createItemsPublicUrl(DatasetListItemsOptions, Long expiresInSecs)` | A public (optionally signed) items URL. Completes with `String`. |

> **Server-side item filters and iteration.** The dataset-items endpoint applies `offset`/`limit` to
> the raw items and then drops those removed by a server-side filter (`skipEmpty`, `skipHidden`,
> `clean`, `simplified`), so a page can contain fewer items than requested. Because `iterateItems`
> advances the offset by the number of items actually returned, combining it with those filters over a
> multi-page dataset has two failure modes: page windows can overlap and **repeat items**, and — more
> severely — if an entire offset window is filtered out the endpoint returns an empty page, which the
> iterator treats as the end, so iteration **stops early and silently skips the remaining data** (an
> all-filtered first page yields nothing at all). Prefer paging without server-side item filters when
> iterating, or fetch pages explicitly with `listItems` and filter client-side.
| `getStatistics()` | Dataset statistics. Completes with `JsonNode`; throws `NotFoundError` if the dataset is gone (no separate "statistics absent" state). |
| `createItemsPublicUrl(DatasetListItemsOptions, Long expiresInSecs)` | A public (optionally signed) items URL, served as `json`. Completes with `String`. |
| `createItemsPublicUrl(DatasetListItemsOptions, Long expiresInSecs, DownloadItemsFormat format)` | As above, with the URL's serialization `format` set explicitly. |

> **Server-side item filters, `unwind`, and iteration.** The dataset-items endpoint applies
> `offset`/`limit` to the raw rows and then transforms them: a filter (`skipEmpty`, `skipHidden`,
> `clean`, `simplified`) can drop rows (fewer items returned than scanned), and `unwind` can split a
> row's array field into several items (more items returned than scanned). Where the API reports the
> number of rows it scanned (`X-Apify-Pagination-Count`), `iterateItems` advances by that number
> rather than by the number of items returned, so neither case repeats already-seen items, skips
> rows, nor ends iteration early. (A reported `0` is trusted only when the page also returned
> nothing, since scanning zero rows can never produce items; any other combination falls back to the
> returned count, as if the header were absent.)

```java
Dataset ds = client.datasets().getOrCreate("my-dataset").join();
Expand All @@ -83,10 +84,11 @@ adds `attachment` (`Boolean`), `bom` (`Boolean`), `delimiter` (`String`), `skipH
(`Boolean`), `xmlRoot` (`String`), `xmlRow` (`String`), `feedTitle` (`String`), `feedDescription`
(`String`).

`createItemsPublicUrl(DatasetListItemsOptions, Long expiresInSecs)` completes with a `String` URL.
If the dataset is private, the client fetches it, reads its URL-signing secret, and appends an
HMAC-SHA256 signature (bounded by `expiresInSecs`, or non-expiring when `null`); for public datasets
the URL is unsigned.
`createItemsPublicUrl(DatasetListItemsOptions, Long expiresInSecs[, DownloadItemsFormat format])`
completes with a `String` URL. If the dataset is private, the client fetches it, reads its
URL-signing secret, and appends an HMAC-SHA256 signature (bounded by `expiresInSecs`, or
non-expiring when `null`); for public datasets the URL is unsigned. The optional `format` sets the
serialization the URL serves items in (default `json` when omitted/`null`).

> **Public-URL limitation (matches the JavaScript reference client).** The public-URL builders
> (`createItemsPublicUrl`, and the key-value-store `getRecordPublicUrl` / `createKeysPublicUrl`)
Expand All @@ -108,7 +110,7 @@ datasets.

| Method | Description |
|---|---|
| `get()` / `update(Object)` / `delete()` | Metadata CRUD. |
| `get()` / `update(Object)` / `delete()` | Metadata CRUD. On a client reached through a run/task with no store id of its own (e.g. `run.keyValueStore()`), a 404 is ambiguous (parent vs. sub-resource gone) and throws `NotFoundError` instead of resolving to empty/no-op — see [Fetching single resources](../README.md#fetching-single-resources). |
| `listKeys(ListKeysOptions)` | List keys. Completes with `KeyValueStoreKeysPage`. |
| `iterateKeys(ListKeysOptions)` / `iterateKeys(ListKeysOptions, Long chunkSize)` | A lazy `Flow.Publisher<KeyValueStoreKey>` over all keys, paging with the cursor (`exclusiveStartKey`). Note: here the options' `limit` caps the **total** number of keys yielded (`null`/unset or non-positive = all), whereas for `listKeys`/`createKeysPublicUrl` the same `ListKeysOptions.limit` is a single-request page size. `chunkSize` sets the per-request page size (`null` = server default). |
| `recordExists(String key)` | Whether a record exists. Completes with `boolean`. |
Expand Down Expand Up @@ -158,7 +160,7 @@ overload here — the request-queue creation endpoint does not accept a creation

| Method | Description |
|---|---|
| `get()` / `update(Object)` / `delete()` | Metadata CRUD. |
| `get()` / `update(Object)` / `delete()` | Metadata CRUD. On a client reached through a run/task with no queue id of its own (e.g. `run.requestQueue()`), a 404 is ambiguous (parent vs. sub-resource gone) and throws `NotFoundError` instead of resolving to empty/no-op — see [Fetching single resources](../README.md#fetching-single-resources). |
| `withClientKey(String)` | A copy that identifies its requests with a stable client key (required for lock operations). |
| `listHead(Long limit)` | Requests at the head. Completes with `RequestQueueHead`. |
| `addRequest(RequestQueueRequest, boolean forefront)` | Add a request. Completes with `RequestQueueOperationInfo`. |
Expand Down
2 changes: 1 addition & 1 deletion docs/tasks.md
Original file line number Diff line number Diff line change
Expand Up @@ -28,7 +28,7 @@ Task task = client.tasks().create(Map.of(
| `start(Object input, TaskStartOptions)` | Start a task run (input overrides stored input; `null` uses it). Completes with `ActorRun`. |
| `call(Object input, TaskStartOptions, Long waitSecs)` | Start and poll until finished; does **not** stream the run's log. Completes with `ActorRun`. |
| `call(Object input, TaskCallOptions, Long waitSecs)` | As above, additionally streaming the run's log for the duration of the wait by default (matching the reference client's `call` defaulting `options.log` to `'default'`). Use `TaskCallOptions.disableLogStreaming()` to opt out, or `logOptions(StreamedLogOptions)` for a custom destination. |
| `getInput()` | The stored input. Completes with `Optional<JsonNode>`. |
| `getInput()` | The stored input. Completes with `JsonNode`; throws `NotFoundError` if the task is gone (no separate "input absent" state). |
| `updateInput(Object)` | Replace the stored input. Completes with `JsonNode`. |
| `lastRun(String status)` / `lastRun(LastRunOptions)` | A `RunClient` for the last run (see [`LastRunOptions`](actors.md#actorclient)). |
| `runs()` | Nested run collection client. |
Expand Down
2 changes: 1 addition & 1 deletion pom.xml
Original file line number Diff line number Diff line change
Expand Up @@ -6,7 +6,7 @@

<groupId>com.apify</groupId>
<artifactId>apify-client</artifactId>
<version>0.6.5</version>
<version>0.7.0</version>
<packaging>jar</packaging>

<name>Apify Java Client</name>
Expand Down
9 changes: 9 additions & 0 deletions src/main/java/com/apify/client/ApifyClient.java
Original file line number Diff line number Diff line change
Expand Up @@ -116,6 +116,15 @@ String getApiBaseUrl() {
return baseUrl;
}

/**
* Returns the fully-qualified public API base URL this client builds shareable URLs against
* (including the {@code /v2} suffix). Not part of the public API, for the same reason as {@link
* #getUserAgent()}.
*/
String getPublicApiBaseUrl() {
return publicBaseUrl;
}

// ----- Actor accessors -----------------------------------------------------

/** A client for the Actor collection (list &amp; create Actors). */
Expand Down
Loading
Loading