Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
29 changes: 23 additions & 6 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -17,7 +17,8 @@ from VDDK 8 NBD traffic; see `docs/`.

## Status

Supported and tested on **vCenter 8 / ESXi 8** (lab: 8.0.1). The
Supported and tested on **vCenter 8 / ESXi 8** (lab: 8.0.1), including a
standalone ESXi host with no vCenter. The
VixDiskLib compatibility mode is `8.0` only. VIM login requests
pyVmomi's vim25 **8.x** versions, so a newer host such as vSphere 9 stays
on 8.x SOAP instead of 9.x types.
Expand All @@ -27,14 +28,30 @@ moment.

Default transport is `nbdssl` (`nbd` is still available):

- `VixDiskLib_ConnectEx` (UID credentials)
- `VixDiskLib_ConnectEx` (UID credentials; vCenter or direct ESXi)
- `VixDiskLib_Open` (datastore path, read-only or read-write)
- `VixDiskLib_Read` (optional ``skip_decompression`` packs FastLZ extras)
- `VixDiskLib_Write`

Not implemented: compression open flags other than FastLZ, CBT /
allocated-block queries, disk geometry (`DDB_GET`), encrypted disks,
and direct ESXi `ha-nfc` without vCenter `vpxa-nfc`.
- `VixDiskLib_GetInfo` (capacity and physical geometry from the `Open`
reply; `biosGeo`/`adapterType`/`uuid` from `DDB_GET`, matching real
VDDK's cost and behavior)
- `VixDiskLib_QueryAllocatedBlocks` (allocated-block bitmap; see
`docs/nfc_read.md`)
- Changed Block Tracking: `openvixdisklib.nfc_auth.enable_change_tracking`
/ `disk_change_id` / `query_changed_disk_areas` (public VIM API, not
part of VixDiskLib itself; see `docs/cbt.md`)

Reading/writing a snapshot delta file directly (and running
`query_allocated_blocks` against it) already works — `NFC_DELTA_DISK`
turned out to be an optional VMFS-only VDDK client optimization, not a
correctness requirement (see `docs/reverse_engineering_procedure.md`).

An NFC session also survives a live vMotion of the VM being read, with
no code changes needed — even one that relocates the disk itself to a
datastore the original host couldn't reach (see `docs/host_switch.md`).

Not implemented: compression open flags other than FastLZ, and
encrypted disks.

Requires Python 3.10 or later.

Expand Down
110 changes: 110 additions & 0 deletions docs/cbt.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,110 @@
# Changed Block Tracking (CBT)

This is not a reverse-engineered NFC feature. VixDiskLib does not expose
CBT itself: `VixDiskLib_QueryAllocatedBlocks` (implemented separately;
see `docs/nfc_read.md`) reports which blocks are *allocated*
(non-sparse) within a single NFC-opened disk, not which byte ranges
*changed* between two points in time. Real backup tools
get changed-range information from vSphere's public
`VirtualMachine.QueryChangedDiskAreas` VIM call instead, used alongside
VDDK/NFC reads for the actual bytes. `openvixdisklib.nfc_auth` wraps
that public pyVmomi call directly — no capture, no wire format to
document, per the project rule to reuse pyVmomi for anything it already
exposes.

## Workflow

1. `nfc_auth.enable_change_tracking(vm)` — sets
`VirtualMachineConfigSpec.changeTrackingEnabled = True` via
`ReconfigVM_Task`. Takes effect for writes from that point forward;
it does not retroactively track earlier changes.
2. Take a snapshot (or power-cycle the VM). A disk's `changeId` is
empty until this happens.
3. `nfc_auth.disk_change_id(vm, device_key)` — reads the current
`changeId` off `VirtualDisk.backing.changeId` (for example
`"52 f3 b6 37 30 8d ea 3e-58 70 c0 fd 61 44 26 62/2"`).
4. Do backup work (VDDK/NFC reads of the disk at that point).
5. Later, take another snapshot.
6. `nfc_auth.query_changed_disk_areas(vm, new_snapshot, device_key,
change_id_from_step_3)` — returns the byte ranges written between
the two snapshots.
7. Read only those ranges via VDDK/NFC on the new snapshot's disk
chain for an incremental backup.

For an initial full backup, pass `change_id="*"` in step 6 without a
prior snapshot. **Correction from an earlier draft of this doc:**
this does *not* report the entire disk as one changed extent — see
"Wildcard `changeId='*'` reports allocated regions, not the whole
disk" below.

## Validated in this lab

Confirmed end-to-end against a temporary VM on the standalone ESXi
8.0.3 lab host (no vCenter): enabled CBT, snapshotted, wrote one
sector via `openvixdisklib.openvixdisklib` at a known offset,
snapshotted again, and called `query_changed_disk_areas` with the
first snapshot's `changeId`. The single reported extent
(`start=3932160, length=65536`, i.e. sectors 7680–7807) correctly
covered the written sector (7777). Extents were 64 KiB-aligned in this
lab's observations; that granularity is server-defined, not part of
the function's contract.

### Wildcard `changeId="*"` reports allocated regions, not the whole disk

Tested `query_changed_disk_areas(vm, snapshot, device_key, "*")` (the
initial-full-backup path, no prior snapshot needed) against a fresh
10 GiB thin-provisioned temp-VM disk. `result.length` correctly
reports the full declared virtual capacity (10737418240 bytes), but
`result.changed_areas` only covered **1 MiB** total — not the whole
disk. For a thin-provisioned disk, `"*"` reports the regions that are
actually *allocated* (backed by real data on the datastore), not the
full sparse virtual capacity; unwritten/unallocated regions have
nothing to back up regardless. A backup tool doing an initial full
backup with `"*"` should read exactly the reported extents, not assume
it needs to read `result.length` bytes.

### One large contiguous write is one extent; scattered writes are not

Wrote a single 4 MiB contiguous region plus three separate one-sector
writes at scattered offsets (same disk, one CBT interval), then
queried changed areas:

```
4 extents reported:
start= 196608 length= 65536 (64 KiB)
start= 51183616 length= 4259840 (4160 KiB) <-- covers the whole 4 MiB write as ONE extent
start= 460783616 length= 65536 (64 KiB)
start= 921567232 length= 65536 (64 KiB)
```

The 4 MiB write came back as a single extent (padded slightly beyond
4 MiB — 4259840 bytes vs. the exact 4194304 written — to the 64 KiB
tail-end granularity). Each scattered single-sector write produced its
own separate 64 KiB extent. **The extent list scales with the number
of discontiguous changed regions, not with the total volume of changed
data.** A multi-hundred-GB sequential write is still one small extent
record; thousands of scattered small writes (e.g. a busy database VM
doing random I/O across a large disk) produce thousands of extent
records in one `QueryChangedDiskAreas` response, since the API has no
pagination. Real backup tools facing that scenario typically chunk the
query with `start_offset` over fixed-size windows rather than querying
the whole disk in one call — `query_changed_disk_areas`'s
`start_offset` parameter exists for this, but nothing in this module
does the chunking loop itself; that is caller responsibility.

Disk-size scaling itself (e.g., whether extent granularity increases
for very large disks) was not tested — only reasoned about above as an
open question, not verified against a large ESXi 8 disk.

## What this does not cover

- `VixDiskLib_QueryAllocatedBlocks` (NFC-level allocated-block bitmap
within a single disk, useful for skipping sparse regions inside a
delta disk) — implemented separately, see `docs/nfc_read.md`. Pairs
naturally with CBT: `query_changed_disk_areas` says which byte
ranges changed, `query_allocated_blocks` says which parts of a
snapshot's delta disk are actually worth reading. Note its "same
still-open write handle" staleness gotcha in `docs/nfc_read.md` if
chaining a CBT-driven write with an allocation check.
- `DDB_GET` fields (`biosGeo`, `adapterType`, `uuid`) — also
implemented, see `docs/nfc_open.md`.
125 changes: 125 additions & 0 deletions docs/host_switch.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,125 @@
# Host-switch (`NFC_AIO_SWITCH_HOST_*`)

This is not a reverse-engineered NFC feature, at least not for either
scenario tested. `strings` on `libvixDiskLib.so` shows this is VDDK's
mechanism for keeping an NFC/backup session alive across a **live
vMotion** of the VM being backed up — a `SWITCHHOST_VADP` string (VADP
being VMware's official backup-API framework) and a `PreSwitchHost
callback` string carrying a full new-host descriptor (`Server IP,
Port, Session ID, SSL Thumbprint, NFC Service Endpoint`). Testing it
needed a second ESXi host in the same cluster with shared storage — a
significant infra build, documented in `docs/host_switch_lab_setup.md`.

## Investigation 1: compute-only vMotion, shared storage

With two hosts sharing an NFS datastore, kept an NFC read session
alive (native VDDK, one-second-interval reads, under the same SSL-hook
technique as every other capture) while triggering a live vMotion of
the VM mid-session via `RelocateVM_Task` (compute only — the disk's
datastore didn't change). Result: **the session was completely
unaffected**. All 40 reads across the ~40-second test succeeded,
including the ones during and immediately after the migration, with
no visible interruption, no reconnect, and no error.

## Investigation 2: combined storage + compute vMotion

Repeated with a much more aggressive scenario, to try to force a real
switch: a VM with its disk on a **host-local** VMFS datastore (only
reachable by that one host, not the target), then a single
`RelocateVM_Task` moving **both** the VM's compute *and* its disk (to
the shared NFS datastore) to the other host simultaneously — the
kind of migration that should, in principle, strand an NFC session
that was talking to the original host, since after the move that host
has no path to the file's new location at all.

Result: **still completely unaffected**. All 90 reads succeeded
through the full migration (which took noticeably longer than the
compute-only case, as expected for a real data copy), including reads
issued after `RelocateVM_Task` had fully completed and the disk was
confirmed to be at its new location on the new datastore.

## Checked the wire, both times

In both investigations, the raw capture shows only **one TCP file
descriptor** used for the NFC connection (port 902) for the entire
session — before, during, and after the migration. Only one classic
NFC handshake sequence appears anywhere in either capture (searched
for a second `PlainText` handshake message, found only the original
one). No `NFC_AIO_SWITCH_HOST_*` traffic, no reconnect, nothing.

## Why this makes sense

NFC disk access is **datastore-based, not VM/host-based** — the
client opens a path like `[datastore] vm/vm.vmdk`, and the already-open
file handle from `OPEN_FILE` apparently stays valid even when the
underlying file relocates to a completely different datastore during
an active session. Whatever redirection is needed happens entirely
below the NFC layer, transparently, for both a VM's compute moving and
its storage moving — as long as everything stays inside the same
vCenter-managed environment with the migration completing normally.

## Why it's not reachable from here at all: not a public API

Went looking for how a third-party client would even participate in a
host switch — VDDK's own strings mention a `PreSwitchHost callback`
that receives the new host's connection details, which sounded like
something OpenVixDiskLib might need to expose too. It doesn't:
`grep` across every header in the VDDK 8.0.3 SDK
(`vixDiskLib.h`, `vixDiskLibPlugin.h`, `vixMntapi.h`) for
`SwitchHost`/`Callback` finds only the documented, unrelated
completion/progress/logging callbacks. **There is no public
registration function for this mechanism anywhere in the SDK.**

This settles the question definitively, without needing to test the
one remaining scenario (the connected host itself becoming
unavailable) by disrupting a real host: whatever `NFC_AIO_SWITCH_HOST_*`
and `PreSwitchHost` actually do, they're wired into VMware's own
internal/first-party tooling (VADP), not exposed through the SDK any
third-party backup vendor — or OpenVixDiskLib — actually links
against. There is no code path by which a normal `VixDiskLib_Open`/
`Read`/`Write` client could ever trigger, observe, or need to
implement this, regardless of what happens to the underlying hosts.
Both empirical tests above already showed it doesn't fire for any
vMotion scenario reachable through the public API; this closes the
remaining theoretical gap by showing there's no public entry point for
it to fire through in the first place.

## What this does not cover

- Direct-ESXi (no vCenter) sessions were not tested for either vMotion
scenario; everything above went through vCenter. Not expected to
matter, since the conclusion (no public API surface for this
mechanism at all) is independent of which connection path is used.

## Also found along the way: NFC needs a snapshot to open a *running* VM's disk, on any datastore

Discovered by accident while setting up these tests, not something the
`NFC_AIO_SWITCH_HOST_*` investigation was looking for, but real,
general, and worth recording carefully since an earlier draft of this
document mischaracterized it as NFS-specific — it is not.

Opening a **running** (powered-on) VM's base disk directly over NFC
fails with an `NFC_ERROR` reply (`NfcFssrvrOpen` permission check, per
the binary's own strings) — confirmed against **three separate VMs**,
on **both VMFS and NFS** datastores, on both hosts, with both
read-only and read-write tickets. The same operation against the same
VM **powered off** works immediately. Taking a snapshot first
(redirecting live writes to a delta file) and opening the **parent**
disk — the same pattern already used elsewhere in this project for
reading a snapshot's parent VMDK — also works immediately, with the VM
still powered on.

This project's other integration tests have never actually hit this,
purely by accident: the shared `lab` pytest fixture creates its
temporary VM but never powers it on, so every prior capture in this
whole project (against `SLES16`, `ovdl-crypto-test`, and the fixture's
own temp VMs) was reading a disk belonging to a **powered-off** VM
without realizing that was load-bearing. Confirmed directly: powering
on `SLES16` (an otherwise ordinary, long-lived lab VM on VMFS) and
attempting the exact same read that works fine while it's off
reproduces the identical `NFC_ERROR`.

Worth remembering for any future capture, on any datastore type: if
VDDK/`openvixdisklib` returns a generic `"Unknown error"` / opaque
`NFC_ERROR` opening a disk that otherwise looks correct, check whether
the VM is powered on and whether a snapshot is needed first.
58 changes: 58 additions & 0 deletions docs/host_switch_lab_setup.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,58 @@
# Requirements for a 2-host vMotion lab

Testing anything host-switch or vMotion related needs a vCenter with two
ESXi hosts in the same cluster, plus storage both hosts can see. See
`docs/encryption_lab_setup.md` for a single-host vCenter lab; this
document only covers what changes when a second host is added. The
environment-specific details (how the hosts or the shared storage were
provisioned) are intentionally left out, as they depend on the lab.

## 1. Shared storage

vMotion needs a datastore that both hosts can see. Any shared storage
works (NFS, iSCSI, vSAN, ...). For NFS, mounting the same
`remoteHost`/`remotePath` on both hosts through
`host.configManager.datastoreSystem.CreateNasDatastore()` makes vCenter
recognize it as a single shared datastore (same `vim.Datastore` moref on
both hosts).

## 2. Second host

Any ESXi host works. When it is itself a nested VM, nested
virtualization must be enabled on the underlying hypervisor (e.g.
`host-passthrough` CPU mode on libvirt/KVM).

Join it to the same cluster as the first one, using the same calls:
`AddStandaloneHost` followed by `MoveInto_Task` on the cluster. Both
hosts should end up `connected` in the same cluster.

## 3. Enable vMotion

For a flat single-subnet lab network, the simplest option is to enable
the vMotion service on each host's existing management VMkernel adapter
(`vmk0`) instead of creating a dedicated portgroup/vSwitch:

```python
host.configManager.virtualNicManager.SelectVnicForNicType("vmotion", "vmk0")
```

Production environments should separate vMotion traffic onto its own
VMkernel adapter/VLAN.

## 4. Verify vMotion

With a VM powered on, relocate it to the other host:

```python
vm.RelocateVM_Task(spec=vim.vm.RelocateSpec(host=<target>))
```

## Gotcha: NFC needs a snapshot for a running VM's disk on NFS

Not vMotion specific, but hit while building the host-switch test case:
opening a **running** VM's disk directly over NFC on an NFS datastore
failed (`NfcFssrvrOpen` permission-check `NFC_ERROR`) regardless of the
ticket type or which host served it. It works with the VM powered off,
or with the VM powered on but reading the **parent** disk after taking a
snapshot. See `docs/host_switch.md` for details, recorded there since
it's a protocol-level finding rather than a lab-setup one.
Loading