Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
13 changes: 10 additions & 3 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -24,10 +24,17 @@ Implemented against vCenter 8 / ESXi 8. Default transport is `nbdssl`
- `VixDiskLib_Open` (datastore path, read-only or read-write)
- `VixDiskLib_Read` (optional ``skip_decompression`` packs FastLZ extras)
- `VixDiskLib_Write`
- `VixDiskLib_QueryAllocatedBlocks` (allocated-block bitmap; see
`docs/nfc_read.md`)

Not implemented: compression open flags other than FastLZ, CBT /
allocated-block queries, disk geometry (`DDB_GET`), encrypted disks,
and direct ESXi `ha-nfc` without vCenter `vpxa-nfc`.
Reading/writing a snapshot delta file directly (and running
`query_allocated_blocks` against it) already works — `NFC_DELTA_DISK`
turned out to be an optional VMFS-only VDDK client optimization, not a
correctness requirement (see `docs/reverse_engineering_procedure.md`).

Not implemented: compression open flags other than FastLZ, CBT,
disk geometry (`DDB_GET`), encrypted disks, and direct ESXi `ha-nfc`
without vCenter `vpxa-nfc`.

Requires Python 3.10 or later.

Expand Down
8 changes: 7 additions & 1 deletion docs/nfc_open.md
Original file line number Diff line number Diff line change
Expand Up @@ -247,9 +247,15 @@ I/O: `docs/nfc_read.md`, `docs/nfc_write.md`, and
## What is still VDDK-only

- `DDB_GET` / geometry / zlib and skipz compression / encryption keys
- `NFC_DELTA_DISK`, change-block tracking
- change-block tracking
- Host-switch (`NFC_AIO_SWITCH_HOST_*`)
- Direct ESXi `ha-nfc` without vCenter `vpxa-nfc`

Reading/writing a snapshot delta file directly, and running
`query_allocated_blocks` against it, both already work with the
existing implementation — `NFC_DELTA_DISK` turned out to be an
optional VMFS-only VDDK client optimization, not a correctness
requirement; see `docs/reverse_engineering_procedure.md`.

Reads after open are in `docs/nfc_read.md`. Writes are in
`docs/nfc_write.md`.
75 changes: 74 additions & 1 deletion docs/nfc_read.md
Original file line number Diff line number Diff line change
Expand Up @@ -189,9 +189,82 @@ uncompressed request (offsets in this read, not on disk)
buf when skip_decompression=True: extras packed densely from offset 0
```

## `VixDiskLib_QueryAllocatedBlocks` (AIO type 13)

Reverse-engineered by extending the SSL/write-hook capture (Step 13/14
technique, `docs/reverse_engineering_procedure.md`) to a ctypes call to
`VixDiskLib_QueryAllocatedBlocks` after `Open`, first with
`startSector=0` then — after the first capture's field guesses turned
out wrong — again with a non-zero `startSector` against a known
already-allocated region, to disambiguate fields that are 0 in the
degenerate zero-start case.

No SOAP or authd traffic; it is one more AIO message type in the
already-open NFC/AIO session (like `DDB_GET`).

Request (48 bytes)::

uint64 handle (from OPEN_FILE)
uint64 reserved (0)
uint64 chunk_size_bytes (chunk_size_sectors * sector_size)
uint64 start_offset_bytes (start_sector * sector_size)
uint64 chunk_count (num_sectors // chunk_size_sectors)
uint64 reserved (0)

**Field-order pitfall:** `start_offset_bytes` is at byte offset 24, not
8 — offset 8 is a reserved/always-zero field. A capture with
`startSector=0` can't tell these two apart (both read 0); only a
capture with a non-zero start distinguishes them. A first
implementation attempt put `start_offset_bytes` at offset 8 and got
`chunk_count`-many all-zero bits back for every non-zero-start query,
even for byte ranges known (from a zero-start, full-range query) to be
allocated — the server was silently ignoring the offset the client
thought it was requesting and returning an artifact of a different
misread field.

Reply: a 48-byte body (offset 32 echoes `chunk_count`) followed by a
bitmap extra, one bit per chunk (LSB-first, `1` = chunk has allocated
data), **padded up to a 4-byte boundary** — `ceil(chunk_count / 8)`
alone is correct only when that value is already a multiple of 4
(true for the `chunk_count=16384` case tested first, which is why the
padding bug wasn't caught immediately; a `chunk_count=16` query
exposed it, since `ceil(16/8)=2` bytes under-reads the real 4-byte
reply and desyncs the connection — the *next* AIO reply's header then
reads as garbage).

Both `start_sector` and `num_sectors` must be exact multiples of
`chunk_size_sectors`; the server returns an `NFC_AIO_MSG_ERROR` (type
1) reply otherwise (hit by accident during validation with a
non-aligned `start_sector`).

`openvixdisklib.nfc_open.NfcDisk.query_allocated_blocks` implements
this and run-length-merges contiguous set bits into
`AllocatedBlock(offset, length)` tuples (sectors, matching VDDK's
`VixDiskLibBlock`), exposed as
`VixDiskLibHandle.query_allocated_blocks`. Validated against the live
ESXi lab: a full-disk query, a query of a known-allocated sub-range,
and an aligned empty range all match native VDDK's own
`VixDiskLib_QueryAllocatedBlocks` output on the same disk.

### Gotcha: query on the same still-open write handle can see stale data

Writing a sector and then immediately calling
`query_allocated_blocks` **on that same open handle, without closing
it first**, can report the just-written region as *not* allocated —
the allocation metadata this call reads apparently isn't guaranteed
current until the write handle is closed. Closing after the write and
reopening (or querying from a separate handle opened after the write
completed) reports it correctly. Confirmed on both native VDDK and
this implementation — same-session-no-close showed the write as
unallocated on both, a fresh handle after close showed it correctly
on both — so this is a real server/VMFS behavior, not a bug in either
client. Real backup tools reading allocation before a read pass
naturally do this anyway (open read-only after the writer's handle
already closed), so it's unlikely to bite in practice, but do not
call `query_allocated_blocks` right after a write on the same handle
and expect it to reflect that write.
## What is still VDDK-only

- zlib and skipz NBD compression flags
- `VixDiskLib_ReadAsync` (same IO messages, different client threading)
- `VixDiskLib_QueryAllocatedBlocks` / allocation bitmaps
- `VixDiskLib_GetInfo` capacity (not required to read a known range)
66 changes: 61 additions & 5 deletions docs/reverse_engineering_procedure.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,8 +8,9 @@ This file is the **sequence of steps**, including dead ends, so later
NFC work can follow the same loop instead of rediscovering it.

Scope so far: `VixDiskLib_ConnectEx` + `VixDiskLib_Open` +
`VixDiskLib_Read` + `VixDiskLib_Write` against lab vCenter 8.0.1 /
ESXi 8, transports `nbd` and `nbdssl`. Validation method:
`VixDiskLib_Read` + `VixDiskLib_Write` + `VixDiskLib_QueryAllocatedBlocks`
against lab vCenter 8.0.1 / ESXi 8, transports `nbd` and `nbdssl`.
Validation method:
`tests/integration/` (the session-scoped `lab` fixture creates a temporary
empty VM with a 10 GiB disk and destroys it when the pytest session ends).

Expand Down Expand Up @@ -381,8 +382,63 @@ not an OPEN_FILE bit. Capture VDDK with that flag (NBD + the port-902
Replay: pip `pyfastlz` via `openvixdisklib/fastlz.py` (NFC extra is
raw FastLZ, without the wrapper's 4-byte length prefix) plus `NfcDisk`
compression on each IO. Proof:
`tests/integration/test_nfc_read_write.py` (`fastlz`) and
`tests/perf/test_compare.py`.
## Step 15 — `VixDiskLib_QueryAllocatedBlocks` (AIO type 13)

Same SSL/write-hook technique, extended to `VixDiskLib_QueryAllocatedBlocks`
after `Open`. A first capture with `startSector=0` produced a request
where two 0-valued 8-byte fields were ambiguous — either could plausibly
be `start_offset_bytes`. Implementing from that guess alone put
`start_offset_bytes` at the wrong offset (8 instead of 24) and passed
every zero-start test while silently returning wrong (all-empty)
results for any non-zero start. A second capture with a **non-zero**
`startSector` against a region already known (from the first capture)
to be allocated broke the tie and found the real field order.

A second, independent bug (bitmap reply length) was caught the same
way: `ceil(chunk_count/8)` matched the observed reply length for
`chunk_count=16384` (already a multiple of 4) but under-read and
desynced the connection for `chunk_count=16` — the reply pads the
bitmap to a 4-byte boundary. Diagnosed by testing candidate `extra_recv`
byte counts against whether the *next* AIO round-trip (a normal
`CLOSE_FILE`) completed cleanly, rather than guessing from a single
capture.

Full protocol detail, the exact request/reply layout, and both bugs:
`docs/nfc_read.md`. Implemented as
`openvixdisklib.nfc_open.NfcDisk.query_allocated_blocks`
(`AllocatedBlock` dataclass) and
`VixDiskLibHandle.query_allocated_blocks`. Validated against the live
ESXi lab: full-disk query, a known-allocated sub-range, and an aligned
empty range all match native VDDK's own output on the same disk.


## Note — `NFC_DELTA_DISK` needed no protocol work either

Investigated the backlog item "`NFC_DELTA_DISK` (reading directly
from a snapshot chain)" expecting a distinct wire message or OPEN_FILE
variant, similar to the CBT/`QueryAllocatedBlocks` split earlier.
`strings` on `libvixDiskLib.so` found `NFC_DELTA_DISK` is a **file-type
value** (like `NFC_DISK`), used by an internal VDDK client-side
heuristic ("`"%s" would probably benefit from bitmap copying, so
overriding file type to NFC_DELTA_DISK`") — a VMFS-only optimization
for very sparse redo logs, skipped entirely on NFS per an adjacent
string, and never observed to trigger in this lab's captures (no such
log line, `strings`-confirmed heuristic notwithstanding).

Verified end-to-end that reading, writing, and `query_allocated_blocks`
against an actual post-snapshot delta file all already work correctly
with the existing NFC_DISK-only implementation — no code change
needed. The one real finding from this investigation was a gotcha, not
a gap: an initial test that wrote a sector and immediately queried
allocated blocks *on the same still-open write handle* reported the
write as unallocated; closing the handle first (or opening a separate
one) reported it correctly. Reproduced identically against **native
VDDK** on the same delta file (two-process capture, since loading
native VDDK in the same process as pyVmomi segfaults on this host's
OpenSSL — see `docs/ssl_hook.md`'s Limits section), so this is real
server/VMFS behavior, not specific to either client. Documented as a
`query_allocated_blocks` caveat in `docs/nfc_read.md`.


## What to write down

Expand All @@ -405,7 +461,7 @@ OpenVixDiskLib.
Not yet reversed, same loop as above:

- `DDB_GET` / disk geometry, zlib/skipz compression, encrypted disks
- `NFC_DELTA_DISK`, CBT / `QueryAllocatedBlocks`
- CBT
- `VixDiskLib_GetInfo` capacity
- Host-switch AIO messages
- Direct ESXi `ha-nfc` without vCenter `vpxa-nfc`
12 changes: 12 additions & 0 deletions docs/ssl_hook.md
Original file line number Diff line number Diff line change
Expand Up @@ -143,3 +143,15 @@ skipped.
that is done offline on the hex log.
- It must not ship in OpenVixDiskLib. Keep it out of
the library path used by `openvixdisklib/nfc_auth.py`.
- `ctypes.CDLL` on `libvixDiskLib.so` **in the same process as
pyVmomi** can segfault, at least on this lab's Python/glibc build:
VDDK's bundled OpenSSL and the system OpenSSL pyVmomi already loaded
(for its own HTTPS) collide. Symptom: `Segmentation fault (core
dumped)`, no Python traceback. Split into two separate processes
instead — one doing pyVmomi/setup work, one doing only
`ctypes.CDLL`/native VDDK calls, handing data between them via a
file (see `docs/reverse_engineering_procedure.md`'s note on
`NFC_DELTA_DISK` for an example). This is the same underlying
conflict as `tests/integration/test_vddk.py` /
`test_crosscheck.py` needing `tox -e integration`'s isolated
subprocess env rather than running inside the main pytest process.
Loading
Loading