Skip to content

Fold path noise before an intercept reads the name - #398

Merged
jserv merged 1 commit into
sysprog21:mainfrom
jotpalch:pr-e1-path-fold
Oct 1, 2026
Merged

jserv merged 1 commit into
sysprog21:mainfrom
jotpalch:pr-e1-path-fold

Conversation

@jotpalch

@jotpalch jotpalch commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Linux steps over // runs and . components in the walk every path syscall shares, so //sys/bus, /./sys/bus and /sys/./bus are /sys/bus to all of them. The intercepts match literal prefixes against the name as the guest wrote it, so a respelled name reaches none of them and the host answers instead.

stat("/proc/self/status")     a regular file
stat("/proc//self/status")    ENOENT
chdir("/dev/./bus/usb/002")   succeeds
stat("/dev/./bus/usb/002")    ENOENT

This is not specific to the USB layer. /proc, /dev/shm, /dev/pts, the CPU subtree of /sys, and the synthesized /etc/passwd and /etc/mtab all answer the same way.

Change

path_translate_at() folds an absolute name once, before anything reads it. No intercept folds for itself, and the next one added does not have to remember to.

Two things stay as written:

  • ... Linux applies it to what the component before it resolved to, which is not a lexical question.
  • A . that is the last component. That is where Linux gives it a meaning of its own: rmdir("d/.") is EINVAL where rmdir("d/") removes d.

What a final . or a trailing slash means to a lookup is a requirement that the name resolve to a directory. The host walk applies it for itself. The intercepts match names literally, so stat("/proc/self/.") and stat("/dev/pts/.") fell through to the host as ENOENT where Linux answers with the directory. proc_intercept_open(), proc_intercept_stat_at() and proc_intercept_readlink() now take the requirement off the name once, dispatch on the bare name, and enforce it on the answer: a served directory answers as itself, anything else answers ENOTDIR, and readlink of a served directory answers EINVAL. The gates, chdir, statfs and the fd magic links read the bare name for the same reason.

A relative name is left alone. The resolvers that join one to a stamped descriptor fold as they join, and the rest goes to the host, which folds for itself.

Validation

tests/test-path-fold.c asks sixteen names, one from each owner the namespace has, at nine entry points, under the canonical spelling and six respellings. Each respelling is held to the canonical answer, so nothing is recorded. Each name is then asked with a final . and with a trailing slash, and held to what Linux makes of them: a directory answers as itself, anything else answers ENOTDIR. The creation side is asserted against the errno Linux gives. Every comparison takes the canonical answer on both sides of the respelled one and samples again when it moved, because /dev/fd and /dev/shm list live state that can change between two adjacent calls.

Holding both endings found four more places where one object had two answers, each repaired where the name is served rather than in the test: chdir onto a served file was ENOENT rather than ENOTDIR, readlink of a served directory was ENOENT rather than EINVAL, stat("/etc/mtab") was ENOENT while open served it, and open("/dev/shm/") was EACCES while stat named the directory.

With an ESP32-S3 attached, three runs each, identical:

                         answers that disagree
bcf6c5a (main)               555 of 1375
this branch                    0 of 1375
Linux 6.18.54-0-virt           0 of 1159

The Linux row is the qemu reference lane. It asks fewer because that guest has no USB device, so the three discovered names are not asked.

The test is registered in tests/test-matrix.sh, and make check runs it under ELFUSE_USB_FIXTURE=1 so the USB names are present on a machine with no device.

Cost

A scan of the name at the fold, and one more at each gate and intercept entry point that reads the bare name. bench-hot-guard, nine rounds interleaved so both builds met the same machine, medians in ns/op:

                 bcf6c5a (main)   this branch
stat-path            2171.7         2210.2
fd-create           20836.4        20877.3
getpid                 31.8           32.8

getpid takes no path and is the control.

Not in this change

A relative lookup from a host-backed directory does not reach the intercepts: fstatat(open("/"), "proc/self/status") is ENOENT on main and here, under the canonical spelling too. And under --sysroot, readlink of a link with an absolute target reads back the relative spelling the sysroot stores, so the test's symlink assertion fails there on main and here alike; the lanes run it without a sysroot. And /dev/fd lists the host process's own descriptors rather than the guest's. All three are separate defects and I will file them on their own.

Context

Split out of the ttyACM/ttyUSB alias work that #319 still tracks. That piece had grown a fold of its own at seven call sites, which is the defect above repaired one site at a time. It follows once this lands, smaller for it.


Summary by cubic

Fixes respelled absolute paths missing the intercepts and getting ENOENT from the host. Linux treats //sys/bus, /./sys/bus, and /sys/./bus as /sys/bus, but intercepts matched the name as the guest wrote it, so a respelling fell through. path_translate_at() now folds // runs and . components once, before anything reads the name; relative names are left alone, the host folding what it gets.

Directory requirement
A trailing / or final . requires the name to resolve to a directory. The three intercepts strip it, dispatch on the bare name, and enforce it: served files answer ENOTDIR, and readlink of a served directory answers EINVAL. This surfaced four places where one object had two answers, each fixed where served: stat("/etc/mtab") now mirrors open(), chdir onto a served file says ENOTDIR, and open("/dev/shm/") matches its stat. statfs and the gates read the bare name too, chdir publishes the directory rather than the spelling that reached it, and open("/dev/stdout/") answers ENOTDIR for a non-directory magic-link file.

Validation
Adds tests/test-path-fold.c: sixteen names at nine entry points, under the canonical spelling and six respellings, plus both endings. Registered under the USB fixture; zero of 1375 answers disagree on an ESP32-S3 (555 disagreed on main), and the qemu reference lane agrees on all 1159 it asks. Because /dev/fd lists live state, the canonical answer is sampled on both sides of each respelling and a moving sample is retaken. .. is not folded, a trailing . keeps its meaning (rmdir("d/.") stays EINVAL), and cost is one scan of the name per translation.

Written for commit ce9158a. Summary will update on new commits.

Review in cubic

Comment thread src/syscall/path.c
* two resolvers above fold what they join, and what they decline goes to
* the host as written.
*/
if (tx->guest_path[0] == '/' && path_has_foldable(tx->guest_path)) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

A . that ends the name still reaches the intercepts literally, so stat("/proc/self/."), open("/sys/devices/.", O_DIRECTORY) and stat("/dev/pts/.") miss the exact-match gates and fall through to the host or sysroot as ENOENT, where Linux answers with the directory. Keeping it in guest_path is right for rmdir("d/."), but intercept_path could drop the final /. (and a trailing /) and carry a must-be-a-directory bit, so a synthetic file answers ENOTDIR and a synthetic directory answers as itself. tests/test-path-fold.c only asks the trailing . of a host-backed directory and /dev/null, so a respelling like <canon>/. for the directory names in names[] would pin this; if it lands separately, a follow-up issue keeps it from being lost.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Confirmed and fixed, in this PR.

The three names answered as you describe, and main answers the same: the intercepts matched /proc/self/. literally and the host got it. proc_intercept_open(), proc_intercept_stat_at() and proc_intercept_readlink() now take the requirement off the name once, dispatch on the bare name, and enforce it on the answer: a served directory answers as itself, anything else answers ENOTDIR, and readlink of a served directory answers EINVAL. The gates read the bare name for the same reason, and the arms that matched /proc/ or /dev/pts/ beside the bare name are gone. guest_path keeps the . for the host walk, so rmdir("d/.") stays EINVAL.

Pinned as you suggested. Every name in names[] is now also asked as <canon>/. and <canon>/, and held to what Linux makes of them: a directory answers as itself, anything else answers ENOTDIR. Doing that found four more places where one object had two answers, each repaired where the name is served rather than in the test:

  • chdir onto a served file was ENOENT; Linux says ENOTDIR.
  • readlink of a served directory was ENOENT; Linux says EINVAL.
  • stat("/etc/mtab") was ENOENT while open served it.
  • open("/dev/shm/") was EACCES while stat named the directory.

Validation, three runs each, identical:

                         answers that disagree
bcf6c5a (main)               555 of 1375
this branch                    0 of 1375
Linux 6.18.54-0-virt           0 of 1159

One CI red on the way, mine. a9cbb5d failed Runtime (Release) on getdents /dev/fd/: 13 entries, want 12. On elfuse /dev/fd lists the host process's own descriptors, and a second elfuse running beside it moves that set, so two adjacent listings can differ for no reason a spelling gives. The test now takes the canonical answer on both sides of every respelled one and samples again when the canonical answer itself moved; a respelling that differs from a canonical answer that held still fails as before. Three instances at a time, interleaved:

                         runs that failed
previous test                11 of 45     (every one on a /dev/fd listing)
this test                     0 of 45

The first push carried the same exposure under the plain respellings and CI did not happen to hit it. That /dev/fd shows host descriptors at all is its own defect, and I will file it separately.

Linux steps over "//" runs and "." components in the walk every path
syscall shares, so //sys/bus, /./sys/bus and /sys/./bus are /sys/bus to
all of them. The intercepts match literal prefixes against the name as
the guest wrote it, so a respelled name reached none of them and the
host answered instead: stat("/proc//self/status") is ENOENT where
stat("/proc/self/status") is a file, and chdir("/dev/./bus/usb/002")
succeeds where stat of the same spelling does not. It is not one
layer's defect. /proc, /dev/shm, /dev/pts, the CPU subtree of /sys, the
synthetic USB trees and the synthesized /etc files all part the same
way.

path_translate_at folds an absolute name once, before anything reads it,
so no intercept folds for itself and the next one added has nothing to
remember. A relative name is left alone: the resolvers that join one to
a stamped descriptor fold as they join, and the rest goes to the host,
which folds for itself.

What it costs is a scan of the name at the fold and one more at each
gate and intercept entry point that reads the bare name.
bench-hot-guard, nine rounds interleaved so both builds met the same
machine, medians: stat-path 2210.2 ns/op against 2171.7 on bcf6c5a,
fd-create 20877.3 against 20836.4, and getpid, which takes no path, 32.8
against 31.8.

Two things stay as written. ".." is not folded, because Linux applies
it to what the component before it resolved to, which is not a lexical
question. And a "." that is the last component keeps its place, because
that is where Linux gives it a meaning of its own: rmdir("d/.") is
EINVAL where rmdir("d/") removes d.

What a final "." or a trailing slash means to a lookup is a requirement
that the name resolve to a directory, a final symlink followed to get
there. The host walk applies it for itself; the intercepts match names
literally, so stat("/proc/self/.") and stat("/dev/pts/.") fell through
to the host as ENOENT where Linux answers with the directory, and
stat("/proc/self/status/") did the same where Linux answers ENOTDIR.
proc_intercept_open, proc_intercept_stat_at and proc_intercept_readlink
now take the requirement off the name once, dispatch on the bare name,
and enforce it on the answer: a served directory answers as itself,
anything else answers ENOTDIR, and readlink of a served directory
answers EINVAL. The arms that matched "/proc/" or "/dev/pts/" beside the
bare name go, having nothing left to match. The gates read the bare name
for the same reason, and so do chdir, which publishes the directory
rather than the spelling that reached it, statfs, and the fd magic
links, where a pipe behind /dev/stdout has no path for the host to apply
the rule to.

Holding both endings to what Linux makes of them found four more places
where one object had two answers, each repaired where the name is
served rather than in the test: chdir onto a served file was ENOENT
rather than ENOTDIR, readlink of a served directory was ENOENT rather
than EINVAL, stat of /etc/mtab was ENOENT while open served it, and
open("/dev/shm/") was EACCES while stat named the directory.

tests/test-path-fold.c asks sixteen names, one from each owner the
namespace has, at nine entry points under the canonical spelling and
under six respellings of it, and holds each respelling to the canonical
answer. Nothing is recorded, so a name a kernel does not have agrees by
answering ENOENT under every spelling. Each name is then asked with a
final "." and with a trailing slash, and held to what Linux makes of
them: a directory answers as itself, anything else answers ENOTDIR. The
creation side is asserted against the errno Linux gives, with a symlink
target beside it, which is stored as written. With an ESP32-S3 attached,
three runs of each: 555 of 1375 answers disagree on bcf6c5a and none of
1375 do here. On Linux 6.18.54-0-virt, through the qemu reference lane,
none of 1159 do, the guest there having no USB device to discover names
from.

Some of those directories list live state. /dev/fd lists the caller's
descriptors, and on elfuse those are the host process's own, which a
second elfuse running beside it moves; /dev/shm lists what every process
of the user has made. A listing compared across two adjacent calls can
therefore differ for no reason a spelling gives, and the first push of
this change went red once on CI that way, counting 13 entries in
/dev/fd/ against 12. The canonical answer is now taken on both sides of
every respelled one, and a sample whose canonical answer moved is taken
again; a respelling that differs from a canonical answer that held still
fails as before. With three instances running at once, the earlier test
failed 11 of 45 runs, every one on a /dev/fd listing, and this one 0 of
45. That /dev/fd shows the host process's descriptors is its own defect
and is not repaired here.

Every assertion is plain Linux path resolution, so the test is
registered in tests/test-matrix.sh, and make check runs it under
ELFUSE_USB_FIXTURE=1 as well, so the synthetic USB names are present on
a machine with no device rather than agreeing by their absence. Both
lanes were run in full with it, and both floors go up by two: one for
this test and one for test-shim-sigreturn-x8, which was registered with
a single binary run behind it and left its floor for a run that
observed a lane.
@jserv
jserv merged commit efad801 into sysprog21:main Oct 1, 2026
15 checks passed
@jserv

jserv commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

Thank @jotpalch for contributing!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants