Since v2.7.5, refresh logs curl: (22) ... 404 at err priority for optional IMDS keys on secondary ENIs
Summary
Since v2.7.5, every run of refresh-policy-routes@<iface>.service for a secondary ENI logs two lines at err priority:
ec2net[27201]: curl: (22) The requested URL returned error: 404
ec2net[27201]: curl: (22) The requested URL returned error: 404
Both come from IMDS lookups for keys that are optional and are legitimately absent on many instances. v2.7.4 and earlier did not log these. get_meta() itself already treats single-attempt lookups as "may legitimately be absent" and deliberately suppresses its own summary error for them, but curl's stderr is now forwarded to syslog unconditionally, so the suppression is bypassed.
Environment
| Item |
Value |
| amazon-ec2-net-utils |
v2.7.7, built from the tag with dpkg-buildpackage |
| OS / systemd |
Ubuntu 26.04, systemd 259 |
| Instance type |
im4gn.8xlarge (also expected on any single-card type) |
| ENIs |
Primary (ens5) plus one secondary ENI (ens6) |
| IPv4 prefix delegation |
None assigned to either ENI |
| IPv6 |
None (no global IPv6 addresses) |
What is returning 404
Querying the per-interface IMDS keys directly on the instance:
| IMDS key |
ens5 (primary) |
ens6 (secondary) |
Why |
ipv4-prefix |
404 |
404 |
No IPv4 prefix is delegated to the ENI, so the key is absent |
network-card |
404 |
404 |
Single network card instance type; IMDS omits the key |
local-ipv4s |
200 |
200 |
|
device-number |
200 |
200 |
|
subnet-ipv4-cidr-block |
200 |
200 |
|
Why only secondary ENIs are affected
Both 404ing lookups are made only for non-primary interfaces:
ipv4-prefix is queried in create_rules(), which setup_interface() calls only when ! _is_primary_interface, via prefixes=$(get_iface_imds ${ether} ${subnet_pd_key} 1 || true).
network-card is queried in _get_network_card(), which returns 0 immediately for the primary interface without calling IMDS.
So the primary ENI refreshes silently, and each secondary ENI logs two err lines per refresh. With the packaged timer that is roughly every 60s per secondary ENI, and more often if the timer interval is lowered.
Both call sites pass max_tries=1 and tolerate failure (|| true, or an empty result handled by the caller), so the 404 is an expected outcome, not an error.
Bisect
get_meta() is byte-identical in v2.7.3 and v2.7.4, and in v2.7.5, v2.7.6 and v2.7.7. The behaviour change was introduced in v2.7.5 by commit 658164e ("Fix bug where transient IMDS failure could revert secondary ip addresses."), merged in #162 for #161.
The relevant part of that change in get_meta():
local curl_opts=(
- -s
+ -sS
--max-time 5
-H "X-aws-ec2-metadata-token:${imds_token}"
-f
-A "$USER_AGENT")
...
while [ $attempts -lt $max_tries ]; do
- meta=$(curl "${curl_opts[@]}" "$url")
+ meta=$(curl "${curl_opts[@]}" "$url" \
+ 2> >(logger --id=$$ --priority "${syslog_facility}.err" --tag "$syslog_tag"))
...
+
+ # Single-attempt callers query keys that may legitimately be absent
+ # don't log those.
+ if [ "$max_tries" -gt 1 ]; then
+ error "[get_meta] ${key} failed after ${max_tries} tries"
+ fi
return 1
Before v2.7.5, curl -s discarded the error text. Since v2.7.5, -sS prints it and the process substitution sends it to syslog at err priority on every attempt, including single-attempt lookups. The new comment and max_tries check show single-attempt lookups were intended not to be logged, but that check only covers the summary error line, not curl's stderr.
Expected behaviour
Consistent with the comment in get_meta(): a failed single-attempt lookup of an optional key should not be logged as an error, as in v2.7.4 and earlier. Failures of retried lookups (max_tries > 1) should keep being logged at err, which is the useful part of #162.
Impact
- Error-priority noise on every secondary ENI, indefinitely. That is roughly 1,440 or more lines per day per secondary ENI with the packaged timer.
- It defeats per-unit log filtering such as
LogLevelMax=warning on refresh-policy-routes@.service, which is the natural way to quieten this unit while keeping real errors.
- It can trigger log-based alerting on
err, for a condition that is not an error.
Steps to reproduce
- Launch an instance with v2.7.5 or later, on a single network card instance type, with no IPv4 prefix delegation.
- Attach a secondary ENI.
- Watch
journalctl -t ec2net -f for one refresh interval.
- Observe two
curl: (22) The requested URL returned error: 404 lines at priority 3 per refresh of the secondary interface, attributed to refresh-policy-routes@<iface>.service, and none for the primary.
Proposed fix
Only forward curl's stderr to syslog when retrying (max_tries > 1), and discard it for single-attempt lookups, matching the existing summary error suppression. PR to follow, with regression tests in tests/imds.bats.
Since v2.7.5, refresh logs
curl: (22) ... 404at err priority for optional IMDS keys on secondary ENIsSummary
Since v2.7.5, every run of
refresh-policy-routes@<iface>.servicefor a secondary ENI logs two lines aterrpriority:Both come from IMDS lookups for keys that are optional and are legitimately absent on many instances. v2.7.4 and earlier did not log these.
get_meta()itself already treats single-attempt lookups as "may legitimately be absent" and deliberately suppresses its own summary error for them, but curl's stderr is now forwarded to syslog unconditionally, so the suppression is bypassed.Environment
dpkg-buildpackageim4gn.8xlarge(also expected on any single-card type)ens5) plus one secondary ENI (ens6)What is returning 404
Querying the per-interface IMDS keys directly on the instance:
ens5(primary)ens6(secondary)ipv4-prefixnetwork-cardlocal-ipv4sdevice-numbersubnet-ipv4-cidr-blockWhy only secondary ENIs are affected
Both 404ing lookups are made only for non-primary interfaces:
ipv4-prefixis queried increate_rules(), whichsetup_interface()calls only when! _is_primary_interface, viaprefixes=$(get_iface_imds ${ether} ${subnet_pd_key} 1 || true).network-cardis queried in_get_network_card(), which returns0immediately for the primary interface without calling IMDS.So the primary ENI refreshes silently, and each secondary ENI logs two
errlines per refresh. With the packaged timer that is roughly every 60s per secondary ENI, and more often if the timer interval is lowered.Both call sites pass
max_tries=1and tolerate failure (|| true, or an empty result handled by the caller), so the 404 is an expected outcome, not an error.Bisect
get_meta()is byte-identical in v2.7.3 and v2.7.4, and in v2.7.5, v2.7.6 and v2.7.7. The behaviour change was introduced in v2.7.5 by commit 658164e ("Fix bug where transient IMDS failure could revert secondary ip addresses."), merged in #162 for #161.The relevant part of that change in
get_meta():Before v2.7.5,
curl -sdiscarded the error text. Since v2.7.5,-sSprints it and the process substitution sends it to syslog aterrpriority on every attempt, including single-attempt lookups. The new comment andmax_triescheck show single-attempt lookups were intended not to be logged, but that check only covers the summaryerrorline, not curl's stderr.Expected behaviour
Consistent with the comment in
get_meta(): a failed single-attempt lookup of an optional key should not be logged as an error, as in v2.7.4 and earlier. Failures of retried lookups (max_tries > 1) should keep being logged aterr, which is the useful part of #162.Impact
LogLevelMax=warningonrefresh-policy-routes@.service, which is the natural way to quieten this unit while keeping real errors.err, for a condition that is not an error.Steps to reproduce
journalctl -t ec2net -ffor one refresh interval.curl: (22) The requested URL returned error: 404lines at priority 3 per refresh of the secondary interface, attributed torefresh-policy-routes@<iface>.service, and none for the primary.Proposed fix
Only forward curl's stderr to syslog when retrying (
max_tries > 1), and discard it for single-attempt lookups, matching the existing summary error suppression. PR to follow, with regression tests intests/imds.bats.