Ship resolve image pull stalls before producer allocation #789

Open
opened 2026-09-08 13:24:46 +00:00 by PerishFire · 3 comments
Owner

Observed canonical Ship run 1651 (run ID 4795): https://git.perish.top/PerishLab/plumb/actions/runs/1651 . Marker v0.37.38-beta.2 at 1280700244; controller configuration v0.37.38-beta.1. Resolve task 11653 on runner 538ee449-eb0e-4dea-ad00-83779dfe8f82 (v12.13.2) has exposed only docker pull since 2026-09-08 13:03:31 UTC. Image is the accepted immutable Images forge digest e07cd20edac134a9e76742c9d4562459b7372f4cdb229c5e4fb0303499f1cb73, linux/amd64, 12 compressed layers totaling 663075742 bytes. No workload/platform job has started; downstream blocked states are not producer failures. Registry manifests read back successfully. Read-only Forgejo service logs show several blob responses completed 200 in 108365/161571/21196 ms and later blob requests, but there is no completed pull or job progress after over 20 minutes. A subsequent roughly 20-second service network sample transmitted only 46556 bytes (aggregate pod evidence, not per-request throughput). Root cause is not established: transport, client backpressure, extraction and runner lifecycle require separation. No runner/cluster mutation, cancellation, marker rewrite, additional dispatch or stable activation was performed. Track actual producer bootstrap evidence and bounded/progressive diagnostics; retry the same marker only after cause is established or the current run concludes. Owner boundary: Plumb canonical Ship, Images job image, Actions runner lifecycle, Hardrig substrate. This issue is not a waiver for required stable distribution targets.

Observed canonical Ship run 1651 (run ID 4795): https://git.perish.top/PerishLab/plumb/actions/runs/1651 . Marker v0.37.38-beta.2 at 128070024402c668ee94f3387b249c5af6982f77; controller configuration v0.37.38-beta.1. Resolve task 11653 on runner 538ee449-eb0e-4dea-ad00-83779dfe8f82 (v12.13.2) has exposed only docker pull since 2026-09-08 13:03:31 UTC. Image is the accepted immutable Images forge digest e07cd20edac134a9e76742c9d4562459b7372f4cdb229c5e4fb0303499f1cb73, linux/amd64, 12 compressed layers totaling 663075742 bytes. No workload/platform job has started; downstream blocked states are not producer failures. Registry manifests read back successfully. Read-only Forgejo service logs show several blob responses completed 200 in 108365/161571/21196 ms and later blob requests, but there is no completed pull or job progress after over 20 minutes. A subsequent roughly 20-second service network sample transmitted only 46556 bytes (aggregate pod evidence, not per-request throughput). Root cause is not established: transport, client backpressure, extraction and runner lifecycle require separation. No runner/cluster mutation, cancellation, marker rewrite, additional dispatch or stable activation was performed. Track actual producer bootstrap evidence and bounded/progressive diagnostics; retry the same marker only after cause is established or the current run concludes. Owner boundary: Plumb canonical Ship, Images job image, Actions runner lifecycle, Hardrig substrate. This issue is not a waiver for required stable distribution targets.
Author
Owner

Read-only transport evidence: Forgejo pod 10.42.0.205 has three established port-3000 sockets to ingress pod 10.42.0.199 with roughly 1.6–4.1 MB transmit queues and timer_active=4. The ingress is the existing Traefik pod; its matching upstream sockets retain roughly 1.6–6.1 MB receive queues. Traefik also has external entrypoint sockets with roughly 1.8–2.2 MB transmit queues and retransmit timers, some with long RTO values. Linux documents timer_active=4 as the zero-window probe timer and 1 as retransmit: https://docs.kernel.org/networking/proc_net_tcp.html . This narrows the observation to downstream transport/backpressure, not absent Cargo/Node tools or failed producer jobs; it does not yet prove the failing host, route, MTU or network policy. No service or host was changed. Current run remains active.

Read-only transport evidence: Forgejo pod 10.42.0.205 has three established port-3000 sockets to ingress pod 10.42.0.199 with roughly 1.6–4.1 MB transmit queues and timer_active=4. The ingress is the existing Traefik pod; its matching upstream sockets retain roughly 1.6–6.1 MB receive queues. Traefik also has external entrypoint sockets with roughly 1.8–2.2 MB transmit queues and retransmit timers, some with long RTO values. Linux documents timer_active=4 as the zero-window probe timer and 1 as retransmit: https://docs.kernel.org/networking/proc_net_tcp.html . This narrows the observation to downstream transport/backpressure, not absent Cargo/Node tools or failed producer jobs; it does not yet prove the failing host, route, MTU or network policy. No service or host was changed. Current run remains active.
Author
Owner

Hardrig read-only status confirms zxiyun.us-02 is the Forgejo substrate (node ser823577842771); doctor reports only the existing SSH password/root exposure, which is not established as related and was not changed. Read-only ss from the validated existing Traefik network namespace shows the three backlogged external connections with about 0.3–2.0 MB retransmitted, sampled RTT about 4–35 seconds, RTO 120 seconds and roughly 1.8–2.2 MB send queues. This is concrete degraded transport evidence, not a missing compiler. Existing managed SSH identity/pin was used; no trust, network, host or service changes. Keep precise route/loss/MTU/root-cause unresolved. The same-marker standard local Ship backend is now being used to advance host-compatible work and its common inventory while canonical resolve remains in image pull; this does not bypass missing platform or stable-completion evidence.

Hardrig read-only status confirms zxiyun.us-02 is the Forgejo substrate (node ser823577842771); doctor reports only the existing SSH password/root exposure, which is not established as related and was not changed. Read-only ss from the validated existing Traefik network namespace shows the three backlogged external connections with about 0.3–2.0 MB retransmitted, sampled RTT about 4–35 seconds, RTO 120 seconds and roughly 1.8–2.2 MB send queues. This is concrete degraded transport evidence, not a missing compiler. Existing managed SSH identity/pin was used; no trust, network, host or service changes. Keep precise route/loss/MTU/root-cause unresolved. The same-marker standard local Ship backend is now being used to advance host-compatible work and its common inventory while canonical resolve remains in image pull; this does not bypass missing platform or stable-completion evidence.
Author
Owner

Update: controller Product Profile bypass was separately fixed and landed in PR790. New v4 marker v0.37.38-beta.3 binds configuration 96e0037dea9667e7b79476c877951818065b65c22f7a705e41486e1e4cbd027a and profile d01b76d2963f044e6f80f0fa31b007bde9f0cb0072ba9703e70af0c1823371bc. Canonical Ship run 4796 / number 1652 uses real atom 9af8f5fab7; same runner UUID is currently pulling the same image before any producer allocation. Old run 1651 remains active after more than an hour. Read-only Forgejo logs show the 23,036,789-byte layer 507fc65b... completed at 14:01:14 UTC after 3,091,153.7 ms, followed by HTTP 206 at 14:01:52; this is severe transfer degradation, not proof of absent credentials or failed Windows/macOS jobs. No runtime network/host/runner changes made. Local Ship correctly reached tool preflight and refused missing sccache; the caller host also has a different Rust toolchain, and its GLIBC_2.39 binary cannot simply be mounted into the canonical Bookworm image. We are not relaxing production contracts or counting dry-run as publication evidence.

Update: controller Product Profile bypass was separately fixed and landed in PR790. New v4 marker v0.37.38-beta.3 binds configuration 96e0037dea9667e7b79476c877951818065b65c22f7a705e41486e1e4cbd027a and profile d01b76d2963f044e6f80f0fa31b007bde9f0cb0072ba9703e70af0c1823371bc. Canonical Ship run 4796 / number 1652 uses real atom 9af8f5fab798a4aebd4e69745cb163e2c44dbd6e; same runner UUID is currently pulling the same image before any producer allocation. Old run 1651 remains active after more than an hour. Read-only Forgejo logs show the 23,036,789-byte layer 507fc65b... completed at 14:01:14 UTC after 3,091,153.7 ms, followed by HTTP 206 at 14:01:52; this is severe transfer degradation, not proof of absent credentials or failed Windows/macOS jobs. No runtime network/host/runner changes made. Local Ship correctly reached tool preflight and refused missing sccache; the caller host also has a different Rust toolchain, and its GLIBC_2.39 binary cannot simply be mounted into the canonical Bookworm image. We are not relaxing production contracts or counting dry-run as publication evidence.
Sign in to join this conversation.
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
PerishLab/plumb#789
No description provided.