Skip to content

The first harness cells

The harness started four cells on 2026-10-08. It measured each cell 11 times over two fresh environments. The target was a stock Redis 7.4.11 container.

The timers were scaled. Keepalive used an idle time of 5 s and 3 probes 2 s apart, and the kernel clock was HZ 250. The prober measured each interval from the records of the kernel.

  • Prediction
  • Measured runs
  • Median
Q1 idle, black hole✓ holds
The prediction: 10.996 s to 11.352 sThe median: 11.156 sRun 1, environment 1: 11.156 sRun 2, environment 1: 11.132 sRun 3, environment 1: 11.269 sRun 4, environment 1: 11.192 sRun 5, environment 1: 11.188 sRun 6, environment 1: 11.14 sRun 7, environment 2: 11.062 sRun 8, environment 2: 11.054 sRun 9, environment 2: 11.172 sRun 10, environment 2: 11.235 sRun 11, environment 2: 11.156 s11 s11.1 s11.2 s11.3 s

11 of 11 runs inside [10.99, 11.358] s; median 11.156039 s, its interval [11.061692, 11.234689] s

Q2 in flight, black hole, tcp_retries2 = 3✓ holds
The prediction: 3 s to 6.616 sThe median: 6.34 sRun 1, environment 1: 6.236 sRun 2, environment 1: 6.42 sRun 3, environment 1: 6.408 sRun 4, environment 1: 6.3 sRun 5, environment 1: 6.416 sRun 6, environment 1: 6.44 sRun 7, environment 2: 6.34 sRun 8, environment 2: 6.312 sRun 9, environment 2: 6.212 sRun 10, environment 2: 6.312 sRun 11, environment 2: 6.424 s3 s4 s5 s6 s7 s

11 of 11 runs inside [2.994, 6.622] s; median 6.339989 s, its interval [6.236006, 6.423981] s

Q2 in flight, black hole, tcp_retries2 = 5✓ holds
The prediction: 12.6 s to 26.52 sThe median: 12.92 sRun 1, environment 1: 12.92 sRun 2, environment 1: 13.076 sRun 3, environment 1: 12.872 sRun 4, environment 1: 13.012 sRun 5, environment 1: 13.032 sRun 6, environment 1: 12.912 sRun 7, environment 2: 12.864 sRun 8, environment 2: 12.936 sRun 9, environment 2: 12.944 sRun 10, environment 2: 12.892 sRun 11, environment 2: 12.86 s15 s20 s25 s

11 of 11 runs inside [12.594, 26.526] s; median 12.920007 s, its interval [12.863958, 13.032021] s

Q3 in flight, frozen process✓ holds

The algebra predicts that no TCP timer fires. The harness measured silence in 11 of 11 runs, over a window of 20 s. In each run the frozen kernel sent 5 acknowledgments.

  • Q1, an idle connection into a black hole. Keepalive found the failure in each run. The probes went out 2.016 s apart, which is 16 ms late for each step.
  • Q2, a request in flight. The retransmission give-up found the failure. At tcp_retries2 = 5, the keepalive time is shorter than the budget. Thus the cell also proves that the first retransmission stops keepalive.
  • Q3, a frozen process. No TCP timer fired in 20 s. The frozen kernel acknowledged the request, and it answered each keepalive probe.
  • An error of the harness. The harness measured from the tail loss probe. The budget of the give-up starts one RTO later, at the first retransmission of the RTO timer.
  • An error of the algebra. The band for tcp_retries2 = 3 ended at 6.2 s, and the runs came at 6.21 s to 6.42 s. The algebra did not have the real RTO or the rounding of the timer wheel. The correction added both.

The client also read the view of the kernel through TCP_INFO, every 2 ms. That view agrees with the eBPF records to within 0.19 ms on keepalive, and to within 4 ms on the retransmissions.